Configure Models
Zyrabit SLM allows configuring multiple models for text generation, embeddings, and audio transcription.
Ports Overview
from typing import Protocol, AsyncIterator
class InferencePort(Protocol):
def generate(self, request: dict) -> dict:
...
class StreamingInferencePort(Protocol):
async def stream_generate(self, request: dict) -> AsyncIterator[str]:
...
Creating a New Inference Adapter
To switch to a different provider (e.g., an OpenAI-compatible endpoint):
- Create Adapter: Create
app/infrastructure/inference/openai_adapter.py. - Implement Interface: Inherit from
InferencePortand/orStreamingInferencePort. - Configure: Update
app/inference_factory.pyto return your new adapter whenINFERENCE_PROVIDER=openai.
[!TIP] You can mix and match providers. For example, use local Ollama for embeddings and cloud Anthropic for chat generation.