.generate(), .generate_structured() and .generate_typed() methods — across Groq, OpenAI, Anthropic Claude, HuggingFace, Novita AI, and 100+ providers via LiteLLM. It also provides a deliberately separate decision-only interface for TypeSafe Jev, whose typed Choice, Noul, and Score results are not text generations.
What Are LLM Integrations?
Thesemantica.llms module provides a unified interface for connecting to Large Language Model providers. Instead of learning different APIs for each generative provider, you use the same methods (.generate(), .generate_structured(), .generate_typed()) regardless of whether you’re calling Groq, OpenAI, Anthropic, or local HuggingFace models. TypeSafe Jev is the decision-only exception: it exposes .decide() because it returns typed decisions rather than generated text.
Unified interface across generative providers: Generative LLM providers in Semantica expose identical methods, so switching from OpenAI to Anthropic requires changing only the provider constructor, not your application code. Jev remains separate so its calibrated probabilities and strict decision types are not flattened into a chat-completion response.
Provider wrappers vs semantic extraction provider strings: The semantica.llms classes (Groq, OpenAI, LiteLLM, HuggingFaceLLM) are Python objects for text generation. The semantica.semantic_extract module accepts provider names as strings for entity and relationship extraction. Both approaches are covered in this guide.
Why Use LLM Integrations?
Provider portability. Test with one provider, deploy with another. Switch from Groq for prototyping to Anthropic for production without code changes. Reduced vendor lock-in. Avoid tying your application to a single LLM provider’s API. If pricing changes or service availability issues arise, switching providers is straightforward. Consistent APIs. Use the same.generate(), .generate_structured(), and .generate_typed() methods across all providers instead of learning provider-specific interfaces.
Multi-provider workflows. Run fast models for initial classification and expensive frontier models for complex reasoning in the same pipeline.
Local vs cloud deployment flexibility. Use cloud providers during development and switch to local HuggingFace models for air-gapped production environments.
When To Use / When Not To Use
Use LLM integrations for:- Text generation, summarization, and question-answering tasks
- Complex reasoning that requires natural language understanding
- Structured data extraction from unstructured text
- Multi-step analysis requiring interpretation and synthesis
- Tasks where context, ambiguity, or domain knowledge matter
- Pattern matching that regular expressions can handle
- Simple rule-based classification with clear criteria
- Mathematical calculations or statistical analysis
- Graph traversal and relationship queries
- Data transformations with known logic
- Simple keyword search or exact string matching
- Deterministic workflows with predefined decision trees
- High-frequency, low-latency operations where inference overhead matters
- Tasks where explainability requires transparent rule-based logic
Generative providers in
semantica.llms (Groq, OpenAI, LiteLLM, HuggingFaceLLM) are for text generation and query_with_reasoning(). Jev and AsyncJev are decision-only and are not drop-in llm_provider= values. For structured entity and relation extraction, semantica.semantic_extract accepts provider names as strings.Choosing a Provider
Four factors drive provider selection, each optimized for different use cases: Latency matters most in real-time SOC triage loops where an analyst is waiting on a triage verdict. Groq’s inference infrastructure typically returns 8B model responses in under 300ms, making it ideal for interactive workflows. Accuracy matters most in high-stakes decisions: clinical contraindication checks, credit committee reasoning, and legal document analysis. Frontier models like Claude or GPT-4 available throughLiteLLM provide the strongest reasoning capabilities.
Data residency constraints eliminate cloud providers for classified or HIPAA-regulated workloads. HuggingFaceLLM with local model paths, or Ollama pointed at a local server, both enable fully air-gapped deployments without network calls.
Cost at scale favors high-throughput providers like Novita AI for bulk extraction pipelines processing thousands of documents per hour where per-token costs accumulate quickly.
The unified interface means you can prototype with Groq for speed, validate accuracy with Claude, and deploy to Azure OpenAI for compliance — without changing application code.
The Shared Generation Interface
Every generative provider exposes the same methods:generate() returns a plain string. generate_structured() instructs the model to respond in JSON and returns the parsed result — a dict for a top-level JSON object, or a list if the model returns a top-level JSON array. generate_typed() takes a Pydantic model, validates the model’s output against it, and retries up to max_retries times with the validation error fed back into the prompt — reach for it when downstream code needs a guaranteed shape rather than best-effort JSON. is_available() lets you health-check the provider before committing to a call — useful in retry logic and warm-up checks.
This means every place in Semantica that accepts a generative LLM — query_with_reasoning(), semantic extraction, custom reasoning loops — accepts any of these providers interchangeably. Jev is not part of this interface because it never returns free text.
Jev — Typed Decisions for Low-Latency Routing
TypeSafe Jev is a System One model for typed decisions over application state. Use it for bounded routing, classification, binary checks, or scoring when the valid outcomes are known in advance. Use a chat LLM instead when the task needs an explanation, synthesis, open-ended text, or multi-step reasoning.TYPESAFE_API_KEY, or pass api_key= explicitly:
Semantica owns the reasoning, policy checks, escalation, and provenance around the decision. The provider does not silently call a fallback model:
AsyncJev with the same arguments and result type:
Groq — Fast Inference for Real-Time Agents
Groq is a cloud provider that specializes in ultra-fast language model inference using custom hardware called Language Processing Units (LPUs). Their infrastructure delivers sub-300ms response times for smaller models, making them ideal for real-time applications where speed matters more than maximum reasoning capability. Groq Cloud runs open models on purpose-built Language Processing Units that deliver sub-300ms latency for 8B parameter models. This makes Groq the right default for any agent loop where the LLM is in the hot path — SOC triage, real-time alert classification, conversational agents.llama-3.1-8b-instant for anything in the hot path, llama-3.3-70b-versatile when you need stronger reasoning but can afford slightly higher latency, mixtral-8x7b-32768 for long-context summarisation tasks.
OpenAI — Function Calling and Vision
OpenAI provides access to the GPT model family, including GPT-4o with advanced capabilities like function calling (structured tool use) and vision processing for images and documents. OpenAI models are well-suited for complex reasoning tasks that require strong language understanding and generation capabilities. TheOpenAI provider wraps the OpenAI API. Use it when you need GPT-4o’s function-calling precision, vision capabilities for document screenshots, or when your team already has an OpenAI contract and wants to stay there.
gpt-3.5-turbo is fine for classification and light extraction. Switch to gpt-4o for complex multi-step regulatory reasoning or document understanding.
Anthropic — Complex Reasoning and Structured Extraction
Anthropic provides the Claude model family, built with an emphasis on careful, instruction-following behavior and strong performance on multi-step reasoning, long-document analysis, and code-related tasks. Claude models tend to be more cautious about ambiguous instructions than other providers. That matters when the cost of a confidently wrong answer is high. TheAnthropic provider wraps the Claude API. Reach for it when the task involves reasoning through several dependent steps (not just single-turn extraction), when you’re processing long source documents that need to stay in context, or when you need schema-validated structured output rather than best-effort JSON.
Install with pip install "semantica[llm-anthropic]" (or just pip install anthropic) before using this provider.
Gemini — Long Context and Multimodal Input
Gemini is Google’s model family, with a context window large enough to hold entire codebases or long regulatory filings in a single call, and native support for image and document input alongside text. Reach for it when a task needs to reference a large amount of source material at once, or when the input isn’t plain text. TheGemini provider tries the newer google-genai SDK first and falls back to the older google-generativeai package if that’s what’s installed. Install with pip install "semantica[llm-gemini]" (or pip install google-genai) before using this provider.
Ollama — Local, Air-Gapped Inference
Ollama runs models entirely on your own machine, with no API key and no outbound network call. It’s the right choice for air-gapped environments, offline development, or any workload where the source data can’t leave the local network. Unlike the other providers here,Ollama takes a base_url instead of an api_key. It talks to a local Ollama server over HTTP. Start the server with ollama serve and pull a model with ollama pull llama2 before using this provider. Install the Python client with pip install "semantica[llm-ollama]" (or pip install ollama).
is_available() for Ollama does a real connectivity check (it calls the server’s list() endpoint), unlike the API-key-based providers above, so a False here usually means the server isn’t running rather than a missing credential.
DeepSeek — Budget Reasoning at Scale
DeepSeek exposes an OpenAI-compatible API at a fraction of the cost of the larger US providers, with reasoning quality that holds up well for extraction and classification work. It’s a reasonable default when you’re processing a large volume of documents and don’t need the deepest reasoning tier. Install withpip install "semantica[llm-deepseek]" (or pip install openai, since DeepSeek is accessed through the OpenAI client pointed at a different base URL).
LiteLLM — One Interface, 100+ Providers
LiteLLM is a universal adapter that provides a single interface to over 100 different LLM providers, including Anthropic Claude, Azure OpenAI, AWS Bedrock, Google Vertex AI, and local Ollama instances. It acts as a translation layer, converting your unified API calls into provider-specific requests, enabling easy switching between providers without code changes.LiteLLM is the Swiss Army knife. It wraps the litellm library, which speaks to every major provider using a unified completion API. The model string encodes both provider and model name: "anthropic/claude-sonnet-5", "azure/gpt-4o", "bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0", "ollama/llama3.2". Change the string, change the provider — no other code changes needed.
ANTHROPIC_API_KEY, AZURE_API_KEY, OPENAI_API_KEY, GOOGLE_APPLICATION_CREDENTIALS, etc. LiteLLM picks them up automatically. For a provider switch driven by deployment environment, you can keep provider selection in a config dict and inject it at startup — no if/else branches in application code:
HuggingFaceLLM — Air-Gapped and On-Premise
HuggingFaceLLM provides access to open-source models from the HuggingFace ecosystem, either downloaded from the HuggingFace Hub or loaded from local file paths. This is the only option for completely offline deployments where no network access is available during inference, such as classified environments or air-gapped systems.HuggingFaceLLM loads a model from the HuggingFace Hub or from a local directory path. No network calls during inference. This is the only option for classified environments, HIPAA-constrained clinical deployments, and any network segment without outbound internet access.
HF_TOKEN in your environment for Hub access to gated or private models. For local paths, no token is needed — the model directory must contain the standard HuggingFace checkpoint files.
Swapping Providers Without Changing Application Code
The real payoff of the unified interface is inquery_with_reasoning(). This is the call that drives graph-grounded reasoning in every AgentContext. Because it accepts any provider object, you can slot in a different LLM at any tier of your pipeline with zero changes to the surrounding code.
context, the graph, the retrieval logic — none of it changes. Only the llm_provider argument differs between the two tiers.
Using Providers for Semantic Extraction
For NER, relation extraction, and triplet extraction, Semantica’ssemantica.semantic_extract module accepts provider names as strings rather than class instances. The module handles provider instantiation internally.
Novita AI — Cost-Efficient Bulk Extraction
Novita AI exposes an OpenAI-compatible API at low per-call cost, making it a reasonable choice for high-volume NER pipelines where cost matters more than getting the single best answer. Install withpip install "semantica[llm-novita]" (or pip install openai, since Novita is accessed through the OpenAI client pointed at a different base URL).
Novita class directly:
Domain Examples
- Defense — CTI/Threat
- Security — SOC/Incident
- Life Science — Clinical/Pharma
- Banking — Risk/Compliance
A classified threat-intelligence analysis unit needs to run entirely air-gapped — no outbound network traffic of any kind. The extraction model and the reasoning model both load from a local NFS share. The knowledge graph accumulates over the analysis session; all inference happens on-premise.
Common Pitfalls
Choosing expensive frontier models for simple extraction tasks. GPT-4o or Claude Sonnet for basic entity extraction is overkill — Groq’s Llama models handle straightforward NER and classification at a fraction of the cost and latency. Reserve frontier models for complex reasoning that requires nuanced interpretation. Ignoring latency differences between providers. Groq typically responds in under 300ms, while Anthropic Claude can take 2-3 seconds for the same query. For real-time agents or interactive workflows, latency differences compound across multiple LLM calls. Profile your provider performance under realistic load. Using LLMs for deterministic pattern matching that regex can handle. If your task is extracting email addresses, phone numbers, or other pattern-based entities, regular expressions are faster, cheaper, and more reliable than LLM extraction. Use LLMs when context, ambiguity, or domain knowledge matter for correct interpretation. Not validating structured outputs. Thegenerate_structured() method returns parsed JSON (a dict, or a list for a top-level array), but LLMs can still produce malformed or incomplete structures. Validate the result against your expected schema before using it downstream — or use generate_typed(), which validates against a Pydantic model for you.
Treating a Noul probability as routing confidence. Noul returns the probability of yes, so both 0.01 and 0.99 are highly certain while 0.5 is maximally uncertain. Use result.probability for the raw yes-probability and result.confidence for Semantica’s derived routing certainty.
Using Jev when the output must explain itself. Jev makes bounded typed decisions and does not generate reasoning text. Build the evidence and policy context in Semantica, use a reasoning LLM or human review when an explanation is required, and record both steps in provenance.
Switching providers without testing prompt behavior. Different models respond differently to the same prompt. A prompt optimized for GPT-4 may produce poor results with Llama or Claude. When switching providers, test your prompts and adjust temperature, instructions, or examples as needed.
Overusing local HuggingFace models for tasks requiring latest knowledge. Local models have a knowledge cutoff from their training date and cannot access current information. For tasks requiring up-to-date knowledge (recent CVEs, current regulations, latest threat intelligence), cloud providers with more recent training data may be necessary.
Related Guides
- Agent Memory — using
query_with_reasoning()with any LLM provider for graph-grounded retrieval - Multi-Agent Systems — wiring different LLM providers to different agent tiers in a shared-graph pipeline
- Semantic Extraction — LLM-powered NER, relation extraction, event detection, and triplet extraction
- GraphRAG — multi-hop graph reasoning with
query_with_reasoning()
