OntologyGenerator derives a formal OWL ontology directly from the entities and relationships already in your knowledge graph — no schema design upfront. Use it to produce a machine-readable contract for your graph’s classes and properties, then export to Turtle, OWL/XML, or JSON-LD for SHACL validation, reasoning engines, and STIX/TAXII toolchains.
What Is an Ontology?
An ontology is a formal specification of the concepts and relationships in a domain. It defines a shared vocabulary for your knowledge graph by declaring: Classes — Types of entities in your domain. For example,ThreatActor, Vulnerability, or Software in cybersecurity. Classes describe what kinds of things exist.
Object Properties — Relationships between entities. Examples: exploits (Threat Actor exploits Vulnerability), targets (Malware targets Organization), or uses (Actor uses Tool).
Datatype Properties — Attributes with literal values. Examples: name (text), severity_score (decimal), published_date (date), or ip_address (string).
An ontology makes your knowledge graph machine-readable by specifying which relationships are valid, what types of values each property can hold, and how concepts relate to each other hierarchically.
Why Use Ontologies?
Consistency. Without an ontology, one team might use “Threat_Actor” while another uses “ThreatActor” for the same concept. Ontologies enforce a single naming convention. Validation. Ontologies enable automatic checking — does this relationship make sense? Should a Vulnerability have a CVSS score as text or number? Reasoning. Exporting to OWL/Turtle lets external reasoning engines (HermiT, Pellet, ELK) infer new facts. If Malware subclasses Software and HAMMERTOSS is Malware, an OWL reasoner concludes HAMMERTOSS is also Software. Semantica exports the ontology; reasoning itself runs in the external tool. Knowledge Graph Quality. Structured schemas catch errors early and ensure new data integrates cleanly with existing entities.When To Use / When Not To Use
Use ontologies for:- Formal knowledge graphs with complex entity relationships
- Data integration across multiple teams or organizations
- Automated reasoning and rule-based systems
- Long-term knowledge bases that need consistency over time
- Integration with external tools that expect OWL/RDF schemas
- Simple semantic search over text documents
- Lightweight RAG where vector similarity is sufficient
- Prototype development where the schema changes rapidly
- Single-use data analysis that doesn’t need reusability
- Cases where the overhead exceeds the complexity of your domain
Semantica’s ontology module derives formal OWL ontologies directly from entities and relationships already in your knowledge graph — no schema design upfront. A 6-stage pipeline infers classes, builds hierarchies, maps OWL types, and serialises to Turtle. The pipeline runs in memory; you do not need a running triple store.
A Simple Example
Start with a familiar business domain to understand the mechanics before diving into cybersecurity:https://company.example.org/ontology/) becomes the namespace prefix for all classes and properties. Person becomes <https://company.example.org/ontology/Person> in the exported RDF.
Object properties connect entities to other entities (works_for, reports_to). Datatype properties connect entities to literal values (name).
The graph that has no schema
Now with a schema understanding, populate a knowledge graph with CTI data to make the pipeline mechanics visible.Generating the ontology
OntologyGenerator reads your graph dict and runs the 6-stage pipeline.
Malware shows the pipeline detected that malware is a sub-type of software based on co-occurrence patterns in the relationship graph. You can override these inferences manually before exporting.
Validating the schema
Run structural validation before exporting.ClassInferrer and PropertyGenerator before the next export cycle.
Growing the ontology as new node types emerge
UseClassInferrer to add new entity types incrementally without regenerating the full ontology.
infer_classes on each new batch, review the results, adjust parent assignments where the pipeline guessed wrong, and merge. The ontology grows with the graph rather than falling behind it.
Generating an ontology from unstructured text
LLMOntologyGenerator extracts classes and properties from prose using an LLM — useful for bootstrapping when no structured graph exists yet.
LLMOntologyGenerator is best for bootstrapping a new domain where no structured graph exists yet. Once you have a graph, prefer OntologyGenerator.generate_from_graph() — it is deterministic, reproducible, and does not consume LLM tokens on every run.Exporting for downstream systems
Export in the format your downstream tools expect.Drafting, Reviewing, and Publishing Ontology Changes
Editing a live ontology is a heavier change than editing graph data. Other systems have already built against those class and property names, so a rename ripples outward. Explorer exposes the draft/proposal flow over HTTP so a change is staged and reviewed before it reaches the graph. The overhead is worth it once more than one person or agent edits the same ontology. While you are still prototyping, regenerating the ontology is usually cheaper than reviewing a diff. The state machinepublish answers 400 for any state other than approved, so a proposal cannot reach the graph without an explicit approval.
The flow
impact_analysis— a structured diff fromVersionManager.diff_ontologies, plusclass_adds,class_removals,property_changesandrestriction_changescounts.shacl_validation— current graph data validated against shapes generated from the proposed ontology. It reads{"status": "skipped", "reason": "No store configured"}when the session has no store, and{"status": "error", ...}when validation itself failed. Neither stops the proposal from being created; both are advisory.
VersionManager.create_version() before it touches the live graph. If version creation fails, the request answers 500 and nothing has been added.
It applies additions only. added_classes and added_properties become owl:Class and owl:ObjectProperty nodes; removals and modifications are recorded in the version history but do not delete or rewrite anything already in the graph.
Version history and comparison
compare returns class_changes, property_changes, restriction_changes and axiom_changes as separate blocks, so you do not have to diff the whole ontology to see what moved.
Common Pitfalls
Over-modeling. Don’t create 50 classes when 10 would suffice. Start simple and add complexity only when you need formal distinctions for reasoning or validation. Having separate classes forMaliciousEmail and PhishingEmail is only useful if they have different properties or relationships.
Ontology drift. When new entity types appear in your graph, the ontology becomes stale unless you regenerate or incrementally update it. Set up monitoring to detect when new entity types appear that aren’t covered by your current ontology.
Inconsistent class naming. Pick a convention (CamelCase, snake_case, or kebab-case) and stick to it. Mixing ThreatActor, threat_actor, and threat-actor in the same ontology creates confusion and breaks tooling that expects consistent naming patterns.
In-memory state. The flow keeps its state on the application object rather than in a store, so a restart loses every draft, proposal and version record, and the state is not shared across worker processes. Anything you intend to keep has to be published before the process restarts.
Unencoded # in a path segment. /api/ontology/drafts/{uri} and /api/ontology/versions/{uri} take the ontology URI as a path segment, and # starts a URL fragment there, so the server only sees the part before it. Encode it as %23, as in the version examples above. OWL ontologies commonly end in #, and the failure is silent: the endpoint answers 200 with an empty list rather than reporting a truncated URI. URIs without a # are unaffected.
Domain Examples
- Defense — CTI/Threat
- Security — SOC/Incident
- Life Science — Clinical/Pharma
- Banking — Risk/Compliance
A defense CTI team ingests raw OSINT reports each morning. The ontology must stay interoperable with STIX 2.1 and the NATO MISP taxonomy, so IRIs follow the DoD namespace and the ontology is exported to OWL/XML for the org’s SIEM reasoning plugin.
Related Guides
- SHACL Validation — generate W3C SHACL constraint shapes from your ontology and validate live graph data against them
- Reasoning & Rules — apply forward/backward-chaining rules over your ontology to derive new facts
- Export & Serialization — export graphs to RDF, GraphML, CSV, and Neo4j Cypher
- Semantic Extraction — extract entities and relationships that feed ontology generation
- Context Graphs — the knowledge graph that ontology generation reads from
