ContextGraph analytics — centrality, community detection, Node2Vec embeddings, link prediction, and structural similarity — answer what your graph means, not just what it contains. Enable them once with advanced_analytics=True and use them to rank the most structurally important nodes, surface hidden operational clusters, and flag implied connections before they are formally observed.
What Is Graph Analytics?
Graph analytics applies mathematical algorithms to discover patterns and properties in your knowledge graph’s structure. It goes beyond storing and retrieving data to analyze the relationships themselves. Graph analytics vs. graph traversal: Traversal follows existing edges to find connected nodes. Analytics examines the entire graph structure to find patterns — which nodes are most influential, which groups of nodes form communities, which connections are missing. Graph analytics vs. reasoning: Reasoning applies logical rules to derive new facts. Analytics applies statistical and topological algorithms to measure structural properties like centrality, clustering, and similarity. Graph analytics helps you understand the shape and importance patterns within your data, revealing insights that aren’t apparent from individual nodes or edges.Why Use Graph Analytics?
Finding influential entities. Not all nodes are equally important. Analytics identifies which entities are most central, most connected, or most strategically positioned in the network structure. Community discovery. Analytics reveals hidden clusters and operational groups by analyzing connection density patterns. Entities that cluster together often share purposes, origins, or behaviors not obvious from metadata alone. Relationship discovery. Link prediction identifies probable connections that haven’t been explicitly observed yet, helping investigators focus on the most likely missing relationships. Investigation support. Analytics provides objective measures of importance and relatedness, helping analysts prioritize which entities to investigate first and which relationships warrant deeper examination. Risk identification. Centrality measures identify entities whose removal would most disrupt the network — useful for understanding single points of failure, key infrastructure, or high-impact targets.When To Use / When Not To Use
Use graph analytics for:- Large graphs (100+ nodes) where patterns aren’t obvious from inspection
- Prioritizing investigation efforts based on structural importance
- Discovering hidden communities and operational clusters
- Understanding network resilience and vulnerability points
- Identifying missing relationships through link prediction
- Following known relationships between specific entities
- Exploring neighborhoods around particular nodes
- Path-finding between known entities
- Applying known logical rules to derive new facts
- Policy enforcement and rule-based decisions
- Situations where relationships follow clear logical patterns
- Small graphs where patterns are visually obvious
- Direct lookups of specific entities or relationships
- Cases where you know exactly what you’re looking for
- You need to understand the overall structure and patterns
- Manual inspection would miss important structural properties
- You want to discover unexpected relationships or communities
- Objective measures of importance or similarity would guide decisions
Key Analytics Concepts
Modularity measures how well-defined communities are in your graph. High modularity (>0.4) means the detected communities have many internal connections and few external ones — indicating real organizational structure. Louvain community detection finds groups of nodes that are more densely connected to each other than to the rest of the graph. It’s useful for discovering operational clusters, organizational units, or functional groups. Node2Vec creates vector embeddings for nodes by simulating random walks through the graph. Nodes that appear in similar contexts during these walks end up with similar vectors, even if they’re not directly connected. Centrality metrics measure different aspects of node importance:- Degree: how many direct connections
- Betweenness: how often a node sits on shortest paths between others
- Eigenvector: importance based on having important neighbors
- PageRank: importance in a random walk through the graph
All analytics require
advanced_analytics=True at construction time. Without it, every method in this guide raises RuntimeError: advanced_analytics not enabled. The flag lazy-initializes five sub-components: CentralityCalculator, CommunityDetector, NodeEmbedder, LinkPredictor, and SimilarityCalculator.Setting Up the Analytical Graph
The One-Call Snapshot
Before diving into individual analyses, get the lay of the land first.analyze_graph_with_kg() runs every analytics pass in sequence and returns a single unified report:
Finding the Kingpins — Centrality Analysis
Not all nodes are equal. Centrality analysis gives each node five different scores, each measuring a different kind of importance. The key intuition: a node can have low degree (few direct connections) but extreme betweenness (everything passes through it). That’s a broker — the node that, if removed, would fragment the graph into disconnected pieces. In threat intelligence, brokers are often infrastructure nodes: bulletproof hosting providers, shared C2 frameworks, or intermediary loaders that every campaign routes through.Scoring a Single Node
When you suspect a particular entity is significant, start with a direct score:Bulk Rankings Across the Full Graph
To rank every node across all five measures at once:Which Measure to Trust?
For attribution analysis, eigenvector centrality often surfaces the most meaningful nodes — high eigenvector means you’re connected to other high-eigenvector nodes, which in threat intelligence tracks with “known-to-be-important” entities clustering together.
Mapping Threat Actor Clusters — Community Detection
Centrality tells you about individual nodes. Community detection tells you about groups — which nodes form tightly-knit clusters with more internal connections than external ones. Imagine your graph has twenty distinct threat actor nodes. Centrality can rank them individually, but it can’t tell you that twelve of them actually share infrastructure, tooling patterns, and target sectors in a way that makes them a single operational cluster, while the remaining eight split into two separate groups. Community detection finds that automatically.COZY BEAR, VENOMOUS BEAR, and FANCY BEAR all appear in the same community because they share C2 infrastructure — despite being traditionally attributed to different GRU units.
Visualize immediately with:
Teaching the Graph to Measure Distance — Node2Vec Embeddings
Centrality and community detection are topological — they work with edges as binary connections. Node2Vec embeddings go deeper: they generate a dense vector for each node by simulating thousands of random walks through the graph, learning which nodes tend to appear in similar neighborhoods. The practical result: nodes that play similar structural roles end up near each other in embedding space — even if they have no direct edge between them. Two C2 domains that both sit between threat actors and victim infrastructure will have similar embeddings, even if they’re in completely different parts of the graph.Decision nodes as node2vec_embedding and used automatically by find_precedents_by_scenario() as the structural similarity component.
What Does This Attack Pattern Remind Me Of?
Once you have embeddings, you can ask: “which nodes in my graph are structurally most similar to APT29?” This isn’t about shared edges — it’s about shared role. Which other nodes play the same position in the graph that APT29 plays?LAZARUS and APT38 both appear at the top with similarity > 0.85, that’s worth noting in your analyst report: these groups are playing structurally identical roles in the graph, which may warrant a unified tracking hypothesis even if traditional attribution keeps them separate.
To compare two nodes directly:
Anticipating the Next Move — Link Prediction
Link prediction answers a different question: not “which nodes are similar” but “which edges are missing?” The graph may lack a direct connection between two nodes not because the connection doesn’t exist, but because you haven’t observed it yet. In practice: you haveAPT29 → uses → SUNBURST and SUNBURST → targets → Windows Server 2019. Link prediction might surface APT29 → targets → Windows Server 2019 with a high score — the connection is implied by transitivity but hasn’t been explicitly drawn.
Link prediction is also available on
Decision nodes through DecisionQuery.predict_decision_relationships(decision_id, top_k). See the Decision Intelligence guide for how to surface causal relationships between past decisions.Understanding Your Decision History
When decisions are stored as nodes in the graph,get_decision_insights() gives you an analytical view across all of them:
"advanced_analytics" key inside insights contains the full analyze_graph_with_kg() result — so calling get_decision_insights() gives you the complete analytical picture without an extra call.
Domain Examples
- Defense — CTI Threat Network
- Security — Active Incident
- Life Science — Clinical Trial Graph
- Banking — Counterparty Risk
Your SOC just merged three threat intel feeds into a single CTI graph. Before this week’s analyst briefing, you need to rank the most dangerous actors, identify hidden operational clusters, and flag implied connections that haven’t been formally tracked.
Common Pitfalls
Treating predictions as facts. Link prediction and similarity scores are probabilistic estimates, not confirmed relationships. A high link prediction score suggests a probable connection but requires human verification before acting on it. Duplicate entities. Having “APT-29”, “APT29”, and “Cozy Bear” as separate nodes artificially reduces their centrality scores and fragments communities. Deduplicate entities before running analytics for accurate results. Inconsistent naming. Mixing “ThreatActor”, “threat_actor”, and “Actor” as node types breaks analytics that group by node type. Use consistent naming conventions across your data sources. Over-interpreting analytics results. A node with high betweenness centrality is structurally important in your current graph — not necessarily important in the real world. Analytics reveals patterns in your data, not universal truths about the domain. Running analytics on graphs that are too small. Community detection and centrality measures are most meaningful on graphs with 100+ nodes and adequate connection density. Results on graphs below that threshold may not provide reliable insights. Poor graph quality. Duplicate entities, inconsistent naming, and missing relationships directly impact analytics accuracy. Deduplicate and normalise your data before running analytics — the algorithms amplify whatever structure exists in your graph, including noise.What the Numbers Mean
Related Guides
- Context Graphs — building and querying the underlying
ContextGraph - Visualization — render centrality rankings and community clusters as interactive dashboards
- Decision Intelligence — link prediction and structural similarity applied to decision nodes
- GraphRAG — using analytics results to ground LLM generation in the most contextually relevant subgraph
