Back to POOJARI's profile

ThreatWeave

From isolated alerts to attack stories. The fundamental innovation: Alert → Evidence → Relationship → Attack chain → Risk → Explanation → Response

POOJARI NAGASHIVA
  • Expert reviewed
ThreatWeave

96.7%

Attack recall (binary detection)

98.3%

ROC-AUC (binary detection)

200

Behavioral campaigns clustered from high-risk flows

Overview

CipherMind AI ’26 was built to address a practical challenge in cybersecurity: detecting network threats is not enough — security analysts also need to understand why an incident was flagged, how severe it is, and which incidents are related.

We developed the system progressively, starting with UNSW-NB15 dataset validation and binary attack detection. We compared Logistic Regression, Random Forest, and CatBoost, with CatBoost achieving the strongest overall performance. We then performed feature-ablation experiments and discovered that TTL-based features could be removed with almost no performance loss, making the final model less dependent on dataset-specific shortcuts. Multiclass attack-family classification was added next, along with uncertainty-aware top-k predictions for attack families that naturally overlap.

The pipeline was then extended with unsupervised anomaly detection, evidence fusion, behavioral correlation, SHAP-based explainability, and a local LLM summary layer. The LLM does not make security decisions or access raw network-flow data; it only converts the evidence already produced by the detection pipeline into an analyst-friendly incident summary.

The final system therefore moves beyond simple classification toward an explainable SOC assistant that can detect, prioritize, correlate, explain, and summarize suspicious network activity.

What I learned

Detection accuracy alone is not enough. A cybersecurity system must also provide evidence that analysts can understand and act upon. Feature leakage and dataset shortcuts matter. The id feature showed strong correlation with the label despite having no meaningful security interpretation, while the TTL-family ablation showed that removing suspicious shortcut signals caused almost no performance degradation. Different models have different strengths. CatBoost provided the best overall binary detection performance, while Random Forest offered slightly higher attack recall. Real-world generalization is harder than validation performance. The validation-to-test performance gap showed that strong validation results do not automatically guarantee equally strong performance on shifted data. Some attack families are intrinsically difficult to distinguish. Analysis, Backdoor, and DoS showed substantial behavioral overlap, so forcing every incident into a confidently stated single class can be misleading. Top-k predictions and ambiguity flags provide a more honest representation of uncertainty. Combining multiple sources of evidence is more useful than relying on one model. Supervised probability, anomaly detection, model agreement, SHAP explanations, and behavioral correlation provide complementary perspectives. Explainability should be evidence-based. SHAP helped identify meaningful behavioral signals such as TCP handshake timing, packet characteristics, services, and connection-context features after TTL features were removed. LLMs are better used as a communication layer than as the detector. Using a local Ollama model keeps the system free and allows the LLM to summarize already-computed evidence without replacing the underlying ML models. A good SOC system should prioritize analyst decision-making. The goal is not simply to output “Attack” or “Normal,” but to answer: What happened? Why was it flagged? How confident are we? What incidents are related? And what should the analyst investigate next?

AI tools used

ClaudeUsed Claude to structure the technical documentation and think through how to clearly explain architecture decisions (model choice, fusion weighting, honest limitations) to non-technical judges.

Links & files

Artifacts

4