CipherMind Sentinel
An AI-powered SOC Copilot detecting network intrusions in real-time using LightGBM, Isolation Forests, and LLM-driven incident correlation.
- Expert reviewed
98.4%
Attack Detection Recall
98.2%
ROC AUC (Receiver Operating Characteristic)
<3ms
End-to-End Inference Latency React, Tailwind CSS, Next.js, Python, TypeScript, Machine Learning, Data Analysis, LightGBM, Socket.io, Bun, Render, Vercel
Overview
CipherMind Sentinel is an AI-powered Security Operations Copilot. It ingests raw network-flow data and pushes it through a layered machine learning architecture to produce a small number of risk-scored, explained, and correlated incidents.
Instead of a black box, Sentinel provides a live SOC interface featuring:
Real-time Replay Streaming: Watch incidents form in real-time as network events stream in. Layered ML Detection: A calibrated binary detector, a multiclass attack classifier, and an unsupervised anomaly detector working in tandem. Explainability Center: Transparent risk scoring backed by SHAP (SHapley Additive exPlanations) so analysts know exactly why an alert fired. AI Analyst Copilot: An LLM-backed chat interface that summarizes incidents and helps analysts investigate, complete with a deterministic, evidence-grounded fallback.
How we built it We split the architecture into a robust ML training pipeline and a high-performance runtime environment:
The ML Pipeline (Python): We trained our models on the UNSW-NB15 dataset. We used LightGBM (Platt-calibrated for binary detection, temperature-scaled for multiclass) and an Isolation Forest for unsupervised anomaly detection. The SOC Engine (TypeScript/Bun): To make the system incredibly fast and portable, we exported the model artifacts as JSON and wrote a custom ML inference engine in TypeScript. It natively executes the LightGBM tree-walking and Isolation Forest scoring without relying on a Python backend at runtime. The Frontend (Next.js 16): We built a sleek, 5-view SOC dashboard using Next.js, React 19, and shadcn/ui. The frontend communicates with the TypeScript engine via REST and Socket.io for real-time telemetry streaming.
Challenges we ran into TypeScript ML Inference Parity: One of our hardest technical challenges was completely decoupling our runtime from Python. We had to write a custom tree-walker in TypeScript that perfectly mirrored the Python LightGBM outputs. After rigorous testing, we achieved exact parity (probabilities matching within a ≤ 1e-4 margin of error). Dataset Imbalance: The UNSW-NB15 dataset is highly imbalanced (e.g., the 'Worms' category had very few rows compared to 'Normal' traffic). We had to carefully implement OOF (out-of-fold) calibration and temperature scaling to ensure the model didn't heavily bias toward majority classes. Explainability vs. Accuracy: We knew security professionals wouldn't trust a black-box AI. We spent a massive amount of time engineering the "Explainability Center"—calculating exact local SHAP attributions and transparent risk formulas so every AI decision is auditable and grounded in evidence
What I learned
Python-trained tree structures into optimized JSON artifacts and writing a custom TypeScript interpreter for them, we achieved incredibly fast, lightweight inference suitable for edge deployment. We also learned how to balance highly imbalanced cybersecurity datasets using SMOTE.
AI tools used
Ai helped us translate complex Scikit-Learn / LightGBM tree structures into a custom TypeScript inference engine, debug Next.js Vercel deployment issues, and design the sleek, glassmorphic UI components using Tailwind CSS and Framer Motion.