Smart Agentic CCTV Camera

An agentic CCTV system that transforms camera footage into searchable visual memory, enabling natural-language questions about events, objects, and activity.
Smart Agentic CCTV Camera

Problem Statement

CCTV cameras continuously record large amounts of footage, but extracting useful information from that footage remains largely a manual process. When an incident occurs—such as a theft, unauthorized vehicle entry, or a missing item—a security officer or camera owner may need to manually search through hours of recordings to find a few relevant seconds.

Traditional CCTV systems generally allow users to search by basic parameters such as date, time, and camera, while motion-based alerts can generate large numbers of false alarms. Even systems with object detection typically rely on predefined categories and filters, making it difficult to search for events using natural descriptions such as “a person carrying a box.”

More importantly, conventional systems do not maintain a structured understanding of past events. They cannot naturally answer questions such as:

“How many times did the delivery van arrive yesterday?”

or

“What happened at the gate around 3:40 PM?”

The result is a surveillance workflow that requires significant human attention and makes potentially useful footage difficult to access and investigate efficiently.

Proposed Solution

The Smart Agentic CCTV Camera addresses this problem by transforming raw CCTV footage into a structured, searchable visual memory and placing a conversational AI agent on top of it.

The system accepts footage from RTSP-based IP cameras or uploaded MP4 files. A computer-vision pipeline processes the video using YOLO for object detection and ByteTrack for tracking, identifying objects and maintaining their identities across frames.

The detected visual information is then converted into semantic representations using SigLIP, allowing the system to connect visual content with natural-language descriptions. These representations are indexed for fast retrieval, while a structured event memory records what was observed, where it occurred, and when it happened.

On top of this memory, an LLM-based agent interprets natural-language questions, determines what information needs to be retrieved, searches the appropriate tools and returns relevant frames, timestamps, and explanations.

This changes the interaction model from:

Watching hours of CCTV → finding an event manually

to:

Ask a question → search the camera’s memory → retrieve the relevant evidence.

The proposed system is designed around a mobile interface supporting live and recorded video, natural-language search, event timelines, and basic alerts.