Intelligent multimodal monitoring service for the surveillance area
Abstract:
This paper presents an approach to developing an intelligent multimodal monitoring service for surveillance areas using large neural network models. The proposed method focuses on analyzing heterogeneous data sources – video streams, environmental sensor signals (e.g., temperature, humidity), and event logs – within the observed domain. The system leverages advanced language and vision models (e.g., LLaMA, MiniCPM-V), deployed locally via the Ollama framework, enabling secure and autonomous processing without cloud dependency. A functional prototype has been implemented and tested to detect critical situations, abnormal patterns, and context-specific events offline. The data processing methodology and experimental evaluation based on predefined scenarios are described. Results demonstrate the effectiveness of multimodal models for activity monitoring and highlight their potential in building adaptive and scalable surveillance systems.
Keywords:
intelligent service, multimodal monitoring, Ollama, Large Language Models, activity tracking, video analytics, artificial intelligence