Video monitoring analysis method and system and storage medium
Through the deep learning model, the visual features in the video frame data and the text features input by the user are extracted, multimodal fusion data is generated, and text information described in natural language is output, which solves the problems of waste of storage and information screening in traditional monitoring systems, and realizes efficient video surveillance analysis.
Patent Information
- Application Number
- CN202510467645.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional monitoring systems have problems such as wasting storage space and difficulty in screening information. Most of the videos recorded for a long time are irrelevant information. Saving the entire video takes up a lot of storage resources and the later screening efficiency is low.
By obtaining the video file to be analyzed, performing frame extraction processing, and using deep learning models to extract visual features and text features, generating multimodal fusion data, and finally outputting text information of natural language descriptions, including target behavior descriptions and event causal analysis descriptions.
It has achieved the improvement of video analysis from detection to understanding and analyzing the causal relationship of events. The video content can be summarized and output in text, reducing the pressure of manual participation and storage and improving efficiency.
Smart Images

Figure CN119992428A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of video analysis and processing technology, and specifically relates to a video surveillance analysis method, system and storage medium. Background Art
[0002] Video surveillance systems are one of the important symbols of the liberation of productivity in modern society, and their development has undergone significant technological evolution. Early surveillance mainly relied on human duty. 24-hour shift-based manual monitoring not only consumed a lot of manpower and financial resources, but also was prone to missed judgments and low efficiency due to energy and attention limitations. With the widespread deployment of surveillance equipment, staff can quickly obtain real-time images, greatly improving the security and management efficiency of social operations. The surveillance system has gradually become the infrastructure for the stable operation of society. The rise of big data and cloud computing has injected new vitality into the surveillance system, enabling it to provide efficient solutions in scenarios such as key event playback and auxiliary law enforcement. However, this has also brought new problems such as waste of storage space, data redundancy, and difficulty in extracting key information.
[0003] Although the intelligence of surveillance systems has been gradually improved, their essential functions are still focused on the acquisition and storage of raw videos. This mode of operation leads to a waste of storage space. Most of the long-recorded videos contain irrelevant information, and only a small number of clips contain key information. Saving the entire video not only takes up a lot of storage resources, but also increases the burden of later screening. Even clips that have been screened and detected may still contain a long duration (such as one hour), and manual frame-by-frame viewing is required to grasp the content, which is inefficient and easy to miss. In addition, traditional indexing methods based on manual annotation or metadata are difficult to meet actual needs when the length of the video and the complexity of the content increase. Summary of the invention
[0004] In view of the above analysis, the embodiments of the present invention aim to provide a video surveillance analysis method, system and storage medium, aiming to solve the problems of storage waste and information screening difficulties in traditional surveillance systems.
[0005] In a first aspect of the present application, a video surveillance analysis method is provided, comprising: Acquire a video file to be analyzed, and perform frame extraction processing on the video file to obtain video frame data; Extracting visual features from the video frame data using a deep learning model; Extract text features from the text prompts input by the user and generate corresponding text features; Inputting the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data, and finally outputting text information described in natural language, wherein the text information includes a target behavior description and an event causal analysis description; Among them, the deep learning model extracts the behavioral characteristics of the target based on the multimodal fusion data and constructs time series event data; uses a graph neural network to construct a time event relationship map, and calculates the time dependency weights between each event in the time series event, and determines the associated events with logical relationships in the time series events; uses a causal reasoning model to calculate the causal relationship between the associated events and constructs a causal chain of event development in the time series events; based on the time event relationship map and the causal chain, uses a large language model to generate the text information.
[0006] Optionally, the adopting of a graph neural network to construct a time event relationship graph, and calculating the time dependency weights between the events in the time series events, and determining the associated events with logical relationships in the time series events includes: Mapping events in the time series event data into nodes in a time graph neural network; Construct edges between events based on temporal order and spatial relationships; Using the attention mechanism of the graph neural network or the correlation calculation based on the time series, the time dependency weights between the events in the time series are calculated; A graph neural network trained based on historical event data is used to determine associated events with logical relationships in time series events.
[0007] Optionally, the using a causal reasoning model to calculate the causal relationship between the associated events and constructing a causal chain of event development in a time series event includes: Constructing a cause-effect diagram based on the time-event relationship diagram; Calculate causal weights between related events; The causal weight sorting method is adopted to select the path with the highest causal weight and filter out the paths below the preset threshold to construct the causal chain.
[0008] Optionally, it also includes: Use a Transformer-based time series prediction model to calculate possible future events; When the confidence level of a predicted event exceeds a preset threshold, an early warning is triggered to notify users of future event risks.
[0009] Optionally, the generating the text information using a large language model based on the time event relationship graph and the causal chain includes: Preprocessing the time event relationship graph and the causal chain to generate structured data suitable for inputting into the large language model; Constructing prompt words for input into the large language model; The large language model is used for reasoning to generate a target behavior description and an event causal analysis description.
[0010] Optionally, the step of inputting the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data, and finally outputting text information described in natural language includes: The visual features and text features are input into a deep learning model, the deep learning model uses a cross-attention mechanism to calculate the attention weights of the visual features and text features, performs multi-layer fusion on the visual features and text features in a multi-layer Transformer module, and outputs multimodal fusion data; Inputting the multimodal fusion data into a Transformer decoder to generate structured text information; The structured text information is input into a large language model, and combined with preset prompt words to generate text information described in natural language.
[0011] Optionally, obtaining the video file to be analyzed includes: A directory monitoring mechanism is used to monitor the specified directory in real time, and when a video file to be analyzed is detected, the video file is added to the video processing queue; The size of the received video file is dynamically monitored to determine whether the video file has been transmitted.
[0012] A second aspect of the present application provides a video surveillance analysis system, comprising: The video decoding and frame extraction unit is configured to obtain a video file to be analyzed, and perform frame extraction processing on the video file to obtain video frame data; The deep learning reasoning and analysis unit is configured to use a deep learning model to extract visual features in the video frame data; extract text features of the text prompts input by the user to generate corresponding text features; input the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data, and finally output text information described in natural language, wherein the text information includes a target behavior description and an event causal analysis description; wherein the deep learning model extracts the target's behavioral features based on the multimodal fusion data and constructs time series event data; a graph neural network is used to construct a time event relationship map, and the time dependency weights between each event in the time series event are calculated to determine the associated events with logical relationships in the time series event; a causal reasoning model is used to calculate the causal relationship between the associated events, and a causal chain of event development in the time series event is constructed; based on the time event relationship map and the causal chain, the text information is generated using a large language model.
[0013] Optionally, the video decoding and frame extraction unit, and the deep learning reasoning and analysis unit are two completely independent units, and data is transmitted between the two units through a file synchronization mechanism.
[0014] According to a third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the video surveillance analysis method described in any one of the above is implemented.
[0015] The video surveillance analysis method provided in the present application obtains the video file to be analyzed, and performs frame extraction processing on the video file to obtain video frame data; uses a deep learning model to extract visual features in the video frame data; extracts text features from the text prompts input by the user to generate corresponding text features; inputs the visual features and the text features into the pre-trained deep learning model to generate multimodal fusion data, and finally outputs text information described in natural language, wherein the text information includes a description of the target behavior and a description of the event causal analysis. The present application can upgrade video analysis from detection to understanding, so that the system can analyze the causal relationship of events, and the video content can be summarized and output in text form. Relevant staff no longer need to face lengthy original videos, but can quickly grasp the core information of the surveillance video, which not only greatly saves manual participation, but also significantly improves efficiency. In addition, the present application also provides a video surveillance analysis system and storage medium with the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0017] Figure 1 A flowchart of a specific implementation of the video surveillance analysis method provided in this application Figure 2 A flowchart of another specific implementation of the video surveillance analysis method provided by this application; Figure 3 A structural block diagram of the video surveillance analysis system provided in this application; Figure 4 A schematic diagram of the front-end operation process of the video surveillance analysis system provided in this application; Figure 5 A schematic diagram of the back-end operation process of the video surveillance analysis system provided in this application; Figure 6A schematic diagram of a specific embodiment of the video surveillance analysis system provided by the present application; Figure 7 This is a schematic diagram of the analysis result page of the video surveillance analysis system provided in this application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. It should be noted that, in the absence of conflict, the embodiments in the present disclosure and the features in the embodiments can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0019] The terms used here are for the purpose of describing specific embodiments, and are not intended to be restrictive. As used here, unless the context clearly indicates otherwise, the singular forms "one (kind, person)" and "said (the)" are also intended to include plural forms. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is explained that there are stated features, integral bodies, steps, operations, parts, assemblies and / or their groups, but it is not excluded that there are or add one or more other features, integral bodies, steps, operations, parts, assemblies and / or their groups. It should also be noted that, as used here, the terms "substantially", "approximately" and other similar terms are used as approximate terms and not as degree terms, so that they are used to explain the inherent deviations of the measured values, calculated values and / or the values provided that will be recognized by those of ordinary skill in the art.
[0020] A flowchart of a specific implementation of the video surveillance analysis method provided in this application is as follows Figure 1 As shown, the method specifically includes: S101: Acquire a video file to be analyzed, and perform frame extraction processing on the video file to obtain video frame data.
[0021] Specifically, obtaining the video file to be analyzed includes: using a directory monitoring mechanism to monitor the specified directory in real time, adding the video file to the video processing queue when the video file to be analyzed is detected; dynamically monitoring the size of the received video file to determine whether the video file has been transmitted. After obtaining the video file to be analyzed, the video file is subjected to frame extraction processing to obtain video frame data.
[0022] In this process, clips related to specific targets can be extracted from the video file to be analyzed, for example, only scenes including specific people or objects, specific behaviors or events, and scenes that meet the user's query conditions can be retained. In this way, the original video can be cropped according to the detected targets and behaviors and the target timestamp, so that only relevant clips can be retained. The video can also be compressed to reduce storage space.
[0023] S102: Extract visual features from the video frame data using a deep learning model.
[0024] Use deep learning models to extract visual features from video frame data.
[0025] S103: extracting text features from the text prompt input by the user to generate corresponding text features.
[0026] Parse the text prompt entered by the user and generate text features.
[0027] S104: Input the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data, and finally output text information described in natural language, wherein the text information includes a target behavior description and an event causal analysis description.
[0028] Among them, the deep learning model extracts the behavioral characteristics of the target based on the multimodal fusion data and constructs time series event data; uses a graph neural network to construct a time event relationship map, and calculates the time dependency weights between each event in the time series event, and determines the associated events with logical relationships in the time series events; uses a causal reasoning model to calculate the causal relationship between the associated events and constructs a causal chain of event development in the time series events; based on the time event relationship map and the causal chain, uses a large language model to generate the text information.
[0029] Among them, the use of graph neural networks to construct a time event relationship graph, and calculate the time dependency weights between each event in the time series events, and determine the associated events with logical relationships in the time series events specifically include: mapping the events in the time series event data into nodes in the time graph neural network; constructing edges between events based on time sequence and spatial relationships; using the attention mechanism of the graph neural network or based on the correlation calculation of the time series to calculate the time dependency weights between each event in the time series events; using the graph neural network trained based on historical event data to determine the associated events with logical relationships in the time series events.
[0030] Using graph neural networks to process video event relationships, the core steps include: using event nodes to represent independent events in the video, such as target behavior, object interaction, and environmental changes. Specifically, events are extracted based on the target detection and behavior recognition model, and the events are mapped to nodes in the graph. Connect events based on time sequence: If event B occurs after event A, connect B and A. Calculate the time interval between events, and do not establish an edge when the time interval exceeds the threshold. Connect events based on spatial relationships: If event B occurs in the same physical area as event A, the connection weight is enhanced. Time dependency weights are used to quantify the degree of association between events, and GNN-based attention mechanisms (GAT) or time series-based correlation calculations can be used. A graph neural network is used to determine associated events with logical relationships in time series events. The graph neural network is pre-trained, learns the relationship between events, and is used to predict potential event sequences.
[0031] In the above embodiment, a causal reasoning model is used to calculate the causal relationship between the related events, and a causal chain of event development in a time series event is constructed, including: constructing a causal graph based on the time event relationship graph; calculating the causal weights between the related events; and using a causal weight sorting method to select the path with the highest causal weight and filter out the paths below a preset threshold to construct a causal chain.
[0032] The event is regarded as a node, the causal relationship as a directed edge, and the preliminary connection is determined through time sequence and correlation analysis to construct a causal graph. Causal weight is used to quantify the strength of the causal relationship between events, not just the temporal order. The causal weight of the event is calculated, and counterfactual reasoning is performed to verify the inferred causal relationship. Through the causal graph, calculated weights, and counterfactual reasoning, the final causal chain can be constructed.
[0033] In addition, based on the above embodiments, the present application also includes: using a Transformer-based time series prediction model to calculate possible future events; when the confidence of the predicted event exceeds a preset threshold, triggering an early warning to notify the user of future event risks.
[0034] In some specific implementations, based on the time-event relationship graph and the causal chain, generating the text information using a large language model includes: preprocessing the time-event relationship graph and the causal chain to generate structured data suitable for input into the large language model; constructing prompt words for input into the large language model; using the large language model for reasoning to generate a target behavior description and an event causal analysis description.
[0035] In some specific implementations, the visual features and the text features are input into a pre-trained deep learning model to generate multimodal fusion data, and finally outputting text information described in natural language, including: inputting the visual features and the text features into the deep learning model, the deep learning model uses a cross-attention mechanism to calculate the attention weights of the visual features and the text features, performing multi-layer fusion of the visual features and the text features in a multi-layer Transformer module, and outputting multimodal fusion data; inputting the multimodal fusion data into a Transformer decoder to generate structured text information; inputting the structured text information into a large language model, and generating text information described in natural language in combination with preset prompt words.
[0036] It is understandable that the present application can output text information described in natural language or structured text information, which does not affect the implementation of the present application.
[0037] In a specific embodiment, Figure 2 As shown, the video surveillance analysis method provided by the present application may include the following steps: S201: Front-end interaction: interacting with the user through the front-end interface, receiving the video files uploaded by the user or specifying the video file directory to be processed.
[0038] S202: Video file monitoring: The backend service monitors the specified directory in real time, detects newly added video files, and automatically determines whether the file transfer is completed.
[0039] In order to avoid reading the video file when it is not completely written, this application defines a wait_for_completion(file_path) function: first get the file size, if the file size does not change within a few seconds (determined by wait_time=5), the file transfer is considered complete.
[0040] S203: Video decoding and frame extraction: For the transmitted video files, the backend service uses an efficient video reading library to extract frames at a specified frame rate to obtain a series of video frame data.
[0041] When using an efficient video reading library for frame extraction, frames are extracted at a specified frame rate (for example, 1 frame per second) through an efficient video reading library such as decorator, and evenly sampled when there are too many frames to avoid insufficient video memory or memory.
[0042] S204: Deep learning reasoning analysis: The extracted video frame data is input into a preloaded deep learning model, and combined with preset question prompts, a semantic analysis of the video content is performed to identify the scene, target object, character action and other information in the video, and generate text information.
[0043] The deep learning model is a large multimodal model. The model calculates the attention weights for the input sequence (text vector + image vector) respectively, and continuously updates and fuses them in the multi-layer Transformer module, and finally obtains an implicit representation containing the comprehensive information of "video frame semantics + question semantics".
[0044] S205: Result storage and visual presentation: The results of deep learning reasoning analysis are stored in a pre-set storage medium, such as an Excel file, and the analysis results are notified to the user in real time through the front-end interface. A visual presentation function is also provided so that the user can intuitively view the analysis results.
[0045] This application uses deep learning combined with natural language processing, and uses deep learning target detection technology to screen and edit original videos, retaining only clips related to specific targets, significantly reducing storage pressure. At the same time, by combining video semantic analysis and text generation technology, video content can be summarized and structured in text form, further improving the convenience of information management and retrieval. In addition, the integration of automated workflows from video upload monitoring, deep learning analysis to text output not only saves manual participation, but also significantly improves efficiency.
[0046] Through the application of this application, staff no longer need to face lengthy original videos and search for key events like looking for a needle in a haystack, but can easily grasp the core information through the structured text content automatically generated by the system. This application realizes a complete technical chain from video acquisition to key content extraction by combining automated monitoring of monitoring directories, deep learning semantic analysis and structured output of results. This not only solves the problems of storage waste and information screening difficulties in traditional monitoring systems, but also provides an efficient solution for video data management in multiple fields.
[0047] The video surveillance analysis method provided in this application can be implemented by a corresponding video surveillance analysis system, such as Figure 3 As shown, the video surveillance analysis system includes the following modules: Front-end module 100: used to interact with users, receive video files uploaded by users or specify a video file directory to be processed, and display analysis results.
[0048] Backend module 200: includes a video file monitoring unit 201, a video decoding and frame extraction unit 202, a deep learning reasoning analysis unit 203, and a result storage and visualization presentation unit 204.
[0049] The video file monitoring unit 201 is used to monitor the designated directory in real time, detect newly added video files, and automatically determine whether the file transfer is completed.
[0050] The video decoding and frame extraction unit 202 is used to extract frames from the video file that has completed transmission.
[0051] The deep learning reasoning analysis unit 203 is used to perform semantic analysis on the extracted image frames. The deep learning model used is a multi-modal large model with functions such as image encoding, multi-modal fusion and answer generation.
[0052] The result storage and visualization unit 204 is used to store the analysis results and notify the user in real time through the front-end interface.
[0053] In addition, the back-end module of the present application also includes an exception handling unit, which is used to perform exception handling when an exception occurs during video reading, reasoning analysis or result storage, to ensure the stability and traceability of the system.
[0054] This application uses websocket technology and http technology to form a simple front-end and back-end separation system, that is, recognition and editing, video understanding are divided into two completely independent units, so that the two will not interfere with each other whether in the development stage or in the actual operation stage. The two units are "softly connected" through a file synchronization mechanism, and a queue mechanism is used to manage the video transmitted from the detection end to the analysis end.
[0055] The communication port is established between the front-end and back-end through websocket, and some requests are interacted using http.
[0056] It is understandable that this application will perform the following preparations after the program is started, namely the INIT initialization process: 1. Model loading: Load and initialize the deep learning model and word segmenter in advance for subsequent video content understanding. The model can support multiple precisions (such as int4 or bf16), which can be selected based on server performance.
[0057] 2. Excel file management: Create or open an Excel spreadsheet in a local or network disk to record the processing information of each video, including the start and end time of the video, and the analysis conclusions generated by the deep learning model.
[0058] 3. Directory monitoring and queue initialization: Watchdog monitors the specified folder in real time. When a new video file is detected, the file path will be placed in the security queue for subsequent video processing threads to obtain.
[0059] 4. Multithreading and asynchronous framework: An HTTP server thread: Provides access to video files and related result files to the front end or other clients on the specified port.
[0060] A video processing thread: It is responsible for taking out the video files to be analyzed from the queue and performing decoding, frame extraction and inference analysis.
[0061] WebSocket server running in the main thread: supports multiple clients connecting in parallel and can send real-time information to the front end during deep learning inference.
[0062] Figure 4 The front-end operation process diagram is shown, and the process includes the following steps: 1. Initialization phase 1. Load dependencies and create the main window Introduce graphical interface library, network library, image processing library, etc.
[0063] Create the main application window and set its title and basic size.
[0064] 2. Interface layout There are several buttons placed on the top, such as "Download Results" and "Broadcast".
[0065] The left side shows the "Unprocessed Video List" and "Video Playback Area".
[0066] The right side displays "Console Information" and "Analysis Results".
[0067] 3. Multithreading preparation Create a voice thread to handle text-to-speech (TTS) requests.
[0068] Establish a network thread and keep connected to the backend WebSocket service.
[0069] Use timed polling in the main thread to check for new messages received from the network thread and update the interface.
[0070] 2. System operation process 1. Monitor backend messages The network thread connects to the backend WebSocket and continuously receives instructions or information, such as "new video arrival", "processing", "analysis results", etc.
[0071] After receiving the message, put it into a thread-safe queue for use by the main thread.
[0072] 2. Main interface update The main thread periodically takes messages from the queue: If it is a "new video", update the list on the left.
[0073] If "Processing", download and play the video.
[0074] If it is "Analysis Results", the text will be inserted into the analysis area on the right and prompted to update Excel.
[0075] If it is "Log", it will be displayed in the "Console Information" on the right.
[0076] 3. Video playback When a video needs to be played, use the video reading library to obtain the image frame by frame, scale it and render it in the interface label.
[0077] After playing to the end, it will automatically loop or stop.
[0078] 4. Text reading Use the voice thread to perform text playback: On macOS, use the system's built-in say command to read Chinese; Other platforms will only print prompts for now or expand other TTS solutions.
[0079] Users can choose to read all analysis results or only the new additions.
[0080] 5. Download analysis results After clicking the "Download Results" button, the program requests the backend server to obtain the result file (Excel) and save it locally.
[0081] If a file with the same name exists, a timestamp will be added after the name to prevent overwriting.
[0082] After the download is complete, you can try to open the result file automatically.
[0083] 3. Branching Logic Example Network connectivity: If the backend message type cannot be identified, no processing will be performed; if the network is abnormal, an error message will be printed, but the interface will continue to run.
[0084] Video download error: If the download is unsuccessful, a dialog box pops up to prompt.
[0085] Judgment of the content to be read aloud: If there are no Chinese characters or insufficient text, the prompt "No text to be read aloud" will appear.
[0086] System shutdown: When the user closes the window, the voice thread is stopped and the main loop ends.
[0087] In a specific embodiment, the following logical structure is included: GUI layout and multithreading Use tkinter to create a graphical interface, including modules such as the video playback window, console area, analysis result area, and unprocessed video list.
[0088] The parallel operation of voice reading and message receiving is realized through SpeechThread and a background WebSocket thread, without blocking the main interface operation.
[0089] WebSocket + Queues The front-end maintains a long connection with the back-end through websocket_handler, and new messages sent by the back-end at any time can be captured by the front-end.
[0090] After the message is written to the queue, the main thread's process_queue() (Tk polling) updates the GUI in a safe environment; this is a common practice for Tkinter thread safety.
[0091] Video playback Use OpenCV to read and update ImageTk.PhotoImage in real time and display it in the interface control Label.
[0092] Simplified processing: If the reading is completed, it will be played in a loop from the beginning.
[0093] Branching Logic The platform (macOS vs. others) determines the TTS method; Parse the message data_type and determine different responses such as "new_video", "processing_video", "analysis_result", etc.
[0094] The file download part will also determine whether the file already exists, whether the download is successful, etc.
[0095] Overall functionality Real-time display: Dynamically update the GUI when receiving "new video / current processing / analysis results" from the backend.
[0096] Auxiliary operations: You can download and view Excel analysis results with one click; you can choose to read all or newly added Chinese content.
[0097] Applicable scenarios: Friendly human-computer interaction, suitable for quickly viewing analyzed videos and their results in monitoring centers or scientific research scenarios.
[0098] Through the organic combination of these logical structures, the video captured and analyzed by the backend is displayed on the frontend. The cooperation of multithreading and Tk event polling enables it to flexibly respond to user operations and keep the interface smooth.
[0099] Figure 5 The schematic diagram of the backend operation process is shown, which mainly includes the initialization preparation stage and the system operation process.
[0100] S1. Initialization preparation phase, including: S1-1. Import dependent libraries os, time, queue, threading: commonly used operating system, time, queue, multi-threading support and other Python built-in libraries.
[0101] PIL.Image: used for image processing (converting video frames into Image objects).
[0102] transformers: Hugging Face's model loading library, used here to load AutoModel and AutoTokenizer.
[0103] decord: An efficient video reading library (can read video frames in batches).
[0104] watchdog: A library used to monitor changes in files in a specified directory.
[0105] openpyxl: for reading and writing Excel files.
[0106] asyncio, websockets, json: build asynchronous WebSocket server and send asynchronous messages to the front end.
[0107] http.server, HTTPServer, SimpleHTTPRequestHandler: Provide an HTTP file service locally.
[0108] torch: for deep learning inference support.
[0109] S1-2. Creation of global variables and key objects Queue video_queue: used to store the video file paths to be processed and provide a secure communication mechanism for subsequent multi-threaded or asynchronous processing.
[0110] Excel file initialization: The program will detect whether video_analysislocat_results.xlsx already exists. If not, a new one will be created and the header will be written. If it already exists, it will be loaded directly. This Excel is mainly used to record the processing results of the video (including the start time, end time and analysis results).
[0111] processed_videos: A list used to record the paths of processed video files to prevent duplicate processing.
[0112] connected_clients: A collection of clients connected via WebSocket. When a new WebSocket client is connected, it is added to the collection and removed when the connection is closed.
[0113] Directory Settings video_directory: The directory of the monitored video files, i.e. " / icislab / volume3 / mitchel / llama.cpp / supervised_file".
[0114] excel_directory: The root directory used when providing the HTTP server, also set to the same directory " / icislab / volume3 / mitchel / llama.cpp / supervised_file".
[0115] S1-3. Model initialization (deep learning part) Load the model and tokenizer through AutoModel and AutoTokenizer.
[0116] Two examples are given in the code: one is the int4 quantized (actually used) loading method, and the other is the bf16 mode (commented, not enabled yet).
[0117] Important "hyperparameters": When using model.chat() in int4 mode, max_inp_length=16384 is set, which is the maximum length of the input sequence.
[0118] In bf16 mode higher memory usage is allowed, but performance and accuracy are slightly improved (commented out).
[0119] MAX_NUM_FRAMES = 12800: The maximum number of frames read when processing a video, used to limit excessive video memory or memory usage.
[0120] S1-4. Threads and asynchronous event loops Main thread: The entry point of the program is in the if __name__ == '__main__': part, which starts the event loop loop = asyncio.get_event_loop() and runs the WebSocket server.
[0121] HTTP server thread: Start a separate thread via threading.Thread to run start_http_server() so that the main thread is not blocked.
[0122] Video processing thread: Also start a daemon thread through threading.Thread to execute process_video(video_queue), which will continue to fetch videos from video_queue for analysis.
[0123] Watchdog observer thread: After observer.start(), the Watchdog library starts a listening thread internally to monitor the file changes in video_directory in real time.
[0124] In particular, the maximum number of threads supported by multiple threads is limited. There is no hard limit on the maximum number of threads in the code. In theory, it depends on the resources that can be allocated by the operating system and the Python runtime environment. In actual use, the main threads are: 1. Main thread (running event loop) 2. HTTP Server Thread 3. Video processing thread 4.Watchdog observer thread S2. System operation process, including: S2-1. Monitoring file events: VideoHandler When the code runs, Watchdog will monitor the new file events under the video_directory path: callback function on_created(): If a file in .mp4, .avi or .mkv format is detected, it is considered a new video. After printing the prompt, put the video path event.src_path into the video_queue queue. At the same time, use asyncio.run_coroutine_threadsafe() to send a JSON message of { "type": "new_video", "video_path": ...} to the front end to inform the client that there is a new video to be processed.
[0125] S2-2. Waiting for file transfer to complete: wait_for_completion() To avoid reading when the video file has not been completely written, a wait_for_completion(file_path) function is defined: 1. Get the file size first. If the file size does not change within a few seconds (determined by wait_time=5), the file transfer is considered complete.
[0126] 2.check_interval=1 means checking the file size every 1 second; if the file size is detected to be unchanged for wait_time consecutive times, 1 is returned.
[0127] S2-3 video processing, calling Minicpm v2.6 multimodal large model to implement video analysis, includes the following steps: 1. Image Encoder: When receiving image input, Minicpm v2.6 usually uses the built-in visual encoding module (VIT) to extract features from the image. For each image frame, the model extracts its high-dimensional visual features and generates a sequence of image feature vectors (Visual Embeddings).
[0128] 2. Text Encoder / Tokenizer: Similar to images, the text part first uses Tokenizer to convert characters or word fragments into Token ID sequences, and then inputs them into the corresponding text encoding network to obtain a series of text vectors (Text Embeddings).
[0129] 3. Multimodal fusion (Cross-Attention / Self-Attention mechanism): In a large multimodal model, visual features and text features interact in the fusion layer (usually self-attention or cross-attention). Specifically, the model calculates the attention weights for the input sequence (text vector + image vector) respectively, and continuously updates and fuses them in the multi-layer Transformer module, and finally obtains the feature encoding containing the comprehensive information of "video frame semantics + question semantics".
[0130] 4. Generate answers (Decoder / Output Head): The fused semantic representation will enter the model's decoder head to generate the final text answer (or annotation information, labels, etc.). Minicpm v2.6 will decode the output word by word (or word by word fragment) according to the existing context during the inference phase until the end condition is met (such as reaching the maximum length or encountering a special stop character).
[0131] 5. Output text results: Finally, a string of answers is generated, which usually includes a description of the scene in the video, the number of targets, guesses about the time and place, and specific answers to the questions.
[0132] S2-4. WebSocket server and message sending Start the service at port 0.0.0.0:8765 via websockets.serve(handler, "0.0.0.0", 8765). All clients connected to this port will trigger handler(websocket, path). When a client connects, the websocket is stored in connected_clients. When the connection is closed, it is removed from it. Send data to the client: send_data(data) S2-5. HTTP server, start_http_server() function is executed in a separate thread: 1.os.chdir(excel_directory): Switch the working directory to the location where Excel and video files are stored.
[0133] 2.HTTPServer(('0.0.0.0',8000), SimpleHTTPRequestHandler).serve_forever(): listens to port 8000. The client can directly access the video or Excel through http: / / server IP:8000 / filename.
[0134] The backend service of this application completes the preparation of directories, Excel, queues, and deep learning models during initialization. It handles file monitoring, file transfer, video analysis, model reasoning, and result storage through multi-threading and asynchronous event loop division of labor, and performs exception or branch processing in multiple branches of the entire process, ensuring the stability and traceability of the service.
[0135] Figure 6 A schematic diagram of a specific embodiment is shown. In this embodiment, a video analysis front-end display page includes three options: broadcast new content, click to broadcast, and download analysis results, as well as a list of unprocessed videos, console information, video playback, and analysis results. The video content being processed and analyzed is displayed on the video playback page, and the program analysis results are displayed on the analysis result page. Figure 7 shown.
[0136] This application realizes the fully automated collection and semantic analysis of video data, and provides real-time processing capabilities for multi-channel surveillance videos or batch video files. Compared with the traditional manual mode, this application has outstanding advantages in the following aspects: (1) Real-time and automation: It can run continuously in the background, automatically detect and process new files without human intervention.
[0137] (2) High concurrency and scalability: The multi-threaded queue structure and asynchronous framework support not only the rapid processing of large batches of videos, but also the flexible expansion of model inference load capacity based on server resources.
[0138] (3) Result visualization and ease of use: The processing progress and final results are sent to the front end via WebSocket, and users can view them in real time on the graphical interface; HTTP services are also provided to directly access videos, Excel and other files.
[0139] (4) Reliability and stability: Complete exception capture and processing logic is set up in each core link to ensure that even if a video processing channel fails, it will not cause the entire system to crash.
[0140] In summary, this application can significantly improve the efficiency and intelligence of video surveillance and analysis. The system is very easy to transplant and deploy, and can be applied to multiple fields such as intelligent security, traffic monitoring, academic research, and natural environment observation, with significant application value and broad market prospects.
[0141] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the video surveillance analysis method according to any one of the above-mentioned methods is implemented.
[0142] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0143] The professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0144] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0145] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A video surveillance analysis method, characterized in that: include: Acquire a video file to be analyzed, and perform frame extraction processing on the video file to obtain video frame data; Extracting visual features from the video frame data using a deep learning model; Extract text features from the text prompts input by the user and generate corresponding text features; Inputting the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data, and finally outputting text information described in natural language, wherein the text information includes a target behavior description and an event causal analysis description; Among them, the deep learning model extracts the behavioral characteristics of the target based on the multimodal fusion data and constructs time series event data; uses a graph neural network to construct a time event relationship map, and calculates the time dependency weights between each event in the time series event, and determines the associated events with logical relationships in the time series events; uses a causal reasoning model to calculate the causal relationship between the associated events and constructs a causal chain of event development in the time series events; based on the time event relationship map and the causal chain, uses a large language model to generate the text information.
2. The video surveillance analysis method according to claim 1, characterized in that: The method of using a graph neural network to construct a time event relationship graph, and calculating the time dependency weights between the events in the time series events, and determining the associated events with logical relationships in the time series events includes: Mapping events in the time series event data into nodes in a time graph neural network; Construct edges between events based on temporal order and spatial relationships; Using the attention mechanism of the graph neural network or the correlation calculation based on the time series, the time dependency weights between the events in the time series are calculated; A graph neural network trained based on historical event data is used to determine associated events with logical relationships in time series events.
3. The video surveillance analysis method according to claim 2, characterized in that: The use of a causal inference model to calculate the causal relationship between the associated events and construct a causal chain of event development in a time series event includes: Constructing a cause-effect diagram based on the time-event relationship diagram; Calculate causal weights between related events; The causal weight sorting method is adopted to select the path with the highest causal weight and filter out the paths below the preset threshold to construct the causal chain.
4. The video surveillance analysis method according to claim 1, characterized in that: Also includes: Use a Transformer-based time series prediction model to calculate possible future events; When the confidence level of a predicted event exceeds a preset threshold, an early warning is triggered to notify users of future event risks.
5. The video surveillance analysis method according to any one of claims 1 to 4, characterized in that: The generating the text information by using a large language model based on the time event relationship graph and the causal chain includes: Preprocessing the time event relationship graph and the causal chain to generate structured data suitable for inputting into the large language model; Constructing prompt words for input into the large language model; The large language model is used for reasoning to generate a target behavior description and an event causal analysis description.
6. The video surveillance analysis method according to any one of claims 1 to 4, characterized in that: The step of inputting the visual features and the text features into a pre-trained deep learning model to generate multimodal fusion data and finally outputting text information described in natural language includes: The visual features and text features are input into a deep learning model, the deep learning model uses a cross-attention mechanism to calculate the attention weights of the visual features and text features, performs multi-layer fusion on the visual features and text features in a multi-layer Transformer module, and outputs multimodal fusion data; Inputting the multimodal fusion data into a Transformer decoder to generate structured text information; The structured text information is input into a large language model, and combined with preset prompt words to generate text information described in natural language.
7. The video surveillance analysis method according to claim 6, characterized in that: The obtaining of the video file to be analyzed comprises: A directory monitoring mechanism is used to monitor the specified directory in real time, and when a video file to be analyzed is detected, the video file is added to the video processing queue; The size of the received video file is dynamically monitored to determine whether the video file has been transmitted.
8. A video surveillance analysis system, characterized in that: include: The video decoding and frame extraction unit is configured to obtain a video file to be analyzed, and perform frame extraction processing on the video file to obtain video frame data; A deep learning reasoning analysis unit, configured to extract visual features from the video frame data using a deep learning model; The text features of the text prompt input by the user are extracted to generate corresponding text features; the visual features and the text features are input into a pre-trained deep learning model to generate multimodal fusion data, and finally the text information described in natural language is output, wherein the text information includes a target behavior description and an event causal analysis description; wherein the deep learning model extracts the target's behavior features based on the multimodal fusion data and constructs time series event data; a graph neural network is used to construct a time event relationship map, and the time dependency weights between each event in the time series event are calculated to determine the associated events with logical relationships in the time series event; a causal reasoning model is used to calculate the causal relationship between the associated events, and a causal chain of event development in the time series event is constructed; based on the time event relationship map and the causal chain, the text information is generated using a large language model.
9. The video monitoring and analysis system according to claim 8, characterized in that: The video decoding and frame extraction unit, and the deep learning reasoning and analysis unit are two completely independent units, and data is transmitted between the two units through a file synchronization mechanism.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the video surveillance analysis method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Video analysis method and device, equipment, storage medium and program product
CN119274105A
Electric power cross-modal knowledge fusion multi-agent cooperative processing method and system
CN119477235A
Event analysis method, system and device based on task generation and multiple modes
CN119557603A
Cited By
Video monitoring method and device and storage medium
CN120812222A
Deep learning-based medical self-service check-in terminal interactor identity binding method and system
CN121389094A