Monitoring video analysis method, device and equipment based on artificial intelligence

By employing an AI-based video surveillance analysis method, which utilizes local and global convolutional network models for feature extraction and target identification, and combines adaptive alarm rules, this approach addresses the shortcomings of traditional video surveillance systems in terms of processing efficiency, accuracy, and timeliness, achieving efficient, accurate, and flexible video surveillance analysis.

CN120953876APending Publication Date: 2025-11-14INNER MONGOLIA ELECTRIC POWER SURVEY & DESIGN INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511062554.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional video surveillance systems are inadequate in terms of processing efficiency, accuracy, and timeliness, especially at night or in inclement weather conditions. The centralized data processing mode leads to bandwidth bottlenecks and network congestion, and the alarm mechanism lacks flexibility and cannot be dynamically adjusted.

Method used

An AI-based video surveillance analysis method is adopted, which uses local and global convolutional network models for feature extraction and target identification. Combined with adaptive alarm rules, a distributed learning algorithm is used to reduce dependence on the central server, thereby achieving adaptive alarm classification and dynamic decision-making.

Benefits of technology

It improves the processing efficiency, accuracy, and timeliness of surveillance video analysis, reduces bandwidth pressure, ensures the accuracy and timeliness of emergency response, and enhances the real-time performance and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953876A_ABST
    Figure CN120953876A_ABST
Patent Text Reader

Abstract

The invention provides a monitoring video analysis method, device and equipment based on artificial intelligence. The method comprises the following steps: acquiring a monitoring video of a target area; performing feature extraction on the monitoring video by using an intelligent identification model to obtain an identification target; according to the recognition target, target alarm information is obtained in combination with a preset alarm rule; pushing the target alarm information to a client; wherein the intelligent identification model is determined according to a local convolutional network model and a global convolutional network model, and the local convolutional network model is obtained by training a preset convolutional network model according to a historical monitoring video. According to the invention, the processing efficiency, accuracy and timeliness of monitoring video analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image analysis technology, and in particular to a method, apparatus and equipment for analyzing surveillance videos based on artificial intelligence. Background Technology Currently, video surveillance systems are widely used in various industrial and public safety fields, especially at construction sites of large-scale engineering projects. However, traditional video surveillance systems typically rely on a single camera for image acquisition, which cannot fully interpret visual signals. Furthermore, the centralized data processing model easily creates bandwidth bottlenecks, affecting the system's real-time performance and reliability. Traditional video surveillance systems mainly rely on high-resolution cameras for image acquisition, which, while providing a clear view, performs poorly at night or in inclement weather conditions and struggles to detect non-visual signals. Most existing monitoring systems employ a centralized data processing model, with all data uploaded to a central server for analysis and storage. The large volume of data transmitted over the network consumes significant bandwidth resources, increasing latency. Since video data acquired by cameras must be transmitted to the central server for processing, network bandwidth pressure is significantly increased. Poor network conditions (such as remote areas or areas with insufficient signal coverage) may lead to data transmission interruptions or quality degradation, thus affecting the normal operation of the entire system. While this model facilitates management and maintenance, the massive amount of data easily causes network congestion, impacting the system's real-time performance. Many monitoring systems use preset fixed thresholds to trigger alarms. This method is simple and easy to implement, but it lacks flexibility and cannot dynamically adjust the alarm level according to the actual situation. It may miss some important anomalies or cause frequent false alarms. In video surveillance systems, the performance of the client has a significant impact on the overall system performance, especially in the playback and processing of multiple video streams. The hardware configuration of the client device (such as CPU, GPU, memory, etc.) directly affects the number of video streams that can be played smoothly at the same time. For decoding and rendering of high-definition or multi-channel video streams, low-performance clients may not be able to meet the requirements, resulting in screen stuttering, latency, or even crashes. Summary of the Invention

[0002] The technical problem this invention aims to solve is to provide a method, apparatus, and device for analyzing surveillance videos based on artificial intelligence. This can improve the processing efficiency, accuracy, and timeliness of surveillance video analysis.

[0003] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: An artificial intelligence-based surveillance video analysis method includes: Acquire surveillance video of the target area; The intelligent recognition model is used to extract features from the surveillance video to obtain the target to be identified; Based on the identified target and combined with pre-set alarm rules, target alarm information is obtained; The target alarm information is pushed to the client. The intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model. The local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

[0004] Optionally, an intelligent recognition model is used to extract features from the surveillance video to obtain the target for recognition, including: The surveillance video is input into the first processing module of the intelligent recognition model and subjected to standard convolution to obtain initial basic features; The initial basic features are input into the second processing module of the intelligent recognition model for enhancement and fusion to obtain intermediate enhanced features; The intermediate enhanced features are input into the third processing module of the intelligent recognition model for recognition and localization to obtain the recognition target.

[0005] Optionally, the training process of the intelligent recognition model includes: Retrieve historical monitoring videos from monitoring nodes; The historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model. The hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain the global convolutional network model. Based on the local convolutional network model and the global convolutional network model, an intelligent recognition model is obtained.

[0006] Optionally, the historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model, including: The local convolutional network model is updated according to a preset objective to obtain local hyperparameters; The intelligent recognition model is obtained based on the local hyperparameters.

[0007] Optionally, the hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain a global convolutional network model, including: On the server side, the hyperparameters of the local convolutional network model are integrated and optimized using an aggregation formula, which is:

[0008] in, This indicates the updated hyperparameters. Indicates the number of iterations. Let represent the number of the k-th iteration, and n represent the total number of samples. Denotes the hyperparameters of the k-th iteration. Indicates learning efficiency. This represents the gradient of the k-th iteration.

[0009] Optionally, based on the identified target and in conjunction with pre-set alarm rules, target alarm information is obtained, including: The identified target is input into a large language model for summary and analysis. The large language model obtains initial warning information based on pre-set alarm rules. The initial warning information is cross-validated and dynamically decided to obtain the target alarm information.

[0010] Optionally, the initial warning information is cross-validated and dynamically decided to obtain target alarm information, including: The initial warning information is subjected to data quality verification, multi-modal consistency test and knowledge base matching to obtain cross-validation data; The cross-validation data is dynamically weighted by confidence level to obtain the target alarm information.

[0011] Embodiments of the present invention also provide an artificial intelligence-based surveillance video analysis device, comprising: The acquisition module is used to acquire surveillance video of the target area; The processing module is used to extract features from the surveillance video using an intelligent recognition model to obtain the target; based on the target and combined with pre-set alarm rules, obtain target alarm information; and push the target alarm information to the client; wherein, the intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model, and the local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

[0012] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the artificial intelligence-based surveillance video analysis method of the present invention.

[0013] Embodiments of the present invention also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the artificial intelligence-based surveillance video analysis method described in the present invention.

[0014] The above-described technical solution of the present invention has at least the following technical effects: The above-described AI-based surveillance video analysis method of the present invention acquires surveillance video of a target area; uses an intelligent recognition model to extract features from the surveillance video to obtain the identified target; based on the identified target and combined with pre-set alarm rules, obtains target alarm information; and pushes the target alarm information to a client. The intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model, and the local convolutional network model is trained on a preset convolutional network model using historical surveillance video. This method addresses the problems of insufficient understanding of visual signals, bandwidth bottlenecks caused by centralized data processing models, and increased latency in existing technologies, thereby improving the processing efficiency, accuracy, and timeliness of surveillance video analysis. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the AI-based video surveillance analysis method of the present invention. Figure 2 This is a system architecture diagram of the artificial intelligence-based surveillance video analysis method of the present invention; Figure 3 This is a flowchart of the distributed learning process of the AI-based surveillance video analysis method of the present invention. Figure 4 This is a schematic diagram of the system hardware layout and data flow of the artificial intelligence-based surveillance video analysis method of the present invention; Figure 5 This is a schematic diagram of the AI-based surveillance video analysis device of the present invention. Detailed Implementation

[0016] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0017] like Figure 1 As shown, an embodiment of the present invention proposes an artificial intelligence-based surveillance video analysis method, including: Step S1: Obtain surveillance video of the target area; Step S2: Using an intelligent recognition model, feature extraction is performed on the surveillance video to obtain the target to be identified; Step S3: Based on the identified target and combined with the pre-set alarm rules, obtain the target alarm information; Step S4: Push the target alarm information to the client; The intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model. The local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

[0018] In this embodiment, as Figure 1 As shown, in the AI-based surveillance video analysis method, firstly, surveillance video of the target area is acquired through a field monitoring network. The field monitoring network system includes multiple monitoring nodes, each equipped with a camera, microprocessor, and communication module. These cameras are installed at key locations on the construction site, and the microprocessor and communication module are installed in the construction site's computer room, forming a dense monitoring network. Then, an intelligent recognition model is used to extract features from the surveillance video to obtain the identified target. Next, based on the identified target and pre-set alarm rules, target alarm information is obtained. Finally, the target alarm information is pushed to the client. The intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model. The local convolutional network model is trained on a preset convolutional network model using historical surveillance videos.

[0019] This invention utilizes a distributed learning algorithm to reduce reliance on a central server, lower bandwidth pressure, and improve data processing efficiency. It achieves comprehensive coverage and real-time monitoring of construction areas in complex environments, significantly improving the safety management level of construction sites. An adaptive alarm grading mechanism ensures the accuracy and timeliness of emergency response, avoiding overreaction or underreporting. Artificial intelligence technology controls the pan-tilt unit to track target images, providing managers with a more intuitive and efficient management tool. Decoder technology is applied to meet the needs of low-bandwidth transmission and low-hardware playback, achieving a match between network bandwidth and terminal hardware. A large language model is used for summarizing and querying early warnings, enhancing the system's user experience.

[0020] In an optional embodiment of the present invention, step S2 involves using an intelligent recognition model to extract features from the surveillance video to obtain the target for recognition, including: Step S21: Input the surveillance video into the first processing module of the intelligent recognition model and perform standard convolution to obtain initial basic features; Step S22: Input the initial basic features into the second processing module of the intelligent recognition model for enhancement and fusion to obtain intermediate enhanced features; Step S23: Input the intermediate enhanced features into the third processing module of the intelligent recognition model for recognition and positioning to obtain the recognition target.

[0021] In this embodiment, an intelligent recognition model is used to extract features from the surveillance video to obtain the target to be identified; First, the surveillance video is input into the first processing module of the intelligent recognition model and subjected to standard convolution to obtain initial basic features; different scale versions of the image are obtained. The model adopts a multi-scale input strategy, that is, it processes images of different sizes at the same time to help the model better capture targets of different sizes.

[0022] The first processing module includes a basic convolutional unit, a feature enhancement unit, a pooling unit, and an attention unit. The basic convolutional unit is a combined unit containing one convolutional layer, one batch normalization layer, and one SiLU activation function. The convolutional layer extracts image features, the batch normalization layer accelerates training and reduces overfitting, and the SiLU activation function introduces non-linearity to the model, enabling it to learn more complex patterns. The feature enhancement unit is a bottleneck structure consisting of three convolutional layers, including one 3x3 convolutional kernel and two 1x1 convolutional kernels. It enhances feature representation capabilities while reducing computational cost. The pooling unit performs feature extraction and dimensionality reduction, employing a spatial pyramid structure. By performing pooling operations at different scales, it captures multi-scale spatial information, which is crucial for object detection tasks as it helps the model better identify targets of different sizes. The attention unit uses a channel attention mechanism to adjust the importance of feature channels. The model can automatically learn which feature channels are more important for the current task and assign them more weights, thereby improving detection accuracy.

[0023] Then, the initial basic features are input into the second processing module of the intelligent recognition model for enhancement and fusion to obtain intermediate enhanced features. The second processing module includes a feature fusion unit and an upsampling unit. The feature fusion unit is an important part of the model. It combines feature maps from different levels, integrating low-level detail information with high-level semantic information to form a richer feature representation, which helps the model better understand image content. The upsampling unit is used to restore low-resolution feature maps to higher resolution for fusion with features from earlier layers. This cross-scale feature fusion strategy helps retain more detail information, which is especially important for detecting small targets.

[0024] Finally, the intermediate enhanced features are input into the third processing module of the intelligent recognition model for identification and localization to obtain the identified target. The third processing module includes multiple output heads. The feature map after feature enhancement and fusion is fed into different output heads, each responsible for generating prediction results at a specific scale, including the target bounding box position, size, confidence level, and class probability. For example, there are three output heads, corresponding to feature maps of 64x80x80, 128x40x40, and 256x20x20 respectively. This means the model can make predictions at multiple scales, improving its ability to detect targets of different sizes.

[0025] In an optional embodiment of the present invention, the training process of the intelligent recognition model includes: Step S51: Obtain historical monitoring videos of the monitoring nodes; Step S52: Input the historical monitoring video of the monitoring node into a preset convolutional network model for training to obtain a local convolutional network model; Step S53: Transmit the hyperparameters of the local convolutional network model to the server for integration and optimization to obtain the global convolutional network model; Step S54: Based on the local convolutional network model and the global convolutional network model, obtain the intelligent recognition model.

[0026] In this embodiment, a shared model can be trained between the monitoring nodes and the server without transmitting monitoring video data. The model includes one server and at least two monitoring nodes, each storing monitoring videos captured by its respective cameras as its local dataset. The server and monitoring nodes share a common network structure, with the server responsible for updating the hyperparameters of the global model and the monitoring nodes responsible for updating local hyperparameters. The model training process is distributed. During the training of the intelligent recognition model, historical monitoring videos from the monitoring nodes are first acquired. Then, a pre-defined convolutional network model deployed on the monitoring nodes is trained. There are multiple monitoring nodes, each holding monitoring videos and capable of independently executing model training tasks in its local environment. Local data can be used to train and update the model. After completing this training process, the monitoring nodes calculate the changes in hyperparameters, obtaining a local convolutional network model.

[0027] The hyperparameters of the local convolutional network model are then transmitted to the server, avoiding direct transmission of the raw data. The server coordinates and integrates hyperparameter updates from multiple monitoring nodes, ultimately generating a new round of global convolutional network model. This process optimizes model performance in distributed learning. The server can employ different model aggregation strategies to integrate the hyperparameter updates received from each monitoring node into a global model.

[0028] In an optional embodiment of the present invention, step S52, inputting the historical monitoring video of the monitoring node into a preset convolutional network model for training to obtain a local convolutional network model, includes: Step S521: The local convolutional network model is updated according to a preset target to obtain local hyperparameters; Step S522: Obtain the intelligent recognition model based on the local hyperparameters.

[0029] In this embodiment, the global model is sent back to each monitoring node from the server. This allows each monitoring node to use the new global model as a basis for training its local model in the next iteration. After the monitoring nodes complete their model training, the server receives the local models sent by each monitoring node again, thus ensuring that the global performance of the model is improved. The network structure in this architecture is globally consistent; updates to the monitoring nodes and the server only adjust the relevant parameters of the network structure. Only network parameters are transmitted between the monitoring nodes and the server; user data is not transmitted.

[0030] The client-side local neural network model can update the local model in different directions while co-training the global model. The goal of updating the local neural network model is:

[0031] in,

[0032] Where v represents the global hyperparameter and λ represents the regularization parameter. This represents the loss function relative to the server. This represents the local hyperparameters of the i-th client. This indicates that the i-th client is in the local model. The original loss function is given by N, where N represents the number of clients and F(v) represents the global loss function. This is a regularization term used to penalize the difference between the i-th client model and the global model. v The difference between them, the role of this regularization term is to allow the client's personalized model and the global model to have a certain distance, but not too far, so as to maintain the similarity of the models among all clients.

[0033] In an optional embodiment of the present invention, step S53 involves transmitting the hyperparameters of the local convolutional network model to the server for integration and optimization to obtain a global convolutional network model, including: On the server side, the hyperparameters of the local convolutional network model are integrated and optimized using an aggregation formula, which is:

[0034] in, This indicates the updated hyperparameters. Indicates the number of iterations. Let represent the number of the k-th iteration, and n represent the total number of samples. Denotes the hyperparameters of the k-th iteration. Indicates learning efficiency. This represents the gradient of the k-th iteration.

[0035] In an optional embodiment of the present invention, step S3, based on the identified target and in conjunction with pre-set alarm rules, obtains target alarm information, including: Step S31: The identified target is input into a large language model for summary and analysis. The large language model obtains initial warning information according to the pre-set alarm rules. Step S32: Perform cross-validation and dynamic decision-making on the initial warning information to obtain target alarm information.

[0036] In this embodiment, targets meeting preset conditions are selected for further processing and alerting. The intelligent recognition model matches the identified targets with pre-configured targets (e.g., specific faces, license plate numbers, etc.). Once a target meeting the conditions is identified, the intelligent recognition model pushes this data (including the target's location, features, etc.) to the large language model for summarizing and analyzing the alert information. To ensure the accuracy of the alert information and avoid false alarms, the large language model further verifies the received data. If the confidence level given by the intelligent recognition model is less than 90%, the large language model needs to perform additional reasoning and verification to determine the credibility of the alert information. If it is not credible, the alert information is discarded. If it is credible, the next step is executed. If the alert information is verified to be credible, the large language model will push the relevant information to the mobile phones of the corresponding management personnel according to the pre-configured security management alert file.

[0037] In an optional embodiment of the present invention, step S32 involves cross-validating and dynamically deciding on the initial warning information to obtain target alarm information, including: Step S321: Perform data quality verification, multi-modal consistency test and knowledge base matching on the initial warning information to obtain cross-validation data; Step S322: Dynamically weight the cross-validation data with confidence level to obtain target alarm information.

[0038] In this embodiment, the large language model combines multi-dimensional cross-validation and dynamic decision-making mechanisms to validate low-confidence warning data.

[0039] 1. Data quality verification Check the integrity of the input data (such as target feature vectors, timestamps, and camera location information), and filter out invalid data caused by transmission packet loss or sensor malfunctions. For example, warnings about blurry images or missing location coordinates are directly marked as untrustworthy.

[0040] 2. Multi-model collaborative reasoning The backup model is invoked to perform independent inference on the same target, and the output results are compared with those of the original large language model. If the differences between multiple models exceed a threshold (e.g., bounding box overlap <70%), a manual review process is triggered.

[0041] 3. Knowledge base matching The system queries a pre-defined rule base (such as license plate numbers in an allowed list) or a real-time database (such as employee attendance records) to verify the legitimacy of the target's identity. For example, it raises the risk level when an unfamiliar license plate is detected entering a restricted area outside of designated access hours.

[0042] 4. Confidence level dynamically weighted The multi-head latent attention (MLA) mechanism is used to assign weights to the results of different validation dimensions, such as image sharpness (30%), multi-model consistency (40%), and knowledge base matching (30%), and the final confidence score is calculated by combining the results.

[0043] The specific implementation process of the above-mentioned method of the present invention is illustrated below with specific examples: The on-site monitoring network consists of multiple monitoring nodes, each equipped with a camera, microprocessor, and communication module. These cameras are installed at key locations on the construction site, while the microprocessors and communication modules are installed in the on-site computer room, forming a dense monitoring network.

[0044] The video platform supports GB / T28181, ehomeX, isup, active registration, and national standard cascading access methods, aggregating all on-site monitoring networks to achieve functions such as real-time video, video playback, and PTZ control.

[0045] The distributed learning architecture features a built-in microprocessor in each monitoring node, enabling it to independently collect and initially process local data. At regular intervals, each node sends the aggregated model parameters to the central node. The central node then summarizes all parameters, generates a global model, and distributes the updated model back to all nodes, thereby protecting data and reducing communication overhead.

[0046] The system employs an adaptive alarm grading mechanism, with multiple alarm rule levels that automatically select the most appropriate response based on the location, time, and severity of the event. For example, minor warnings can be sent to on-site personnel via SMS, while severe emergencies will immediately activate the emergency response plan.

[0047] Decision support tools allow managers to view various monitoring data and analysis results overlaid on the real-time scene through a digital screen, helping them to make quick decisions.

[0048] The decoder employs high-compression video codec technologies (such as H.265 or SVAC) to significantly reduce bandwidth consumption. H.265 can save approximately 50% of bandwidth compared to H.264 while maintaining the same video quality. The decoder enables the conversion of various formats to meet the needs of various scenarios.

[0049] Intelligent recognition models, based on considerations of robustness, detection accuracy, generalization ability, and speed in complex scenarios, are particularly suitable for application scenarios that require real-time processing and high-precision target detection.

[0050] Large Language Model: From the perspective of training and usage costs and efficiency, adopting an open-source model can integrate data such as construction specifications, video surveillance, and intelligent early warning to achieve functions such as training, early warning, early warning information distribution, and early warning information Q&A.

[0051] The camera is a high-resolution infrared camera designed for low-light conditions to ensure all-weather monitoring needs.

[0052] The microprocessor uses a high-performance ARM chip, which has powerful data processing capabilities and low power consumption.

[0053] The communication module uses a wireless communication module that supports 5G or Wi-Fi to ensure high-speed and stable data transmission.

[0054] The camera captures image information, and internal algorithms are used to extract and fuse features to form comprehensive perception data.

[0055] Distributed model training involves each monitoring node independently running a distributed learning algorithm, periodically exchanging parameters with other nodes, and continuously optimizing the global model in conjunction with a large language model.

[0056] Dynamic alarm classification: The system assesses the risk level based on real-time data and triggers alarms of the corresponding level according to preset rules.

[0057] Large-screen visualization allows the management system to display monitoring data and analysis results in a real-world scenario, helping managers quickly grasp the on-site situation.

[0058] Large language model for alarm description: Utilizing a large language model to convert video surveillance footage into natural language descriptions, helping users quickly understand the monitored scenario. For example, real-time alarm description: When abnormal behavior is detected, detailed alarm information is generated, such as "Someone has been detected attempting to climb the fence; please handle immediately." Automated report generation: Integrating log data and structured information from the video surveillance system, structured reports are automatically generated based on video surveillance data for daily operations or security audits. Intelligent recognition model training and optimization: Automatically generating labeled data by parsing video content using a language model to generate preliminary labeled information. Providing training suggestions: For example, "The current model performs poorly in small object detection; it is recommended to increase relevant training samples." Due to the poor network environment in the field, the decoder transmits signals in H.265 format, and then converts the H.265 to H.264 signals for distribution, reducing the pressure on the client. Ordinary machines can play multiple video signals, thus balancing the on-site signal and the client hardware.

[0059] The intelligent recognition model analyzes video footage 24 / 7 on-site, including detection of various hazards such as smoking, open flames, falls, electronic fences, hazardous operations, and shift meetings. Additionally, a pan-tilt-zoom (PTZ) control system centers the target within the video feed, ensuring effective recording of critical information.

[0060] Operating procedures and precautions: S0: Initialize the video middleware platform, generate system parameters, start the service, and distribute real-time video data to clients through the video middleware platform for client use. Recordings are saved locally and are not centrally stored.

[0061] S1: Install monitoring nodes. Install cameras at key locations on the construction site to ensure that each node can cover a certain monitoring range.

[0062] S2: Configure the communication module. Connect the communication module of the monitoring node so that it can access the network and establish communication with the central node.

[0063] S3: Initialize the distributed learning model. Load the initial model parameters on each monitoring node and start the training program.

[0064] S4: Initialize the intelligent recognition model so that it can access video data via RTSP / ONVIF / RTMP.

[0065] S5: Deploy the large language and start the service.

[0066] S6: Set alarm rules. Define different alarm levels according to actual needs and input them into the system database.

[0067] S7: Enable large-screen decision-making tools. Administrators can access the system interface to view real-time monitoring data and analysis results.

[0068] S8: Continuous monitoring and maintenance. Regularly check the working status of each node and replace damaged sensors or upgrade the software version in a timely manner.

[0069] For cameras, different models can be selected according to actual needs, such as high-definition visible light cameras or thermal imaging cameras.

[0070] The communication module supports multiple communication protocols such as 4G / LTE / Wi-Fi, and all modules are interchangeable.

[0071] Microprocessors support hardware replacement with different computing power, such as minicomputers, workstations, and large servers.

[0072] The system implementing this invention mainly comprises the following components: an environmental perception network, a distributed learning architecture, an adaptive alarm grading mechanism, a video middleware platform, an intelligent recognition model, a large language model, and a large-screen decision-making tool. The environmental perception network consists of multiple monitoring nodes, each equipped with several cameras to capture image data. The distributed learning architecture allows each node to independently collect (acquire image data) and process local data (recognize image content), continuously improving the global model performance through periodic parameter exchanges. The adaptive alarm grading mechanism automatically selects the most appropriate response based on the time, location, and severity of the event. The video middleware platform aggregates and distributes videos from various construction sites, allowing each manager to easily control the construction site through reasonable permission allocation. The large language model service, with its powerful knowledge reserves, describes and summarizes alarm information, automatically generates reports within a set time, and, during the training of the intelligent recognition model, determines whether parameters such as model overfitting, model oscillation, precision, and recall are reasonable, while providing appropriate suggestions to assist in the training of the intelligent recognition model. The large-screen decision-making tool enables managers to intuitively see monitoring data superimposed on the real-world scene on a large screen, thereby making decisions more quickly. By introducing multiple types of cameras, the system maintains high sensitivity even in low light or obstructed conditions, enhancing its early warning capabilities. The application of distributed learning ensures data security and enhances the system's intelligence; while the adaptive alarm grading mechanism ensures the speed and accuracy of emergency response. Finally, the large-screen decision-making tool provides management with more intuitive and effective management methods, comprehensively improving the safety and management efficiency of the construction site.

[0073] Video streams are accessed through the video middleware subsystem via GB / T28181, ehome, isup, and active registration. This subsystem provides signaling control and streaming media services, enabling real-time video monitoring, platform cascading, decoders, video playback, and PTZ control. The intelligent recognition model acquires data from the video middleware subsystem via RTSP and RTMP protocols. It features AI access, alarm reporting, status detection, and AI analysis. The large language model service receives alarm information from the intelligent recognition model, analyzes the information, summarizes it, and pushes and saves it according to the early warning rule file requirements. Additionally, the intelligent recognition model is configured for distributed training and supports knowledge-based question answering. The intelligent video monitoring and dispatch platform serves as the user's entry point, providing distributed training, device management, and visualization capabilities, supporting mobile devices, personal computers, and digital large screens.

[0074] The process begins at the starting node, where the system first selects monitoring nodes to participate in distributed learning. The selected monitoring nodes download the latest global model parameters from the central server. The monitoring nodes train the model using local data. During this process, the training results are submitted to the large language model for analysis, providing suggestions and measures to guide the training process. The monitoring nodes do not share the original data; they only upload the model parameters, ensuring data security and efficient transmission of useful data. Local training and model uploading are performed on the microprocessors at the construction site. The central server collects and merges the model parameters uploaded by all monitoring nodes, generating a new global model. If the model evaluation results meet preset standards (such as accuracy, loss function, etc.), the model is published for use by the monitoring nodes; otherwise, it returns to the model synchronization step to continue iterating.

[0075] To ensure the acquisition of high-quality video data in real time, providing a foundation for subsequent identification and analysis, the intelligent recognition model first acquires real-time video stream data from cameras via protocols such as RTSP (Real-Time Streaming Protocol), RTMP (Real-Time Messaging Protocol), or ONVIF (Open Network Video Interface Forum). Ensuring the target remains within the monitoring range and its movement trajectory is continuously tracked, the intelligent recognition model processes the acquired video data to identify specific targets (such as people and vehicles). Simultaneously, if necessary, the intelligent recognition model controls a pan-tilt-zoom (PTZ) camera to track these targets and maintain their position in the frame. Only targets meeting preset criteria are further processed and trigger alarms. The intelligent recognition model matches identified targets against pre-configured target information (e.g., specific faces, license plate numbers). Once a matching target is identified, the intelligent recognition model pushes this data (including target location, features, etc.) to a large language model service for summarizing and analyzing alarm information. To ensure the accuracy of warning information and avoid false alarms, the large language model service further verifies the received data. If the confidence level given by the intelligent recognition model is below 90%, the large language model service needs to perform additional inference and verification to determine the credibility of the warning information. If it is not credible, the warning information is discarded. If it is credible, the next step is executed. If the warning information is verified to be credible, the large language model service will push the relevant information to the mobile phones of the corresponding administrators according to the pre-configured security management warning file.

[0076] like Figure 5 As shown, embodiments of the present invention also provide an artificial intelligence-based surveillance video analysis device 50, comprising: Module 51 is used to acquire surveillance video of the target area; The processing module 52 is used to extract features from the surveillance video using an intelligent recognition model to obtain the target; based on the target and combined with pre-set alarm rules, obtain target alarm information; and push the target alarm information to the client; wherein the intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model, and the local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

[0077] Optionally, an intelligent recognition model is used to extract features from the surveillance video to obtain the target for recognition, including: The surveillance video is input into the first processing module of the intelligent recognition model and subjected to standard convolution to obtain initial basic features; The initial basic features are input into the second processing module of the intelligent recognition model for enhancement and fusion to obtain intermediate enhanced features; The intermediate enhanced features are input into the third processing module of the intelligent recognition model for recognition and localization to obtain the recognition target.

[0078] Optionally, the training process of the intelligent recognition model includes: Retrieve historical monitoring videos from monitoring nodes; The historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model. The hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain the global convolutional network model. Based on the local convolutional network model and the global convolutional network model, an intelligent recognition model is obtained.

[0079] Optionally, the historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model, including: The local convolutional network model is updated according to a preset objective to obtain local hyperparameters; The intelligent recognition model is obtained based on the local hyperparameters.

[0080] Optionally, the hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain a global convolutional network model, including: On the server side, the hyperparameters of the local convolutional network model are integrated and optimized using an aggregation formula, which is:

[0081] in, This indicates the updated hyperparameters. Indicates the number of iterations. Let represent the number of the k-th iteration, and n represent the total number of samples. Denotes the hyperparameters of the k-th iteration. Indicates learning efficiency. This represents the gradient of the k-th iteration.

[0082] Optionally, based on the identified target and in conjunction with pre-set alarm rules, target alarm information is obtained, including: The identified target is input into a large language model for summary and analysis. The large language model obtains initial warning information based on pre-set alarm rules. The initial warning information is cross-validated and dynamically decided to obtain the target alarm information.

[0083] Optionally, the initial warning information is cross-validated and dynamically decided to obtain target alarm information, including: The initial warning information is subjected to data quality verification, multi-modal consistency test and knowledge base matching to obtain cross-validation data; The cross-validation data is dynamically weighted by confidence level to obtain the target alarm information.

[0084] It should be noted that all implementation methods in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effect.

[0085] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the artificial intelligence-based surveillance video analysis method of the present invention. All implementations in the above method embodiments are applicable to the embodiments of this computing device and can achieve the same technical effects.

[0086] Embodiments of the present invention also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the artificial intelligence-based surveillance video analysis method described in this invention. All implementations in the above method embodiments are applicable to the embodiments of this computer-readable storage medium and can achieve the same technical effects.

[0087] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0088] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0089] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0091] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0092] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0093] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of the present invention.

[0094] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code for implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps for performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.

[0095] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for analyzing surveillance video based on artificial intelligence, characterized in that, include: Acquire surveillance video of the target area; The intelligent recognition model is used to extract features from the surveillance video to obtain the target to be identified; Based on the identified target and combined with pre-set alarm rules, target alarm information is obtained; The target alarm information is pushed to the client. The intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model. The local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

2. The AI-based surveillance video analysis method according to claim 1, characterized in that, Using an intelligent recognition model, features are extracted from the surveillance video to obtain the target for recognition, including: The surveillance video is input into the first processing module of the intelligent recognition model and subjected to standard convolution to obtain initial basic features; The initial basic features are input into the second processing module of the intelligent recognition model for enhancement and fusion to obtain intermediate enhanced features; The intermediate enhanced features are input into the third processing module of the intelligent recognition model for recognition and localization to obtain the recognition target.

3. The AI-based surveillance video analysis method according to claim 1, characterized in that, The training process of the intelligent recognition model includes: Retrieve historical monitoring videos from monitoring nodes; The historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model. The hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain the global convolutional network model. Based on the local convolutional network model and the global convolutional network model, an intelligent recognition model is obtained.

4. The AI-based surveillance video analysis method according to claim 3, characterized in that, The historical monitoring videos of the monitoring nodes are input into a preset convolutional network model for training to obtain a local convolutional network model, including: The local convolutional network model is updated according to a preset objective to obtain local hyperparameters; The intelligent recognition model is obtained based on the local hyperparameters.

5. The AI-based surveillance video analysis method according to claim 3, characterized in that, The hyperparameters of the local convolutional network model are transmitted to the server for integration and optimization to obtain the global convolutional network model, including: On the server side, the hyperparameters of the local convolutional network model are integrated and optimized using an aggregation formula, which is: in, This indicates the updated hyperparameters. Indicates the number of iterations. Let represent the number of the k-th iteration, and n represent the total number of samples. Denotes the hyperparameters of the k-th iteration. Indicates learning efficiency. This represents the gradient of the k-th iteration.

6. The AI-based surveillance video analysis method according to claim 5, characterized in that, Based on the identified target and combined with pre-set alarm rules, target alarm information is obtained, including: The identified target is input into a large language model for summary and analysis. The large language model obtains initial warning information based on pre-set alarm rules. The initial warning information is cross-validated and dynamically decided to obtain the target alarm information.

7. The AI-based surveillance video analysis method according to claim 6, characterized in that, Cross-validation and dynamic decision-making are performed on the initial warning information to obtain target alarm information, including: The initial warning information is subjected to data quality verification, multi-modal consistency test and knowledge base matching to obtain cross-validation data; The cross-validation data is dynamically weighted by confidence level to obtain the target alarm information.

8. A surveillance video analysis device based on artificial intelligence, characterized in that, include: The acquisition module is used to acquire surveillance video of the target area; The processing module is used to extract features from the surveillance video using an intelligent recognition model to obtain the target; based on the target and combined with pre-set alarm rules, obtain target alarm information; and push the target alarm information to the client; wherein, the intelligent recognition model is determined based on a local convolutional network model and a global convolutional network model, and the local convolutional network model is trained on a preset convolutional network model based on historical surveillance videos.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Mobile security communication performance intelligent prediction method

    CN117395163A

  • Intelligent campus safety monitoring and early warning system and method thereof

    CN120014809A