AI behavior abnormity capturing and alarming method for panoramic video
By performing hierarchical segmentation and parallel recognition of panoramic videos, and utilizing Mask R-CNN and YOLOv8 models, the problems of high computational overhead and high false alarm rate in panoramic video abnormal behavior recognition are solved, achieving efficient and accurate abnormal behavior capture and real-time alarm.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DEZHOU POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing panoramic video abnormal behavior recognition methods suffer from high computational overhead, recognition delays, and high false alarm rates when processing large-scale videos. They also lack targeted segmentation and partitioning of panoramic images, making it difficult to meet the needs of real-time monitoring and rapid alarm.
The Mask R-CNN model is used for the first segmentation, dividing the panoramic video frames into areas for people, towers, lines, and background. The YOLOv8 model is used for the second segmentation, further subdividing the data into functional sub-regions such as the head, body skeleton, and hands. Anomaly probabilities are identified in parallel, and model parameters are dynamically optimized to improve the accuracy and timeliness of the identification.
By employing a hierarchical identification method that combines global segmentation, local subdivision, and parallel recognition, redundant information is eliminated, enabling efficient capture of abnormal behaviors, improving the accuracy and timeliness of identification, and reducing false alarm rates and delays.
Smart Images

Figure CN121963293A_ABST
Abstract
Description
A method for AI-based behavioral anomaly detection and alerting in panoramic video. Technical Field
[0001] This invention relates to the field of image recognition and processing, and specifically to an AI-based method for capturing and alerting abnormal behavior in panoramic video. Background Technology
[0002] With the rapid development of artificial intelligence and computer vision technologies, AI-based anomaly monitoring using panoramic video has been widely applied in scenarios such as power construction, public safety, and industrial production. Panoramic video can provide complete scene coverage and a global perspective, helping to improve environmental perception and risk prevention capabilities. However, panoramic video itself has a large frame and redundant information. Directly identifying abnormal behavior across the entire frame often leads to excessive computational load and insufficient response speed, making it difficult to meet the needs of real-time monitoring and rapid alarm.
[0003] In existing panoramic video anomaly behavior recognition technologies, personnel detection and behavior discrimination are the core components for achieving intelligent monitoring and alarms, and have been applied in scenarios such as power construction safety, public place security, and industrial operation supervision. However, existing methods have the following shortcomings when processing panoramic videos: First, they lack targeted segmentation and partitioning of the panoramic image, causing the recognition model to often need to run directly on the large-scale video, resulting in excessive computational overhead and a high risk of recognition delays and false positives / missed negatives. Furthermore, over-reliance on a single model for recognition, while improving feature extraction capabilities to some extent, is prone to false alarms and performance degradation in complex backgrounds. Second, traditional model optimization methods lack precise localization of recognition errors, typically requiring complete training of the entire network. This not only increases training costs and computational burden but also makes it difficult to promptly and specifically correct recognition defects, resulting in insufficient robustness of the system in actual operation and an inability to balance real-time performance and accuracy.
[0004] An existing method for managing abnormal behavior in power distribution network operations based on panoramic video analysis first acquires panoramic video of the operation scene and personnel information, generating scene information and personnel appearance information accordingly. Then, relevant power distribution network operation specifications are collected to construct an operation standard information database. Based on this standard information, the personnel information and appearance information in the scene are analyzed to generate first behavioral information. Simultaneously, personnel posture structure is generated based on the personnel information, and gradient histogram direction features are extracted from the panoramic video to obtain predicted personnel behavior information. Further, the predicted behavior information is analyzed based on the operation standard information to generate second behavioral information. Finally, by comparing the first and second behavioral information, an alert is issued when an anomaly is detected, allowing management personnel to review and intervene in the operation process, thereby achieving the identification and management of abnormal behavior in power distribution network operations. While the existing technology can manage personnel behavior using panoramic video and work specifications, it has significant shortcomings: First, it lacks targeted segmentation and partitioning of panoramic images, requiring the recognition model to directly operate on large-scale videos, resulting in high computational costs, delays, and missed detections. It also relies too heavily on a single model, leading to a high false positive rate. Second, the model optimization lacks precise localization of recognition errors. Once a false positive or missed detection occurs, the entire network often needs to be retrained, resulting in high training costs and difficulty in timely repair of defects. Consequently, the system suffers from deficiencies in both real-time performance and robustness.
[0005] Existing technology also includes a real-time early warning method for regional anomalies based on high-point panoramic intelligent inspection. This method first acquires visible light video streams and infrared thermal imaging video streams of the target area to form a panoramic video sequence, and then obtains anomaly feature-enhanced video sequences through intelligent analysis. Based on this, target feature description information is generated, and this information is used to construct a spatiotemporal topology map of targets within the region. The spatiotemporal correlation characteristics between targets are analyzed to detect abnormal behavior patterns in the region. Further, the detected abnormal behavior patterns are embedded as regional anomaly features into the spatiotemporal topology map, and combined with the target feature description information, environmental risk prediction results are generated. Finally, based on the spatiotemporal topology map integrating regional abnormal behavior patterns and the environmental risk prediction results, a spatiotemporal evolution model is constructed, and hierarchical early warning information for regional anomalies is generated according to this model, realizing multi-dimensional dynamic assessment and real-time early warning of regional risks. While the existing technology can predict and warn of regional anomalies after introducing multi-source data and spatiotemporal topology modeling, it still has some shortcomings: On the one hand, it relies too much on global video streams and topology modeling, without fine-grained segmentation and local feature extraction of the target, resulting in limited efficiency and insufficient real-time performance in anomaly recognition; on the other hand, in the model optimization stage, its anomaly detection results mainly rely on global risk prediction, lacking a mechanism for locating and adaptively updating specific identification errors. Once false alarms or missed alarms occur, it is difficult to correct them in time, which can easily lead to inaccurate warning information and affect the reliability and practicality of the system.
[0006] Therefore, it can be seen that most existing panoramic video abnormal behavior recognition methods rely on the entire frame for feature extraction and recognition, resulting in severe background information redundancy. This leads to high model computational complexity, slow response speed, and difficulty in meeting real-time alarm requirements. Furthermore, directly processing full-frame information with a single model can easily lead to decreased recognition accuracy and delays in abnormal behavior capture. Summary of the Invention
[0007] To overcome the shortcomings of the above technologies, this invention provides an alarm method that effectively improves the accuracy and timeliness of abnormal behavior capture through hierarchical identification of "global segmentation - local subdivision - parallel identification".
[0008] The technical solution adopted by this invention to overcome its technical problems is: an AI behavior anomaly capture and alarm method for panoramic video, comprising: S1. extracting from panoramic video... Frame video frame, number Frame video frame is , S2. Using a single-cut region mask For the first Frame video frames Cut to obtain the first Frame Personnel Area , No. Frame tower area , No. Frame line area , No. Frame background area S3. Utilize a secondary cutting region mask For the first Frame Personnel Area Perform a second cut to obtain Body parts and areas ,in For the first Each body part / area S4. By identifying the model, the first Body parts and areas Perform anomaly identification and obtain anomaly probability. S5. Utilizing anomaly probabilities Determine if there is an anomaly; if so, output an alarm and store the abnormal video clip; S6. Optimize the recognition model to obtain the optimized recognition model.
[0009] Furthermore, in step S2, through Get the first Frame Personnel Area , No. Frame tower area , No. Frame line area , No. Frame background area ,in For video cutting models, These are the parameters for the video cutting model.
[0010] Preferably, the video segmentation model described above is the Mask R-CNN model.
[0011] Furthermore, in step S3, through get Each body part area, among which For video cutting models, These are the parameters for the video cutting model.
[0012] Preferably, the video segmentation model described above is the Mask R-CNN model.
[0013] In step S4, through Obtain the probability of anomalies ,in To identify the model, For the first Body parts and areas The network parameters of the recognition model.
[0014] Preferably, the recognition model mentioned above is the YOLOv8 model.
[0015] Furthermore, step S5 includes the following steps: S5-1. Determine the anomaly probability. Is it greater than the first? Abnormal thresholds for individual body parts If so, it will make Output an alarm and execute step S5-2, where For the first The first frame of the video frame Alarm variables for individual body parts / regions, otherwise make S5-2. Extract the first... Frame video frames To the Frame video frames Composition of real-time video clips As abnormal video segments, these segments are stored. The length of the video window. Set to 5-10 seconds.
[0016] Furthermore, step S6 includes the following steps: S6-1. Using the formula Calculation yields the first Body parts and areas With the Auxiliary importance value of each region , of which The region includes the first Body parts and areas With the Frame Personnel Area Corresponding tower area , No. Frame Personnel Area Corresponding line area , No. Frame Personnel Area Corresponding background area The union of, For the first Body parts and areas The central spatial coordinates, For the first The central spatial coordinates of each region for and Spatial distance; S6-2. Select the highest auxiliary importance value, the highest auxiliary importance value corresponds to the first Each body part area is , S6-3. The first Body parts and areas With the Body parts and areas Merging to obtain regions S6-4. Utilize the recognition model IDE for region Re-identify and determine if the identification result is correct. If not, proceed to step S6-5; otherwise, use the Adam optimizer to update the secondary cutting mask part of the IED model parameters. , In the formula, The update step size for the segmentation model, For video cutting model parameters Model parameters related to secondary cutting For the segmentation loss function, For the segmentation loss function for The gradient; S6-5. Utilizing the region The replacement step S6-3 Body parts and areas Then repeat steps S6-3 to S6-4; S6-6. If the recognition result of the recognition model IDE re-recognizing the region after merging all regions and all body part regions is incorrect, then execute step S6-7; S6-7. [The text abruptly ends here, likely due to an incomplete sentence or missing information.] With the Frame Personnel Area Corresponding tower area , No. Frame Personnel Area Corresponding line area , No. Frame Personnel Area Corresponding background area Merging to obtain regions S6-8. Utilize the recognition model IDE for region Re-identify and determine if the identification result is correct. If not, proceed to steps S6-8; otherwise, use the Adam optimizer to update the first-stage cut mask portion of the IED model parameters. , In the formula, For video cutting model parameters The model parameters related to a single cut are as follows: For the segmentation loss function, For the segmentation loss function for The gradient; S6-9. Utilizing the region Replace the area in step S6-7 Then repeat steps S6-7 to S6-8; S6-10. When the identification model IDE identifies the tower areas corresponding to all regions. Line area Background area If the re-identification result of the merged region is incorrect, proceed to step S6-11; S6-11. Update the result using the Adam optimizer. Body parts and areas Network parameters of the recognition model , , To identify the update step size of the model IDE, Let cross-entropy be the loss function. Cross-entropy loss function The gradient.
[0017] Preferably, the segmentation loss function in step S6-4 is Dice Loss.
[0018] The beneficial effects of this invention are as follows: First, using a Mask R-CNN multi-mask video segmentation model, the panoramic video frame is segmented once under a single segmentation mask, dividing the image into areas such as personnel, towers, lines, and background, eliminating redundant information and achieving efficient separation of multiple target areas. Second, under a second segmentation mask, the personnel area is further segmented, subdivided into functional sub-regions such as head, torso, and hands according to safety management needs, providing structured input for subsequent targeted identification. Finally, a lightweight model-based parallel recognition of abnormal behavior partitions is proposed. Multiple parallel regional abnormal behavior recognition models output the abnormal probability of each part. When the probability exceeds a threshold, an alarm is triggered immediately, and video segments are captured and stored in real time. This method, through hierarchical recognition of "global segmentation—local subdivision—parallel recognition," effectively improves the accuracy and timeliness of abnormal behavior capture.
[0019] Secondly, this invention proposes an optimization method for a rapid abnormal behavior recognition model based on panoramic information assistance. First, the importance of auxiliary information is calculated based on the spatial distance between different regions, and this is used as the priority for the merging order. Second, region merging is performed according to a hierarchical and step-by-step strategy: first, body part regions are gradually merged and repeatedly identified; if the result is correct, the model's secondary cutting parameters are identified as incorrect; if not, environmental regions are merged; if correct at this point, the model's primary cutting parameters are identified as incorrect; if the result is still not correct after merging all regions, the recognition model is deemed to have insufficient performance. Finally, based on the localization results, the cutting model parameters or recognition model parameters for the corresponding mask parts are updated accordingly. This method dynamically introduces panoramic information when the model malfunctions, gradually locates the source of the error, and performs adaptive optimization, achieving efficient model updates and adaptive correction of cutting rules, thereby significantly improving the accuracy and robustness of abnormal behavior recognition. Attached Figure Description
[0020] Figure 1 is a flowchart of the method of the present invention. The left box in the figure represents the method for rapid capture of abnormal behavior based on dynamic segmentation of panoramic video, and the right box in the figure represents the method for rapid optimization of abnormal behavior model based on panoramic information assistance. Detailed Implementation
[0021] The present invention will be further described below with reference to Figure 1.
[0022] A method for AI-based behavior anomaly detection and alerting in panoramic video includes: S1. Extracting from the panoramic video... Frame video frame, number Frame video frame is , .
[0023] S2. Using a single-cut region mask For the first Frame video frames Cut to obtain the first Frame Personnel Area , No. Frame tower area , No. Frame line area , No. Frame background area .
[0024] S3. Utilize a secondary region mask. For the first Frame Personnel Area Perform a second cut to obtain Body parts and areas ,in For the first Each body part / area .
[0025] S4. By identifying the model, the first Body parts and areas Perform anomaly identification and obtain anomaly probability. .
[0026] S5. Utilizing outlier probability Determine if there is an abnormality; if so, output an alarm and store the abnormal video clip.
[0027] S6. Optimize the recognition model to obtain the optimized recognition model.
[0028] By first segmenting panoramic video frames into regions such as personnel, power poles, lines, and background, redundant information is eliminated and efficient separation of multiple target areas is achieved. Secondly, under a secondary segmentation mask, the personnel region is further segmented into functional sub-regions such as head, torso, and hands, based on safety management requirements, providing structured input for subsequent targeted identification. Finally, a lightweight model-based parallel recognition of abnormal behavior is proposed. Multiple parallel regional abnormal behavior recognition models output the probability of anomalies for each part of the body. When the probability exceeds a threshold, an alarm is triggered immediately, and video segments are captured and stored in real time. This method, through hierarchical recognition of "global segmentation—local subdivision—parallel recognition," effectively improves the accuracy and timeliness of abnormal behavior capture.
[0029] In existing methods, once false positives or false negatives occur in model recognition, there is a lack of an auxiliary optimization mechanism based on panoramic information, often requiring manual correction or offline retraining. This not only increases model maintenance costs but also makes it difficult for the system to correct errors in a timely manner during actual operation, affecting the reliability and stability of abnormal behavior recognition. Therefore, step S6 of this invention proposes a method for optimizing a rapid abnormal behavior recognition model based on panoramic information.
[0030] In one embodiment of the present invention, step S2 is performed by... Get the first Frame Personnel Area , No. Frame tower area , No. Frame line area , No. Frame background area ,in For video cutting models, These are the parameters for the video cutting model.
[0031] In this embodiment, the preferred video segmentation model is the Mask R-CNN model.
[0032] In one embodiment of the present invention, in order to further improve recognition efficiency and accuracy, step S3 is performed by... get Each body part area, among which For video cutting models, These are the parameters for the video cutting model.
[0033] In this embodiment, the preferred video segmentation model is the Mask R-CNN model. Specifically, a multi-task mask prediction module is introduced into the shared feature extraction layer of the Mask R-CNN model. This module can perform different segmentation tasks on the input video frames based on different network parameter configurations and mask types. By switching the parameter set and mask set within the same structural framework, hierarchical processing of primary and secondary segmentation is achieved, improving the segmentation efficiency of panoramic videos and the accuracy of target region identification.
[0034] In one embodiment of the present invention, step S4 is performed by... Obtain the probability of anomalies ,in To identify the model, For the first Body parts and areas The network parameters of the recognition model are represented as follows: .
[0035] In this embodiment, the preferred recognition model is the YOLOv8 model, which has the advantages of fast detection speed and low inference overhead, and can meet the requirements of real-time anomaly capture while ensuring accuracy. However, the present invention does not limit the recognition model used and can support the use of recognition models other than YOLOv8.
[0036] In one embodiment of the present invention, step S5 includes the following steps: S5-1. Determine the anomaly probability. Is it greater than the first? Abnormal thresholds for individual body parts If so, it will make Output an alarm and execute step S5-2, where For the first The first frame of the video frame Alarm variables for individual body parts / regions, otherwise make .
[0037] S5-2. Extract the first... Frame video frames To the Frame video frames Composition of real-time video clips As abnormal video segments, these segments are stored. The length of the video window. Set to 5-10 seconds to cover the critical time period before and after the occurrence of abnormal behavior. Video clips can serve as evidence for subsequent verification by management personnel, and also as training samples for the identification model, enabling further model optimization and performance improvement. These steps allow for parallel detection of various abnormal behaviors in different areas, avoiding the latency of processing large models across panoramic images, and effectively improving the efficiency and accuracy of anomaly identification.
[0038] In one embodiment of the present invention, step S6 includes the following step: S6-1. Using the formula Calculation yields the first Body parts and areas With the Auxiliary importance value of each region , of which The region includes the first Body parts and areas With the Frame Personnel Area Corresponding tower area , No. Frame Personnel Area Corresponding line area , No. Frame Personnel Area Corresponding background area The union of, For the first Body parts and areas The central spatial coordinates, For the first The central spatial coordinates of each region for and The smaller the spatial distance, the closer the two regions are in spatial distribution, the stronger the correlation between the regions, and the greater the auxiliary importance to the current recognition model; conversely, the larger the distance, the weaker the correlation between the regions, and the limited contribution to recognition.
[0039] S6-2. Select the highest auxiliary importance value. The highest auxiliary importance value corresponds to the [number]th [position / level]. Each body part area is , .
[0040] S6-3. The first Body parts and areas With the Body parts and areas Merging to obtain regions .
[0041] S6-4. Using the recognition model IDE to analyze the region Re-identify and determine if the identification result is correct. If not, proceed to step S6-5. If the identification is correct, it indicates that the anomaly stems from an improper secondary segmentation method in the personnel area. Then, use the Adam optimizer to update the secondary segmentation mask portion of the IED model parameters obtained from the identification model. , In the formula, The update step size for the segmentation model, For video cutting model parameters Model parameters related to secondary cutting For the segmentation loss function, For the segmentation loss function for The gradient.
[0042] S6-5. Utilizing the area The replacement step S6-3 Body parts and areas Then repeat steps S6-3 to S6-4.
[0043] S6-6. If the recognition model IDE does not perform the recognition result correctly when re-recognizing the region after merging all regions and all body part regions, then proceed to step S6-7.
[0044] S6-7. Region With the Frame Personnel Area Corresponding tower area , No. Frame Personnel Area Corresponding line area , No. Frame Personnel Area Corresponding background area Merging to obtain regions .
[0045] S6-8. Using the recognition model IDE to identify regions Re-identify and determine if the identification result is correct. If not, proceed to steps S6-8. If it is correct, it indicates an error in the primary segmentation model. Then, use the Adam optimizer to update the primary cut mask portion of the IED model parameters obtained from the identification model. , In the formula, For video cutting model parameters The model parameters related to a single cut are as follows: For the segmentation loss function, For the segmentation loss function for The gradient.
[0046] S6-9. Utilizing the area Replace the area in step S6-7 Then repeat steps S6-7 to S6-8.
[0047] S6-10. When the identification model IDE identifies the corresponding tower areas for all regions. Line area Background area If the re-identification result of the merged region is incorrect, it indicates that the recognition model is not performing well, so proceed to step S6-11.
[0048] S6-11. Update the result using the Adam optimizer to obtain the... Body parts and areas Network parameters of the recognition model , , To identify the update step size of the model IDE, Let cross-entropy be the loss function. Cross-entropy loss function The gradient.
[0049] First, the importance of auxiliary information is calculated based on the spatial distance between different regions, and this is used as the priority for merging. Second, region merging is performed according to a hierarchical and step-by-step strategy: body part regions are merged step by step and re-identified. If the result is correct, the error is identified as a secondary cutting parameter error in the model; if not, the environment regions are merged. If the result is correct at this point, the error is identified as a primary cutting parameter error in the model; if the result is still not correct after merging all regions, the recognition model is deemed to have insufficient performance. Finally, based on the localization results, the cutting model parameters or recognition model parameters for the corresponding mask parts are updated accordingly. This method dynamically introduces panoramic information when the model malfunctions, gradually locates the source of the error, and performs adaptive optimization, achieving efficient model updates and adaptive correction of cutting rules, thereby significantly improving the accuracy and robustness of abnormal behavior recognition.
[0050] In this embodiment, preferably, the segmentation loss function in step S6-4 is Dice Loss.
[0051] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for AI-based behavioral anomaly detection and alarm in panoramic video, characterized in that, include: S1. Extract from panoramic video Frame video frame, number Frame video frame is , S2. Using a single-cut region mask For the first Frame video frames Cut to obtain the first Frame Personnel Area , the Frame tower area , the Frame line area , the Frame background area ; S3. Utilize a secondary region mask. For the first Frame Personnel Area Perform a second cut to obtain Body parts and areas ,in For the first Each body part / area ; S4. By identifying the model, the first Body parts and areas Perform anomaly identification and obtain anomaly probability. S5. Utilizing anomaly probabilities Determine if there is an anomaly; if so, output an alarm and store the abnormal video clip; S6. Optimize the recognition model to obtain the optimized recognition model.
2. The AI behavior anomaly capture and alarm method for panoramic video according to claim 1, characterized in that: In step S2, through Get the first Frame Personnel Area , the Frame tower area , the Frame line area , the Frame background area ,in For video cutting models, These are the parameters for the video cutting model.
3. The AI behavior anomaly capture and alarm method for panoramic video according to claim 2, characterized in that: The video segmentation model is the Mask R-CNN model.
4. The AI behavior anomaly capture and alarm method for panoramic video according to claim 1, characterized in that: In step S3, through get Each body part area, among which For video cutting models, These are the parameters for the video cutting model.
5. The AI behavior anomaly capture and alarm method for panoramic video according to claim 4, characterized in that: The video segmentation model is the Mask R-CNN model.
6. The AI behavior anomaly capture and alarm method for panoramic video according to claim 4, characterized in that: In step S4, through Obtain the probability of anomalies ,in To identify the model, For the first Body parts and areas The network parameters of the recognition model.
7. The AI behavior anomaly capture and alarm method for panoramic video according to claim 6, characterized in that: The recognition model is the YOLOv8 model.
8. The AI behavior anomaly capture and alarm method for panoramic video according to claim 1, characterized in that, Step S5 includes the following steps: S5-1. Determine the anomaly probability. Is it greater than the first? Abnormal thresholds for individual body parts If so, it will make Output an alarm and execute step S5-2, where For the first The first frame of the video frame Alarm variables for individual body parts / regions, otherwise make S5-2. Extract the first... Frame video frames To the Frame video frames Composition of real-time video clips As abnormal video segments, these segments are stored. The length of the video window. Set to 5-10 seconds.
9. The AI behavior anomaly capture and alarm method for panoramic video according to claim 6, characterized in that, Step S6 includes the following steps: S6-1. Using the formula Calculation yields the first Body parts and areas With the Auxiliary importance value of each region , of which The region includes the first Body parts and areas With the Frame Personnel Area Corresponding tower area , the Frame Personnel Area Corresponding line area , the Frame Personnel Area Corresponding background area union, For the first Body parts and areas The central spatial coordinates, For the first The central spatial coordinates of each region for and Spatial distance; S6-2. Select the highest auxiliary importance value, the highest auxiliary importance value corresponds to the first Each body part area is , S6-3. The first Body parts and areas With the Body parts and areas Merging to obtain regions S6-4. Utilize the recognition model IDE for region Re-identify and determine if the identification result is correct. If not, proceed to step S6-5; otherwise, use the Adam optimizer to update the secondary cutting mask part of the IED model parameters. , In the formula, The update step size for the segmentation model, For video cutting model parameters Model parameters related to secondary cutting For the segmentation loss function, For the segmentation loss function for The gradient; S6-5. Utilizing the region The replacement step S6-3 Body parts and areas Then repeat steps S6-3 to S6-4; S6-6. If the recognition result of the recognition model IDE re-recognizing the region after merging all regions and all body part regions is incorrect, then execute step S6-7; S6-7. [The text abruptly ends here, likely due to an incomplete sentence or missing information.] With the Frame Personnel Area Corresponding tower area , the Frame Personnel Area Corresponding line area , the Frame Personnel Area Corresponding background area Merging to obtain regions S6-8. Utilize the recognition model IDE for region Re-identify and determine if the identification result is correct. If not, proceed to steps S6-8; otherwise, use the Adam optimizer to update the first-stage cut mask portion of the IED model parameters. , In the formula, For video cutting model parameters Model parameters related to a single cut For the segmentation loss function, For the segmentation loss function for The gradient; S6-9. Utilizing the region Replace the area in step S6-7 Then repeat steps S6-7 to S6-8; S6-10. When the identification model IDE identifies the tower areas corresponding to all regions. Line area Background area If the re-identification result of the merged region is incorrect, proceed to step S6-11; S6-11. Update the result using the Adam optimizer. Body parts and areas Network parameters of the recognition model , , To identify the update step size of the model IDE, Let cross-entropy be the loss function. Cross-entropy loss function The gradient.
10. The AI behavior anomaly capture and alarm method for panoramic video according to claim 9, characterized in that: In step S6-4, the segmentation loss function is Dice Loss.