Video anomaly event detection method, apparatus and device, and storage medium

Through the video abnormal event detection method based on the dual-stream autoencoder network, fatigue and misdetection problems caused by manual detection are solved, and higher detection accuracy and cost-effectiveness are achieved.

WO2025113145A1PCT designated stage expired Publication Date: 2025-06-05CHINA MOBILE ZIJIN INNOVATION INST CO LTD +2

Patent Information

Application Number
PCT/CN2024/131058
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-08
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing video abnormal event detection methods cause fatigue due to long-term human eyes gaze on the screen, resulting in missing or mischecking major abnormal events.

Method used

The video abnormal event detection method based on the dual-stream autoencoder network is adopted. By acquiring the video frames at the industrial site and reconstructing the current reconstruction frame corresponding to the current video frame, the current dynamic threshold of the current video frame is calculated, and the abnormal event is detected based on the dynamic threshold and the reconstruction frame.

Benefits of technology

It reduces the problem of mis-checking or missed detection, improves the accuracy of video abnormal events detection, and avoids the fatigue and cost of manual inspections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131058_05062025_PF_FP_ABST
    Figure CN2024131058_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of videos. Disclosed are a video anomaly event detection method, apparatus and device, and a storage medium. The method comprises: acquiring the current video frame of an industrial site, and reconstructing, by means of an anomaly detection model, a current reconstructed frame corresponding to the current video frame, wherein the anomaly detection model is a model constructed on the basis of a two-stream autoencoder network, the two-stream autoencoder network being used for extracting a spatial feature and a temporal feature; calculating the current dynamic threshold of the current video frame, wherein the current dynamic threshold is used for performing anomaly detection on the current reconstructed frame; and on the basis of the current dynamic threshold and the current reconstructed frame, detecting abnormal events in the industrial site. In the present disclosure, whether abnormal events occur is detected by means of an anomaly detection model and a dynamic threshold value; therefore, the reconstruction capability of the model on the abnormal events can be suppressed, and an auto-encoder is encouraged to produce a higher reconstruction error for the abnormal events, thus reducing the problems of false detection or missed detection, and improving the accuracy of video anomaly event detection.
Need to check novelty before this filing date? Find Prior Art

Description

Video abnormal event detection method, device, equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure is based on and claims the priority of Chinese patent application with application number 202311620286.9 and application date November 29, 2023. The entire content of the Chinese patent application is hereby incorporated into this disclosure by reference. Technical Field

[0003] The present disclosure relates to the field of video technology, and in particular to a method, apparatus, device, and storage medium for detecting abnormal video events. Background Art

[0004] With the continuous development of the Industrial Internet, various advanced artificial intelligence technologies are being widely applied in industrial sites. Many large factories are also moving towards intelligent, secure, and cost-effective construction. Currently, safety management in production lines and warehouses mainly involves detecting abnormal events among production personnel, including falls, running, and entering restricted areas.

[0005] Abnormal event detection, also known as outlier detection or novelty detection, refers to the process of detecting data instances that deviate significantly from the majority. Abnormal events in industrial sites present several challenges: first, the low probability of occurrence makes data collection difficult; second, the definition of abnormal events is vague; different factory environments in the video may require the detected objects to be in different states; and finally, abnormal events are unpredictable. For example, in the same environment, various abnormal events such as falls, running, and packet loss may occur, making it impossible to list all abnormal situations.

[0006] However, currently, most people still rely on manual internal inspections or long-term patrol recordings, which is not only time-consuming and costly, but also may cause eye fatigue caused by long-term staring at the screen, leading to omission or misdetection of major abnormal events.

[0007] Summary of the Invention

[0008] The main purpose of the present disclosure is to provide a method, device, equipment and storage medium for detecting video abnormal events, aiming to solve the technical problem that existing video abnormal event detection methods may miss or misdetect major abnormal events due to fatigue of the human eye caused by long-term staring at the screen.

[0009] To achieve the above objectives, the present disclosure provides a method for detecting abnormal video events, comprising:

[0010] Obtaining a current video frame of the industrial site and reconstructing a current reconstructed frame corresponding to the current video frame using an anomaly detection model, wherein the anomaly detection model is a model built based on a two-stream autoencoder network for extracting spatial and temporal features;

[0011] Calculating a current dynamic threshold of the current video frame, wherein the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame;

[0012] An abnormal event at the industrial site is detected according to the current dynamic threshold and the current reconstructed frame.

[0013] Optionally, calculating the current dynamic threshold of the current video frame includes:

[0014] Determining a quality index of the current video frame by using a peak signal-to-noise ratio, and performing normalization processing on the quality index to obtain an abnormality score value of the current video frame;

[0015] Determining a normal value range of the current video frame based on the quartile box plot method, and determining an initial threshold value according to the normal value range;

[0016] The local statistical properties of the abnormal score value are calculated through a fixed sliding window, and the initial threshold is adjusted according to the local statistical properties to obtain a current dynamic threshold of the current video frame.

[0017] Optionally, before acquiring the video frame of the industrial site and reconstructing the reconstructed frame corresponding to the video frame using the anomaly detection model, the method further includes:

[0018] Acquiring video data of an industrial site and preprocessing the video data to obtain video frames;

[0019] synthesizing a multimodal graph and a pseudo abnormal video frame based on the video frame;

[0020] An anomaly detection model is constructed based on the video frame, the multimodal graph, and the pseudo-anomaly video frame.

[0021] Optionally, constructing an anomaly detection model according to the video frame, the multimodal graph, and the pseudo-anomaly video frame includes:

[0022] Constructing a two-stream autoencoder network, the two-stream autoencoder network including an appearance network and a motion network, the appearance network is used to train the encoder from the perspective of spatial stream, and the motion network is used to train the encoder from the perspective of temporal stream;

[0023] Inputting the video frame and the pseudo abnormal video frame into the appearance network to obtain a spatial flow bottleneck layer vector;

[0024] Inputting the multimodal graph into the motion network to obtain a time flow bottleneck layer vector;

[0025] Performing feature fusion on the time stream bottleneck layer vector and the spatial stream bottleneck layer vector to obtain a fusion vector;

[0026] Based on the fusion vector, a network is trained by back propagation of a joint reconstruction error loss function to obtain an anomaly detection model.

[0027] Optionally, the joint reconstruction error loss function is used to maximize the loss of pseudo-abnormal reconstructed frames and pseudo-abnormal original frames and minimize the loss of reconstructed frames and original frames during training, and the motion network is trained through multimodal graph loss constraints.

[0028] Optionally, the multimodal graph includes at least one of a dynamic graph, an optical flow graph, and a motion history graph, and synthesizing the multimodal graph based on the video frame includes:

[0029] Encoding the frame sequence of the video frame based on the parameters of a preset sorting function, and sorting and learning the pixels of the video frame according to the encoding result to obtain a dynamic image;

[0030] Calculating the changes between adjacent video frames in the video frame in the time domain to obtain an optical flow map;

[0031] The background and the motion foreground are segmented based on the motion energy map, and the video frames are compressed into static images to obtain a motion history map.

[0032] Optionally, synthesizing a pseudo abnormal video frame based on the video frame includes:

[0033] Obtaining a skip frame parameter of the video frame;

[0034] The video frames are controlled to skip frames at fixed intervals based on the skip frame parameters to synthesize pseudo abnormal frames.

[0035] In addition, to achieve the above objectives, the present disclosure also proposes a video abnormal event detection device, comprising:

[0036] An acquisition module is used to acquire a current video frame of the industrial site and reconstruct a current reconstructed frame corresponding to the current video frame using an anomaly detection model. The anomaly detection model is a model built based on a two-stream autoencoder network, which is used to extract spatial and temporal features.

[0037] a calculation module, configured to calculate a current dynamic threshold of the current video frame, wherein the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame;

[0038] A detection module is configured to detect abnormal events at the industrial site based on the current dynamic threshold and the current reconstructed frame.

[0039] In addition, to achieve the above-mentioned purpose, the present disclosure also proposes a video abnormal event detection device, including a memory, a processor, and a video abnormal event detection program stored on the memory and executable on the processor, wherein the video abnormal event detection program is configured to implement the video abnormal event detection method as described above.

[0040] In addition, to achieve the above-mentioned purpose, the present disclosure also proposes a storage medium, on which a video abnormal event detection program is stored. When the video abnormal event detection program is executed by a processor, the video abnormal event detection method described above is implemented.

[0041] In the present disclosure, a method of obtaining a current video frame of an industrial site and reconstructing a current reconstructed frame corresponding to the current video frame through an anomaly detection model is disclosed. The anomaly detection model is a model constructed based on a dual-stream autoencoder network. The dual-stream autoencoder network is used to extract spatial features and temporal features, calculate the current dynamic threshold of the current video frame, and the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame. Abnormal events at the industrial site are detected based on the current dynamic threshold and the current reconstructed frame. Since the present disclosure detects whether an abnormal event occurs through an anomaly detection model and a dynamic threshold, it can suppress the model's ability to reconstruct abnormal events, encourage the autoencoder to produce a higher reconstruction error for abnormal events, reduce false detection or missed detection problems, and improve the accuracy of video abnormal event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of the structure of a video abnormal event detection device in a hardware operating environment according to an embodiment of the present disclosure;

[0043] FIG2 is a flow chart of a first embodiment of a method for detecting abnormal video events according to the present disclosure;

[0044] FIG3 is a flow chart of a second embodiment of the method for detecting abnormal video events disclosed herein;

[0045] FIG4 is a specific flow chart of a method for detecting abnormal video events according to an embodiment of the present disclosure;

[0046] FIG5 is a specific flow chart of constructing an anomaly detection model according to an embodiment of the method for detecting abnormal events in videos disclosed herein;

[0047] FIG6 is a diagram of a network architecture of an anomaly detection model according to an embodiment of a method for detecting abnormal video events of the present disclosure;

[0048] FIG7 is a flow chart of a third embodiment of the method for detecting abnormal video events disclosed herein;

[0049] FIG8 is a structural block diagram of the first embodiment of the apparatus for detecting abnormal video events disclosed herein.

[0050] The realization of the objectives, functional features and advantages of the present disclosure will be further explained with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0051] It should be understood that the specific embodiments described herein are only used to illustrate the present disclosure and are not intended to limit the present disclosure.

[0052] 1 , which is a schematic diagram of the structure of a video abnormal event detection device in a hardware operating environment according to an embodiment of the present disclosure.

[0053] As shown in Figure 1, the video abnormal event detection device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), and optionally the user interface 1003 may also include a standard wired interface and a wireless interface. The wired interface of the user interface 1003 may be a USB interface in the present disclosure. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable memory (NVM), such as a disk storage. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0054] Those skilled in the art will understand that the structure shown in FIG1 does not constitute a limitation on the video abnormal event detection device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0055] As shown in FIG. 1 , the memory 1005 , which is identified as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a video abnormal event detection program.

[0056] In the video abnormal event detection device shown in Figure 1, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the user device; the video abnormal event detection device calls the video abnormal event detection program stored in the memory 1005 through the processor 1001, and executes the video abnormal event detection method provided by the embodiment of the present disclosure.

[0057] Based on the above hardware structure, an embodiment of the video abnormal event detection method of the present disclosure is proposed.

[0058] 2 , which is a flow chart of a first embodiment of a method for detecting abnormal video events according to the present disclosure, provides a first embodiment of a method for detecting abnormal video events according to the present disclosure.

[0059] It's important to understand that with the continuous development of the Industrial Internet, various advanced artificial intelligence technologies are being widely applied in industrial settings, and many large factories are also moving towards intelligent, secure, and cost-effective development. Currently, safety management in production lines and warehouses primarily involves detecting abnormal events among production personnel, including falls, running, or entering restricted areas.

[0060] Abnormal event detection, also known as outlier detection or novelty detection, refers to the process of detecting data instances that deviate significantly from the majority of data instances. Abnormal events in industrial sites present several challenges: first, the low probability of occurrence makes data collection difficult; second, the definition of abnormal events is vague; different factory environments in the video may require the detected objects to be in different states; and third, abnormal events are unpredictable. For example, in the same environment, various abnormal events such as falls, running, and packet loss may occur, making it impossible to enumerate all abnormal conditions. However, currently, manual internal inspections or long-term video recordings are still widely used. This is not only time-consuming and costly, but also can lead to omissions or false detections of significant abnormal events due to eye fatigue caused by prolonged staring at the screen. Due to the inefficiency and high cost of traditional manual operations, deep learning-based video abnormal event detection methods have gradually become a hot research topic.

[0061] There are three main types of deep learning-based video anomaly event detection methods: supervised learning-based, weakly supervised learning-based, and unsupervised learning-based video anomaly event detection methods:

[0062] 1. The video anomaly event detection method based on supervised learning is a binary classification task that infers whether each video frame is abnormal from industrial field training data with labeled normal or abnormal data.

[0063] 2. The abnormal event detection method based on weakly supervised learning uses normal and abnormal sample labels at the video level during training. The specific time when the abnormality occurs is unknown. The specific time period when the abnormal event occurs in the factory environment can be found through testing.

[0064] 3. The abnormal event detection method based on unsupervised learning adopts the idea of ​​generative model. By using normal samples to train the model, the reconstruction error between the normal samples generated during the test and the original data is small, while the reconstruction error of the abnormal samples that did not participate in the model training is large. The abnormality of the reconstruction error is detected to alert the production personnel. For example, the related art discloses a "two-stage unsupervised detection method". This method first uses the optical flow branch network and the image branch network to input the image sequence and optical flow sequence of the video respectively for reconstruction, and then uses the reconstructed image sequence and optical flow sequence, as well as the memory network feature information of the optical flow, to input the optical flow feature fusion autoencoder network module, output the predicted video image, and reconstruct the error between the optical flow and the real image based on the error between the predicted image and the real image to detect video abnormalities.

[0065] 1. Disadvantages of supervised learning methods:

[0066] This type of method requires a large amount of video data of abnormal events in industrial sites during training. Due to the ambiguity of the definition of abnormal events, it is difficult to label the training data, resulting in a large overall workload and high cost.

[0067] 2. Disadvantages of weakly supervised learning methods:

[0068] This approach trains at a video-level granularity, relying solely on video-level labels as prior information. This prevents the full utilization of more useful information, limiting the optimization of the network model. Furthermore, the network model's final output only identifies the time period and location of abnormal events involving production personnel, failing to pinpoint the abnormal area at a specific point in time.

[0069] 3. Disadvantages of unsupervised learning methods:

[0070] Convolutional autoencoders are commonly used in this type of method. However, simply stacking encoders and decoders through a multi-layer convolutional neural network can lead to insufficient learning of spatiotemporal features and low-level detail features, affecting the accuracy of model detection. In addition, the ability to generalize normal events becomes increasingly stronger during training, resulting in abnormal events being reconstructed as normal events during testing, causing more missed detections and hindering the detection of abnormal events. For example, although related methods first use dual-path feature extraction, they rely solely on optical flow maps to extract motion features, resulting in insufficient extraction of multimodal information. Secondly, the two-stage training process in this method increases the complexity of the algorithm, and the overall model convergence speed and running time will deteriorate.

[0071] The present disclosure proposes a method for detecting abnormal video events suitable for the industrial Internet, which can detect and alarm abnormal events related to production personnel in a cloud-edge-end collaborative industrial Internet control system. Its characteristics lie in the improvement of the model algorithm level, and its main purpose is to more effectively mine the prior information in the video frame and reduce the occurrence of missed abnormal events. First, the video data is collected at the end layer, and then the data is processed and integrated at the edge layer. The data is then uploaded to the cloud, and the abnormal event detection model and dynamic threshold are used to detect whether an abnormal event has occurred. Finally, the processing results are sent to the industrial gateway intelligent management and control platform for real-time monitoring and alarm.

[0072] In this embodiment, the method for detecting abnormal video events includes:

[0073] Step S10: Obtain a current video frame of the industrial site, and reconstruct a current reconstructed frame corresponding to the current video frame through an anomaly detection model. The anomaly detection model is a model built based on a dual-stream autoencoder network, and the dual-stream autoencoder network is used to extract spatial features and temporal features.

[0074] It can be understood that the execution subject of this embodiment can be a video abnormal event detection device with data processing, network communication and program running functions, such as an industrial Internet control system, etc., or other electronic devices that can achieve the same or similar functions. This embodiment does not limit this.

[0075] It should be understood that this embodiment proposes a new dual-stream autoencoder network for unsupervised learning, which extracts low-level information of each modality in the initial stage of the network, realizes the complementarity of low-level spatial features and temporal features in the encoder stage, does not require labels and is suitable for the detection of low-probability events such as abnormal events, realizes the full learning of spatiotemporal features and low-level detail features, improves the ability to distinguish the location of abnormal events and the specific type of abnormalities, and helps to identify and alarm subtle abnormal behaviors of production personnel in industrial sites.

[0076] It can be understood that obtaining the current video frame of the industrial site may be obtaining the current video data of the industrial site and preprocessing the current video data to obtain the current video frame.

[0077] Step S20: Calculating a current dynamic threshold of the current video frame, where the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame.

[0078] It should be understood that the dynamic threshold means that for each video frame, there is a calculated threshold to determine whether each frame is normal or not. Because this threshold changes continuously with consecutive video frames, the industrial Internet control system has higher sensitivity and anti-interference ability.

[0079] This embodiment completes anomaly detection of continuous video frames by setting a dynamic threshold, reducing the problem of system false detection or missed detection, and increasing the practical value of the model.

[0080] Step S30: detecting abnormal events at the industrial site according to the current dynamic threshold and the current reconstructed frame.

[0081] It can be understood that detecting abnormal events at the industrial site based on the current dynamic threshold and the current reconstructed frame can be that when the reconstruction error of the current reconstructed frame is greater than or equal to the current dynamic threshold, it is judged that an abnormal event at the industrial site is detected in the current video frame; when the reconstruction error of the current reconstructed frame is less than the current dynamic threshold, it is judged that no abnormal event at the industrial site is detected in the current video frame.

[0082] In this embodiment, it is disclosed to obtain the current video frame of the industrial site, and reconstruct the current reconstructed frame corresponding to the current video frame through an anomaly detection model. The anomaly detection model is a model constructed based on a dual-stream autoencoder network. The dual-stream autoencoder network is used to extract spatial features and temporal features, calculate the current dynamic threshold of the current video frame, and the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame. Abnormal events at the industrial site are detected based on the current dynamic threshold and the current reconstructed frame. Since this embodiment uses the anomaly detection model and the dynamic threshold to detect whether an abnormal event has occurred, it can suppress the model's ability to reconstruct abnormal events, encourage the autoencoder to produce a higher reconstruction error for abnormal events, reduce false detection or missed detection problems, and improve the accuracy of video abnormal event detection.

[0083] 3 , which is a flow chart of a second embodiment of a method for detecting abnormal video events according to the present disclosure, the second embodiment of the method for detecting abnormal video events according to the present disclosure is proposed based on the first embodiment shown in FIG. 2 .

[0084] In the second embodiment, before step S10, the method further includes:

[0085] Step S01: Acquire video data of an industrial site, and preprocess the video data to obtain video frames.

[0086] It should be understood that in this embodiment, video data of the industrial site is first obtained, and the video data is preprocessed to obtain video frames, multimodal graphs and pseudo-abnormal video frames are synthesized based on the video frames, and an anomaly detection model is constructed according to the video frames, multimodal graphs and pseudo-abnormal video frames. Since this embodiment effectively utilizes the prior information of the input data by synthesizing pseudo-abnormal frames and multimodal graphs, and at the same time uses the spatiotemporal feature fusion module to reduce the amount of calculation and the interference of irrelevant information, the accuracy of the anomaly detection model is improved, and the real-time detection effect of the system is guaranteed.

[0087] For ease of understanding, reference is made to FIG4 for explanation, but this does not limit the present disclosure. FIG4 is a specific flow chart of a video abnormal event detection method according to an embodiment of the present disclosure. In the figure, the video abnormal event detection method includes the following steps:

[0088] (110 Industrial site video data collection: Industrial site video data collection: Use video acquisition equipment to collect video data of the industrial site that only contains normal industrial operations of production personnel as the first set, and collect video data that contains both normal and abnormal events as the second set. The first set is the training set and the second set is the test set. At the same time, in order to enhance the generalization ability of the model, other types of data are also collected synchronously to expand the first set and the second set, such as different times (day and night) and different environments (different factories or production lines).

[0089] (120 Video Data Preprocessing: The two datasets were preprocessed using the middle layer of the industrial intelligent management and control platform in the 5G new industrial gateway. First, each video was divided into video frames and reshaped to a uniform size. To reduce computational complexity, the video frames were converted to grayscale images to reduce dimensionality. To enhance image contrast and make detailed information clearer, histogram equalization was used to enhance the image. Second, to ensure that all video frames were of the same size, their pixel values ​​were reduced to between 0 and 1. Finally, to eliminate bias in the dataset, the training and test sets were normalized by calculating the mean image of the training dataset. In addition, for special scenes captured at different times due to excessive brightness or darkness, such as sunny backlighting or night scenes captured in factories, adaptive nonlinear color image enhancement technology was used to reduce image feature loss. It mainly uses nonlinear functions to enhance the image. The image neighborhood information is obtained by convolving the Gaussian kernel function with the image. The neighborhood information is used to adaptively enhance the image and restore the color saturation.

[0090] Step S02: synthesizing a multimodal image and a pseudo abnormal video frame based on the video frame.

[0091] It should be understood that in this embodiment, by synthesizing pseudo-abnormal frames and multimodal images, the prior information of the input data is effectively utilized, and at the same time, the spatiotemporal feature fusion module is used to reduce the amount of calculation and the interference of irrelevant information, thereby ensuring the real-time detection effect of the system.

[0092] Furthermore, in order to better capture the long-term motion characteristics of production personnel and reduce the interference of background information, the multimodal graph includes at least one of a dynamic graph, an optical flow graph, and a motion history graph (MHI). The synthesis of the multimodal graph based on the video frame includes: encoding the frame sequence of the video frame based on the parameters of a preset sorting function, and sorting and learning the pixels of the video frame according to the encoding result to obtain a dynamic graph; calculating the changes in the time domain between adjacent video frames in the video frame to obtain an optical flow graph; segmenting the background and the motion foreground based on the motion energy graph, and compressing the video frame into a static image to obtain a motion history graph.

[0093] For ease of understanding, reference is made to FIG4 for explanation, but this does not limit the present disclosure. FIG4 is a specific flow chart of a video abnormal event detection method according to an embodiment of the present disclosure. In the figure, the video abnormal event detection method further includes the following steps:

[0094] (130 Synthesizing Multimodal Graphs: To better capture the long-term motion characteristics of production personnel and reduce the interference of background information, this embodiment proposes using dynamic graphs, optical flow graphs, and motion history graphs as prior motion information for input data. Dynamic graphs are primarily based on the theory of sorting fusion. The parameters of a sorting function are used to encode the video frame sequence, and the pixels of the captured original continuous video frames are directly sorted and learned. In effect, multiple frames in a video are merged into a single image, which represents the appearance and motion information of a video segment. This embodiment uses a weighted sorting pooling algorithm to synthesize dynamic graphs. Optical flow graphs are obtained by calculating the temporal changes in the pixels of two adjacent frames. This embodiment uses the FlowNet deep learning method to synthesize optical flow graphs. The motion history graph, based on the motion energy graph, segments the background and foreground, then compresses a video sequence into a static image. This can be used to represent the displacement changes of the moving object over time. This embodiment uses the inter-frame difference method to extract the motion energy graph.

[0095] Furthermore, in order to simplify the generation of pseudo-abnormal frames and improve the efficiency and ability of the edge computing layer to process data, the synthesis of pseudo-abnormal video frames based on the video frames includes: obtaining skip frame parameters of the video frames; and controlling the video frames to synthesize pseudo-abnormal frames in a fixed interval skipping manner based on the skip frame parameters.

[0096] It can be understood that in this embodiment, during the training phase, pseudo-abnormal frames are synthesized by skipping normal frames at fixed intervals, thereby simplifying the generation of pseudo-abnormal frames and improving the efficiency and ability of the edge computing layer to process data.

[0097] For ease of understanding, reference is made to FIG. 4 for illustration, but this does not limit the present disclosure. FIG. 4 is a specific flowchart of the video anomaly event detection method according to an embodiment of the video anomaly event detection method of the present disclosure. In the figure, the video anomaly event detection method further includes the following steps:

[0098] (140 Synthesize pseudo-anomaly video frames: The idea of the pseudo-anomaly synthesis method comes from the common characteristics when most anomaly events occur, that is, some abnormal targets will show rapid or sudden changes in motion. Therefore, in order to fit more realistic video frames and simply and efficiently generate pseudo-anomaly video frames, a pseudo-anomaly generator is used to simulate abnormal motion from normal data and generate a pseudo-anomaly sequence by skipping frames at a fixed interval. Let every T frames be used as the input of the model, then the selected training video frame sequence is Ω, and its size is T×C×H×W, where T, C, H, and W represent the number of frames in the input sequence, the number of channels, the frame height, and the frame width respectively. Taking the probability p to represent the proportion of the pseudo-anomaly data set used in the entire training set Ω, first randomly select a frame with an index i from the i-th training video segment video i ={x1,x2,L x n} and use this as the starting point to extract T frames to construct a normal training data frame sequence Ω N ={x i ,x i+1 ,L,x i+T-1}={x i+t} 0≤t<T, i+T-1≤n; in addition, introduce a jump frame parameter s to control the generation of a pseudo-anomaly frame sequence Ω P ={x i ,x i+s ,L,x i+(T-1)s}={x i+ts} 0≤t<T, i+(T-1)s≤n, s>1.

[0099] Step S03: Construct an anomaly detection model according to the video frame, the multi-modal graph, and the pseudo-anomaly video frame.

[0100] For ease of understanding, reference is made to FIG. 4 for illustration, but this does not limit the present disclosure. FIG. 4 is a specific flowchart of the video anomaly event detection method according to an embodiment of the video anomaly event detection method of the present disclosure. In the figure, the video anomaly event detection method further includes the following steps:

[0101] (150 Construction of anomaly detection model: The overall network model still follows the idea of ​​unsupervised learning. During training, the continuous video frames of the model only include normal events, namely the first set. By minimizing the error loss function, the model parameters are adjusted so that the reconstructed frames can better fit the input frames. After proper training, when the second set is input, the model can learn the characteristics of normal scenes, thereby reconstructing normal video frames with lower errors, and reconstructing video frames composed of abnormal scenes with higher errors, thereby achieving the effect of differentiation.

[0102] Furthermore, in order to improve the model training effect, the step S03 includes: constructing a dual-stream autoencoder network, the dual-stream autoencoder network includes an appearance network and a motion network, the appearance network is used to train the encoder from the perspective of spatial stream, and the motion network is used to train the encoder from the perspective of temporal stream; inputting the video frame and the pseudo-abnormal video frame into the appearance network to obtain a spatial stream bottleneck layer vector; inputting the multimodal image into the motion network to obtain a temporal stream bottleneck layer vector; performing feature fusion on the temporal stream bottleneck layer vector and the spatial stream bottleneck layer vector to obtain a fusion vector; and training the network based on the fusion vector by back-propagation of a joint reconstruction error loss function to obtain an anomaly detection model.

[0103] Furthermore, the joint reconstruction error loss function is used to maximize the loss of pseudo-abnormal reconstructed frames and pseudo-abnormal original frames and minimize the loss of reconstructed frames and original frames during training, and the motion network is trained through multimodal graph loss constraints.

[0104] For ease of understanding, reference is made to FIG5 for illustration, but this does not limit the present disclosure. FIG5 is a specific flow chart of constructing an anomaly detection model according to an embodiment of the method for detecting abnormal events in a video according to the present disclosure. In the figure, the specific steps of constructing an anomaly detection model are as follows:

[0105] (151 Constructing a two-stream autoencoder network: This embodiment proposes a new two-stream autoencoder network framework. The proposed architecture first trains the encoder part from the perspectives of time stream and space stream respectively. The spatial stream inputs continuous video frames and pseudo-abnormal frames into the appearance network to learn appearance patterns. In order to focus on more low-level contours and edge features, the encoder uses 5 3D convolutional layers to extract spatial features of continuous video frames. The size of the convolution kernel of each convolution layer is 3×3×3, the step size is 2, and the padding edge size is 1. From the first layer to the sixth layer, there are 64, 96, 128, 256, and 256 convolution kernels respectively. Therefore, 64 feature maps of size 32×128×128, 96 of size 16×64×64, 128 of size 8×32×32, 256 of size 4×16×16, and 256 of size 2×8×8 are output respectively. The time stream also uses the same convolutional layer to form a motion network. Since the dynamic image is obtained by taking a convolution kernel every 15 frames, the convolution kernel is used to extract the spatial features of the continuous video frames. The window synthesizes an RGB image, moving the window one frame at a time. Therefore, each synthesized dynamic image is a description of the corresponding motion changes added to each frame of the image. 16 frames of dynamic images can be synthesized continuously in a training batch. The optical flow map uses two consecutive frames of images as input, and the CNN network directly predicts the optical flow to output the corresponding optical flow map. 8 optical flow maps are synthesized continuously in a training batch. The motion history map is similar to the dynamic map. The inter-frame difference method is used to synthesize an MHI map every 7 frames. The image is moved one frame at a time. To match the number of video frames input to the appearance network, 8 frames of motion history map are generated continuously. Next, the learned spatiotemporal features are fused, and a decoder symmetrical to the encoder is used to restore the fused features to the same size as the initial input, generating a reconstructed video frame sequence. The specific network architecture is shown in Figure 6, which is a network architecture diagram of the anomaly detection model of an embodiment of the video anomaly event detection method disclosed in the present disclosure.

[0106] (152 Spatiotemporal feature fusion module: In order to better integrate spatiotemporal information and reduce the amount of computation and interference from irrelevant information, the time stream and spatial stream bottleneck layer vectors after the dual-stream autoencoder are fused. The SENet network is used to integrate the information of the time stream and spatial stream, and different weights are assigned to different feature channels, further improving the detection performance of the algorithm.

[0107] (153 Joint reconstruction error loss function back propagation training network: The two-stream network is used to achieve the complementarity of low-level spatial features and temporal features. In order to jointly train the spatial stream and the temporal stream, a loss function called joint reconstruction error is proposed.

[0108] S1, first set the video frame of the input spatial stream to x, and set the reconstructed frame after training the dual-stream convolutional self-encoder and decoder network to Then set the multimodal graph input sequence The reconstruction sequence is The proposed joint reconstruction error loss function is specifically defined by minimizing the Euclidean loss of the reconstructed frame, the original frame and the multimodal image. First, a special case needs to be considered - pseudo anomaly Ω P The loss function is as follows:

[0109] The error term is expected to maximize the loss of the pseudo-anomaly reconstructed frame and the pseudo-anomaly original frame during training, so that the network learns the characteristic of increasing the loss for abnormal samples, which helps to limit the reconstruction ability of the autoencoder on abnormal inputs. The negative sign is added to better match the overall loss function of the network for training.

[0110] S2, then in the appearance network, in order to learn the normal event Ω N , optimize the network parameters, minimize the loss of the reconstructed frame and the original frame, the specific formula is as follows:

[0111] S3, a multimodal graph is composed of a dynamic graph, an optical flow graph, and a motion history graph. It can capture information in the video from multiple angles and improve the accuracy of anomaly detection. The overall consideration is to use a multimodal graph loss constraint to train the motion network. The specific formula is as follows:

[0112] S4. Finally, the overall joint reconstruction error loss during training is expressed as follows:

[0113] Among them, α and β are weighted coefficients of different loss functions, which are used to measure their relative importance in model training and inference, and α+β=1. When pseudo-abnormal frames are involved in training, L P Because the existence of the negative sign makes the overall loss function smaller and smaller during training, L M Since it only contains normal events, the smaller the loss function, the better. When normal frames participate in training, L N and L M After multiple iterations, the error gradually decreases. In summary, the joint reconstruction error loss function is back-propagated and updates the network gradient in the direction of decreasing loss function. This allows the model to promote the reconstruction of normal events while also encouraging the autoencoder to produce a higher reconstruction error for abnormal events, effectively improving the ability to distinguish between normal and abnormal events in industrial sites.

[0114] Since this embodiment uses a new joint error loss function adapted to the dual-stream autoencoder, it can suppress the cloud model's ability to reconstruct abnormal events and encourage the autoencoder to produce lower and higher reconstruction errors for normal and abnormal events in the industrial site, respectively.

[0115] In this embodiment, video data of the industrial site is first acquired, and the video data is preprocessed to obtain video frames. Multimodal graphs and pseudo-abnormal video frames are synthesized based on the video frames, and an anomaly detection model is constructed according to the video frames, multimodal graphs and pseudo-abnormal video frames. Since this embodiment effectively utilizes the prior information of the input data by synthesizing pseudo-abnormal frames and multimodal graphs, and at the same time uses the spatiotemporal feature fusion module to reduce the amount of calculation and the interference of irrelevant information, the accuracy of the anomaly detection model is improved, and the real-time detection effect of the system is guaranteed.

[0116] 7 , which is a flow chart of a third embodiment of a method for detecting abnormal video events according to the present disclosure, the third embodiment of the method for detecting abnormal video events according to the present disclosure is proposed based on the first embodiment shown in FIG. 2 .

[0117] In the third embodiment, step S20 includes:

[0118] Step S201: Determine the quality index of the current video frame by using the peak signal-to-noise ratio, and normalize the quality index to obtain an abnormality score value of the current video frame.

[0119] It should be understood that in order to improve the accuracy of the current dynamic threshold, in this embodiment, the quality index of the current video frame is first determined by the peak signal-to-noise ratio, and the quality index is normalized to obtain the abnormal score value of the current video frame. Then, the normal value range of the current video frame is determined based on the quartile box plot method, and the initial threshold is determined according to the normal value range. Then, the local statistical properties of the abnormal score value are calculated through a fixed sliding window, and the initial threshold is adjusted through the local statistical properties to obtain the current dynamic threshold of the current video frame.

[0120] Step S202: determining a normal value range of the current video frame based on the quartile box plot method, and determining an initial threshold value according to the normal value range.

[0121] Step S203: calculating local statistical properties of the anomaly score value through a fixed sliding window, and adjusting the initial threshold value according to the local statistical properties to obtain a current dynamic threshold value of the current video frame.

[0122] For ease of understanding, reference is made to FIG4 for explanation, but this does not limit the present disclosure. FIG4 is a specific flow chart of a video abnormal event detection method according to an embodiment of the present disclosure. In the figure, the video abnormal event detection method further includes the following steps:

[0123] (160 Use dynamic threshold to detect anomalies: Dynamic threshold means that for each video frame, there is a threshold calculated to judge whether each frame is normal or not. Therefore, this threshold changes continuously with the continuous video frames, making the monitoring system have higher sensitivity and anti-interference ability.

[0124] S1, first use the Peak Signal-to-Noise Ratio (PSNR) formula to obtain the value as the generation quality indicator of the reconstructed frame in the evaluation test set, and then normalize the value to the maximum and minimum to obtain the normal score S of each frame r (x t ) and the anomaly score S with a complementary relationship a (x t ).

[0125] S2, then determine the initial threshold, and use the box-plot method (quartile box-line plot) to determine the range of normal values ​​for the abnormal score values ​​of the normal data set [B l ,B u ], the specific formula is as follows: B l =Q1-1.5IQR,B u =Q3+1.5IQR. Q1 refers to the 25% of the data in all video frames with the abnormal score value from small to large, Q3 refers to the 75% of the data, and IQR refers to the difference between Q3 and Q1. u As an initial threshold, it is used to roughly identify abnormal events.

[0126] S3 then uses the mean and standard deviation to find an appropriate dynamic threshold. To eliminate noise in the anomaly score, we use a fixed sliding window to calculate local statistical properties of the anomaly score and use them to modify the threshold at test time. Given the current t-th frame, the window is set to [t-(T+1), t-1], and the mean and standard deviation within the sliding window are formulated as follows:

[0127] S4, the final dynamic threshold is defined as: Where a is the adjustment parameter.

[0128] (170) Output results to the industrial gateway intelligent management and control platform: The results of the judgment using dynamic thresholds are fed back to the industrial gateway intelligent management and control platform in real time. The platform supports a visual user interface and simple operation procedures to alarm abnormal events in the industrial field that may endanger the personal safety of production personnel, and provides security solutions based on 5G new industrial gateway related services.

[0129] In this embodiment, the quality index of the current video frame is first determined by the peak signal-to-noise ratio, and the quality index is normalized to obtain the abnormality score value of the current video frame. Then, the normal value range of the current video frame is determined based on the quartile box plot method, and the initial threshold is determined according to the normal value range. Then, the local statistical properties of the abnormality score value are calculated through a fixed sliding window, and the initial threshold is adjusted according to the local statistical properties to obtain the current dynamic threshold of the current video frame, thereby improving the accuracy of the current dynamic threshold.

[0130] In addition, referring to FIG8 , the embodiment of the present disclosure further provides a video abnormal event detection device, the video abnormal event detection device comprising:

[0131] An acquisition module 10 is configured to acquire a current video frame of the industrial site and reconstruct a current reconstructed frame corresponding to the current video frame using an anomaly detection model, wherein the anomaly detection model is a model constructed based on a dual-stream autoencoder network for extracting spatial and temporal features;

[0132] A calculation module 20, configured to calculate a current dynamic threshold of the current video frame, wherein the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame;

[0133] The detection module 30 is configured to detect abnormal events at the industrial site according to the current dynamic threshold and the current reconstructed frame.

[0134] In this embodiment, it is disclosed to obtain the current video frame of the industrial site, and reconstruct the current reconstructed frame corresponding to the current video frame through an anomaly detection model. The anomaly detection model is a model constructed based on a dual-stream autoencoder network. The dual-stream autoencoder network is used to extract spatial features and temporal features, calculate the current dynamic threshold of the current video frame, and the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame. Abnormal events at the industrial site are detected based on the current dynamic threshold and the current reconstructed frame. Since this embodiment uses the anomaly detection model and the dynamic threshold to detect whether an abnormal event has occurred, it can suppress the model's ability to reconstruct abnormal events, encourage the autoencoder to produce a higher reconstruction error for abnormal events, reduce false detection or missed detection problems, and improve the accuracy of video abnormal event detection.

[0135] In one embodiment, the calculation module 20 is further used to determine the quality index of the current video frame through the peak signal-to-noise ratio, and normalize the quality index to obtain the abnormality score value of the current video frame; determine the normal value range of the current video frame based on the quartile box plot method, and determine the initial threshold according to the normal value range; calculate the local statistical properties of the abnormality score value through a fixed sliding window, and adjust the initial threshold through the local statistical properties to obtain the current dynamic threshold of the current video frame.

[0136] In one embodiment, the apparatus for detecting abnormal video events further includes:

[0137] The training module is used to obtain video data from an industrial site and preprocess the video data to obtain video frames; synthesize multimodal graphs and pseudo-abnormal video frames based on the video frames; and construct an anomaly detection model based on the video frames, the multimodal graphs, and the pseudo-abnormal video frames.

[0138] In one embodiment, the training module is also used to construct a dual-stream autoencoder network, which includes an appearance network and a motion network. The appearance network is used to train the encoder from the perspective of spatial stream, and the motion network is used to train the encoder from the perspective of temporal stream; the video frame and the pseudo-abnormal video frame are input into the appearance network to obtain a spatial stream bottleneck layer vector; the multimodal image is input into the motion network to obtain a temporal stream bottleneck layer vector; the temporal stream bottleneck layer vector and the spatial stream bottleneck layer vector are feature fused to obtain a fused vector; based on the fused vector, the network is trained by back-propagation of a joint reconstruction error loss function to obtain an anomaly detection model.

[0139] In one embodiment, the joint reconstruction error loss function is used to maximize the loss of pseudo-abnormal reconstructed frames and pseudo-abnormal original frames and minimize the loss of reconstructed frames and original frames during training, and the motion network is trained through multimodal graph loss constraints.

[0140] In one embodiment, the multimodal graph includes at least one of a dynamic graph, an optical flow graph, and a motion history graph. The training module is further used to encode the frame sequence of the video frame based on the parameters of a preset sorting function, and sort and learn the pixels of the video frame according to the encoding result to obtain a dynamic graph; calculate the changes in the time domain between adjacent video frames in the video frame to obtain an optical flow graph; segment the background and the motion foreground based on the motion energy graph, and compress the video frame into a static image to obtain a motion history graph.

[0141] In one embodiment, the training module is further configured to obtain skip frame parameters of the video frames; and control the video frames to skip frames at fixed intervals to synthesize pseudo abnormal frames based on the skip frame parameters.

[0142] Other embodiments or specific implementations of the video abnormal event detection device disclosed in the present invention can refer to the above-mentioned method embodiments and will not be repeated here.

[0143] In addition, an embodiment of the present disclosure further provides a storage medium, on which a video abnormal event detection program is stored. When the video abnormal event detection program is executed by a processor, the video abnormal event detection method described above is implemented.

[0144] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0145] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.

[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory image (ROM) / random access memory (RAM), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present disclosure.

[0147] The above are only preferred embodiments of the present disclosure and are not intended to limit the patent scope of the present disclosure. Any equivalent structure or equivalent process transformation made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.

Claims

1. A method for detecting abnormal video events, comprising: Acquire a current video frame of the industrial site, and reconstruct a current reconstructed frame corresponding to the current video frame through an anomaly detection model, wherein the anomaly detection model is a model built based on a two-stream autoencoder network, and the two-stream autoencoder network is used to extract spatial features and temporal features; Calculating a current dynamic threshold of the current video frame, wherein the current dynamic threshold is used to perform anomaly detection on the current reconstructed frame; An abnormal event at the industrial site is detected according to the current dynamic threshold and the current reconstructed frame.

2. The method for detecting abnormal video events according to claim 1, wherein: The calculating the current dynamic threshold of the current video frame includes: Determine a quality index of the current video frame by a peak signal-to-noise ratio, and perform normalization processing on the quality index to obtain an abnormality score value of the current video frame; Determining a normal value range of the current video frame based on the quartile box plot method, and determining an initial threshold value according to the normal value range; The local statistical property of the abnormal score value is calculated through a fixed sliding window, and the initial threshold is adjusted according to the local statistical property to obtain a current dynamic threshold of the current video frame.

3. The method for detecting abnormal video events according to claim 1, wherein: Before acquiring the video frame of the industrial site and reconstructing the reconstructed frame corresponding to the video frame through the anomaly detection model, the method further includes: Acquire video data of an industrial site, and preprocess the video data to obtain video frames; synthesizing a multimodal graph and a pseudo abnormal video frame based on the video frame; An anomaly detection model is constructed according to the video frame, the multimodal graph, and the pseudo-abnormal video frame.

4. The method for detecting abnormal video events according to claim 3, wherein: The constructing an anomaly detection model according to the video frame, the multimodal graph, and the pseudo-anomaly video frame includes: Constructing a two-stream autoencoder network, the two-stream autoencoder network comprising an appearance network and a motion network, the appearance network is used to train the encoder from a spatial stream perspective, and the motion network is used to train the encoder from a temporal stream perspective; Inputting the video frame and the pseudo abnormal video frame into the appearance network to obtain a spatial flow bottleneck layer vector; Inputting the multimodal graph into the motion network to obtain a time stream bottleneck layer vector; Performing feature fusion on the time stream bottleneck layer vector and the space stream bottleneck layer vector to obtain a fusion vector; Based on the fusion vector, a network is trained by back propagation of a joint reconstruction error loss function to obtain an anomaly detection model.

5. The method for detecting abnormal video events according to claim 4, wherein: The joint reconstruction error loss function is used to maximize the loss of pseudo-abnormal reconstructed frames and pseudo-abnormal original frames and minimize the loss of reconstructed frames and original frames during training, and train the motion network through multi-modal graph loss constraints.

6. The method for detecting abnormal video events according to claim 3, wherein: The multimodal graph includes at least one of a dynamic graph, an optical flow graph, and a motion history graph, and synthesizing the multimodal graph based on the video frame includes: Encoding the frame sequence of the video frame based on the parameters of the preset sorting function, and sorting and learning the pixels of the video frame according to the encoding result to obtain a dynamic image; Calculating the changes of adjacent video frames in the video frame in the time domain to obtain an optical flow map; The background and the motion foreground are segmented based on the motion energy map, and the video frames are compressed into static images to obtain a motion history map.

7. The method for detecting abnormal video events according to claim 3, wherein: The synthesizing a pseudo abnormal video frame based on the video frame comprises: Obtaining a skip frame parameter of the video frame; The video frames are controlled to synthesize pseudo abnormal frames in a fixed interval frame skipping manner based on the skip frame parameters.

8. A video abnormal event detection device, comprising: An acquisition module is used to acquire a current video frame of the industrial site, and reconstruct a current reconstructed frame corresponding to the current video frame through an anomaly detection model, wherein the anomaly detection model is a model constructed based on a two-stream autoencoder network, and the two-stream autoencoder network is used to extract spatial features and temporal features; A calculation module, used for calculating a current dynamic threshold of the current video frame, wherein the current dynamic threshold is used for performing anomaly detection on the current reconstructed frame; A detection module is used to detect abnormal events in the industrial site according to the current dynamic threshold and the current reconstructed frame.

9. A video abnormal event detection device, comprising: A memory, a processor, and a video abnormal event detection program stored in the memory and executable on the processor, wherein the video abnormal event detection program, when executed by the processor, implements the video abnormal event detection method according to any one of claims 1 to 7.

10. A storage medium having a video abnormal event detection program stored thereon, wherein the video abnormal event detection program, when executed by a processor, implements the video abnormal event detection method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Communication machine room anomaly detection method and device and computing equipment

    CN112199977A

  • Video abnormal event detection method based on double-flow space-time auto-encoder

    CN115830541A

  • Video anomaly detection method based on multi-layer memory enhancement and secondary prediction

    CN116665098A

  • Video abnormal event detection method and device, equipment and storage medium

    CN117671560A

  • Method and System for Zero-Shot Cross Domain Video Anomaly Detection

    US20230281986A1

Cited By

  • Real-time quality detection method and system based on 5G edge calculation

    CN120355707A

  • A real-time quality detection method and system based on 5G edge computing

    CN120355707B

  • Big data anomaly detection method and system based on cloud computing

    CN120744706A

  • Personnel safety detection method and device based on dynamic resource allocation

    CN120876201A

  • Embedded behavior anomaly detection system based on time sequence image analysis

    CN120977017A