AI-based factory abnormal behavior identification monitoring method and system

By adopting an abnormal behavior recognition method based on AI in the industrial plant monitoring system and using multimodal feature fusion to identify abnormal behaviors, the shortcomings of the existing system in terms of accuracy and real-time performance are solved, and accurate identification and real-time response of complex abnormal behaviors are achieved.

CN120088737AActive Publication Date: 2025-06-03GUIZHOU YIQI TONGWU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510565723.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing industrial plant monitoring system has shortcomings in accuracy and real-time performance, making it difficult to effectively identify complex abnormal behaviors, especially in high-risk operation scenarios.

Method used

Using an AI-based factory abnormal behavior recognition monitoring method, a multi-frame behavior feature set is extracted by obtaining real-time video stream data, including bone key points, environmental object contours and action trajectory features, and cross-modal correlation and fusion is performed through the multi-modal feature encoding layer to generate target behavior feature vectors, and abnormal behavior recognition and early warning are performed.

Benefits of technology

It realizes accurate identification and real-time response to complex abnormal behaviors, significantly improves the accuracy and response speed of judgment of violations, and reduces the incidence of safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088737A_ABST
    Figure CN120088737A_ABST
Patent Text Reader

Abstract

The invention provides an AI-based factory abnormal behavior identification monitoring method and system. The method comprises the steps of obtaining continuous video stream data collected by a plurality of monitoring devices in a target factory in real time; performing time dimension segmentation and space dimension preprocessing on the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence, and inputting the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, performing cross-modal association fusion on the skeleton key point features, the environment associated object contour features and the action trajectory features to generate a target behavior feature vector, performing similarity matching with a preset abnormal behavior feature library, determining an abnormal behavior category, and triggering a corresponding early warning signal according to the behavior type when it is determined that the target behavior feature vector belongs to the abnormal behavior category. And real-time alarm display is carried out. According to the invention, the problem that the existing monitoring system is insufficient in accuracy and real-time performance can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and artificial intelligence, and particularly to an AI-based method and system for identifying and monitoring abnormal behaviors in a factory area. Background Art

[0002] In the field of industrial factory area safety production management, traditional monitoring systems mostly rely on manual inspections or simple rule-based judgments of abnormal behaviors based on single sensors (such as infrared detection, area intrusion alarms), which have significant defects. Single sensors or visual features are vulnerable to light changes and occlusion interference, resulting in a high false alarm rate. Secondly, static rule libraries cannot adapt to dynamic factory area environments (such as production line layout adjustments, introduction of new equipment), and manual threshold updates are required frequently, with high operation and maintenance costs. Moreover, there is a lack of in-depth analysis of behavior continuity and multi-modal associations, making it difficult to identify complex risks that require cross-time and cross-space linkages (such as the temporal association between a smoking action and contact with flammable substances). Especially in high-risk operation scenarios, the existing systems have a significant missed detection rate for complex behaviors that require a combination of object recognition and trajectory analysis, such as illegal power connection and collaborative operations in dangerous areas, and cannot meet the requirements of intelligent manufacturing for proactive safety prevention and control. Therefore, there is an urgent need for a method that can solve the problems of insufficient accuracy and real-time performance of existing monitoring systems. Summary of the Invention

[0003] In view of this, embodiments of the present invention at least provide an AI-based method and system for identifying and monitoring abnormal behaviors in a factory area. The technical solutions of the embodiments of the present invention are implemented as follows: On the one hand, an embodiment of the present invention provides an AI-based method for identifying and monitoring abnormal behaviors in a factory area, the method comprising: Obtaining continuous video stream data collected in real time by a plurality of monitoring devices in a target factory area, the continuous video stream data comprising a dynamic behavior sequence of at least one target object; Performing time dimension segmentation and space dimension preprocessing on the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence, the multi-frame behavior feature set comprising skeletal key point features, environment-associated object contour features, and action trajectory features of the target object in each video frame; Inputting the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, and performing cross-modal association fusion on the skeletal key point features, environment-associated object contour features, and action trajectory features through a multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector; Based on similarity matching between the target behavior feature vector and a preset abnormal behavior feature library, determining whether the behavior type of the target object belongs to a predefined abnormal behavior category; If it is determined that it belongs to the abnormal behavior category, the corresponding warning signal is triggered according to the behavior type, and the warning signal and the associated video frames are sent to the target terminal device for real-time alarm display.

[0004] Further, the time-dimensional segmentation and spatial-dimensional preprocessing of the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence includes: The continuous video stream data is divided into multiple video segments at a preset time interval, and each video segment contains a fixed number of consecutive video frames; For each video frame, a pre-trained skeletal key point detection network is called to extract the skeletal key point features of the target object, and the skeletal key point features include the three-dimensional coordinates of each joint point and the motion speed parameters between adjacent joint points; Synchronously call an object contour segmentation network to perform contour recognition on the objects in the video frame that have spatial interaction with the target object, and extract the environmental associated object contour features, where the environmental associated object contour features include object shape descriptors and relative distance parameters from the target object; Based on the temporal relationship of the video segments, track the position changes of the target object in adjacent video frames to generate the action trajectory features, where the action trajectory features include a displacement direction sequence and an acceleration change curve; Combine the skeletal key point features, environmental associated object contour features, and action trajectory features corresponding to the same video frame to form the multi-frame behavior feature set.

[0005] Further, before calling the pre-trained skeletal key point detection network to extract the skeletal key point features of the target object, it also includes the step of training the skeletal key point detection network: Obtain a historical monitoring video data set in the factory area scene, and perform manual annotation on each video frame in the historical monitoring video data set. The annotation content includes the joint point positions and joint point motion labels of the target object; Construct an initial key point detection network, where the initial key point detection network includes a feature pyramid module and a three-dimensional coordinate regression module; Input the video frames in the historical monitoring video data set into the initial key point detection network, extract multi-scale spatial features through the feature pyramid module, and predict the three-dimensional coordinates and motion speed parameters of each joint point through the three-dimensional coordinate regression module; Calculate the Euclidean distance loss value between the predicted coordinates and the annotated coordinates, and combine the difference loss value between the motion speed parameters and the joint point motion labels to generate a total training loss value; Backpropagation optimization is performed on the initial keypoint detection network based on the total training loss value until the total training loss value converges to a preset threshold, and the pre-trained skeletal keypoint detection network is obtained.

[0006] Further, inputting the multi-frame behavior feature set into the pre-trained abnormal behavior recognition model, and cross-modal correlation fusion of the skeletal keypoint features, environment-related object contour features, and action trajectory features is performed through the multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector, including: Input the skeletal keypoint features into the first feature encoding sub-network in the multi-modal feature encoding layer to generate a first encoded feature vector, and the first feature encoding sub-network adopts a temporal convolutional structure to capture the joint motion pattern; Input the environment-related object contour features into the second feature encoding sub-network in the multi-modal feature encoding layer to generate a second encoded feature vector, and the second feature encoding sub-network adopts an attention mechanism to focus on the object contour with the highest interaction frequency with the target object; Input the action trajectory features into the third feature encoding sub-network in the multi-modal feature encoding layer to generate a third encoded feature vector, and the third feature encoding sub-network adopts a recurrent neural network structure to analyze the continuity of the displacement direction sequence; Concatenate the first encoded feature vector, the second encoded feature vector, and the third encoded feature vector, and perform dimensionality reduction fusion through a fully connected layer to obtain the target behavior feature vector.

[0007] Further, the training process of the abnormal behavior recognition model includes: Collect historical abnormal behavior video samples and normal behavior video samples in the factory area, extract the multi-frame behavior feature set for each video sample, and label the corresponding abnormal behavior category label; Construct an initial abnormal behavior recognition model, and the initial abnormal behavior recognition model includes the multi-modal feature encoding layer and the classifier layer; Input the multi-frame behavior feature set into the initial abnormal behavior recognition model, output the target behavior feature vector through the multi-modal feature encoding layer, and calculate the matching probability between the target behavior feature vector and the feature templates corresponding to each abnormal behavior category through the classifier layer; Based on the cross-entropy loss value between the matching probability and the abnormal behavior category label, update the parameters of the initial abnormal behavior recognition model; Introduce an adversarial training strategy, add noise perturbation data to the multi-frame behavior feature set, calculate the classification consistency loss value between the perturbed feature set and the original feature set, and perform joint optimization in combination with the cross-entropy loss value until the model converges.

[0008] Further, the method of determining whether the behavior type of the target object belongs to a predefined abnormal behavior category by performing similarity matching between the target behavior feature vector and a predefined abnormal behavior feature library includes: Read all the benchmark feature vectors of the predefined abnormal behavior categories from the abnormal behavior feature library, where the benchmark feature vectors are obtained by centralizing the features extracted from historical abnormal behavior data through a clustering algorithm; Calculate the cosine similarity between the target behavior feature vector and each benchmark feature vector, and select the abnormal behavior categories corresponding to the top K benchmark feature vectors with the highest similarity as candidate categories; Perform time - dimension alignment verification on the benchmark feature vectors of the candidate categories, and determine whether the change trend of the target behavior feature vector within a continuous time window is consistent with the behavior pattern of the candidate categories; If the verification passes, determine the abnormal behavior category with the highest similarity in the candidate categories as the behavior type of the target object; otherwise, mark the behavior type as an unknown abnormal category and trigger an artificial review process.

[0009] Further, the method for constructing the abnormal behavior feature library includes: Collect video cases of abnormal behaviors that have occurred within the target factory area, where the abnormal behavior cases include illegal smoking, illegal use of electronic devices, and intrusion into dangerous areas; Segment each video case of abnormal behavior, extract multi - frame behavior feature sets within each time period, and label the corresponding abnormal behavior categories and severity levels; Use a hierarchical clustering algorithm to group the multi - frame behavior feature sets of the same abnormal behavior category, and calculate the feature mean of each group as the benchmark feature vector; Assign a priority weight to each benchmark feature vector according to the severity level, and store the benchmark feature vector and its priority weight in the abnormal behavior feature library.

[0010] Further, if it is determined to belong to the abnormal behavior category, triggering a corresponding warning signal according to the behavior type includes: When the behavior type is the first - type abnormal behavior, generate a first - level warning signal, control the start of the audible and visual alarm in the target factory area, and at the same time send an emergency warning message containing behavior screenshots and location information to the terminal device of the safety management personnel; When the behavior type is the second - type abnormal behavior, generate a second - level warning signal and send the behavior video clip and recommended disposal measures to the terminal device; When the behavior type is an unknown exception category, generate a third-level warning signal and upload the associated video stream to the cloud analysis platform; Split the video stream data corresponding to the unknown exception category into multiple annotation tasks and allocate them to multiple expert terminals for independent annotation; Receive the annotation results returned by each expert terminal, where the annotation results include behavior type suggestions and feature correction parameters; Use the majority voting algorithm to perform consistency verification on the annotation results. If the behavior types suggested by more than a preset proportion of experts are the same, add the behavior type to the new category in the abnormal behavior feature library; If the annotation results do not reach consistency, initiate a secondary annotation process until an annotation result that meets the consistency condition is generated.

[0011] Furthermore, the method further includes the step of performing incremental learning on the abnormal behavior recognition model: Regularly collect the warning processing records and manual review results feedback by the target terminal device, and filter out the target case data where the model recognition is incorrect or the confidence level is lower than the threshold; Re-extract the features of the video stream in the target case data to generate an incremental training data set, and add the corrected behavior type label to each data sample; Freeze the parameters of the multi-modal feature encoding layer in the abnormal behavior recognition model, and only fine-tune the weights of the classifier layer; During the fine-tuning training process, use the elastic weight consolidation algorithm to retain the important parameters of the original classifier layer.

[0012] On the other hand, the present invention provides a computer system, including a memory and a processor, where the memory stores a computer program that can run on the processor, and the processor implements the steps in the above method when executing the program.

[0013] The AI-based abnormal behavior recognition and monitoring method provided by the present invention obtains real-time video stream data of the target factory area and divides it into multiple frames of behavior feature sets. Through the multi-modal feature fusion of skeletal key points, environmental object contours, and action trajectories, a target behavior feature vector that can comprehensively represent the behavior pattern is generated. Then, through the matching and warning mechanism of the abnormal behavior feature library, accurate and efficient abnormal behavior recognition and response are achieved. By capturing the details of human postures through skeletal key point features, positioning dangerous item interaction scenarios through environmental associated object contour features, and analyzing displacement rules through action trajectory features, the cross-modal correlation fusion of the three can break through the limitations of single visual features and construct a complete feature portrait of abnormal behaviors from three dimensions: spatial interaction, object association, and motion pattern, significantly improving the discrimination accuracy of complex abnormal behaviors such as illegal smoking and dangerous operations. In addition, by segmenting in the time dimension to retain the continuity of behaviors and preprocessing in the spatial dimension to extract key features with high information density, both redundant data interference is avoided and the dynamic evolution law of the behavior sequence is ensured to be completely captured, enabling the abnormal behavior recognition model to quickly lock the risk time period and spatial area and improving the real-time response speed to sudden abnormal behaviors. Further, based on the feature matching results, differential warning signals are dynamically triggered, and the feature library and model parameters are continuously optimized in combination with the feedback data of the terminal device, forming a closed-loop control mechanism of "recognition-warning-verification-iteration", which not only reduces the interference of false alarms on the production process but also can intervene in advance at the initial stage of illegal behaviors through the prediction mechanism, greatly reducing the incidence of factory area safety accidents. Skeletal key point features can accurately identify illegal postures (such as holding dangerous items), environmental object contour features can associate specific equipment illegal operations (such as unauthorized power connection equipment identification), and action trajectory features can detect abnormal movement paths (such as intrusion into dangerous areas). The synergistic effect of the three ensures stable recognition of various risk behaviors in complex scenarios such as light changes and occlusion interference. Through the dynamic update strategy of the benchmark feature vector of the abnormal behavior feature library, it can adapt to changes in different production line layouts and equipment configurations, avoid the lag defects of traditional static rule libraries, ensure that the recognition model continuously adapts to the evolution of the factory area environment, and form a long-term available intelligent security system.

[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solution of the present invention.

[0016] Figure 1 It is a schematic implementation flowchart of an AI-based abnormal behavior recognition and monitoring method provided by an embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of the composition structure of an abnormal behavior recognition and monitoring device provided by an embodiment of the present invention.

[0018] Figure 3 It is a schematic diagram of the hardware entity of a computer system provided by an embodiment of the present invention. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein are only for the purpose of describing the present invention and are not intended to limit the present invention.

[0021] An embodiment of the present invention provides an AI-based abnormal behavior recognition and monitoring method in a factory area, and this method can be executed by a processor of a computer system. Among them, the computer system may refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, mobile devices (such as mobile phones, personal digital assistants, dedicated messaging devices), etc.

[0022] Figure 1 It is a schematic diagram of the implementation process of an AI-based abnormal behavior recognition and monitoring method in a factory area provided by an embodiment of the present invention. As Figure 1 shown, this method includes the following steps: Step S110: Obtain continuous video stream data collected in real time by multiple monitoring devices in the target factory area, and the continuous video stream data includes a dynamic behavior sequence of at least one target object.

[0023] The target factory area is, for example, the scope of an industrial factory area where monitoring devices are deployed and behavior monitoring is required. Its physical boundary is determined by a fence, an access control system, or a geographical coordinate range, such as the general assembly workshop area or the fermentation area of a liquor production factory. The monitoring devices cover hardware devices with video acquisition functions, including but not limited to fixed cameras, rotatable pan-tilt cameras, infrared thermal imagers, and three-dimensional vision devices with depth sensors. These devices are installed at key positions such as factory entrances and exits, production line workstations, high-risk operation areas, and logistics channels according to preset layout rules to cover the full monitoring perspective of the target factory area. The continuous video stream data is, for example, an uncompressed or compressed-encoded video signal sequence continuously output by the monitoring devices at a constant frame rate. Its temporal continuity is manifested as a fixed and uninterrupted time interval between adjacent video frames. For example, it is transmitted in real time using the H.264 encoding format at 25 frames per second. The target object is the person whose behavior characteristics need to be identified in the video stream data. The dynamic behavior sequence is, for example, the process of spatial position change and action state evolution of the target object within a continuous time window. For example, the action sequence of a worker switching from a walking state to climbing a shelf. In a specific implementation, the security system of the target factory area real-time aggregates the video stream data of each monitoring device through a wired or wireless network, and preliminarily detects and tracks the target object in the video stream to form dynamic behavior sequence metadata including timestamps, device numbers, and target IDs, providing a structured input for subsequent processing.

[0024] Step S120: Perform temporal dimension segmentation and spatial dimension preprocessing on the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence. The multi-frame behavior feature set includes the skeletal key point features, environmental associated object contour features, and action trajectory features of the target object in each video frame.

[0025] The time dimension segmentation is, for example, dividing a continuous video stream into a set of discrete video frames at fixed time intervals. For example, single-frame images are extracted at intervals of every 0.04 seconds (corresponding to 25 frames per second), or adaptive segmentation is performed between the start and end points of key actions in dynamic behaviors to ensure that each video segment completely records a single behavior cycle of the target object. The spatial dimension preprocessing involves pixel-level processing and feature enhancement operations on single-frame images, including filtering algorithms for eliminating motion blur, histogram equalization methods for adjusting uneven illumination, and the application of noise reduction convolutional kernels for reducing noise, so as to improve the robustness of subsequent feature extraction. The multi-frame behavior feature set is the result of aggregating spatio-temporal features extracted from the preprocessed video frame set. Among them, the skeletal key point features are extracted through deep learning-based human pose estimation algorithms (such as OpenPose or AlphaPose), and are specifically represented as a set of two-dimensional or three-dimensional coordinates of human joint points. For example, the position data of the head, shoulders, elbows, and knee joints of an operator in the image coordinate system; the environmental associated object contour features identify the object contours in the video frames that have an interaction relationship with the target object through edge detection algorithms (such as the Canny operator or deep convolutional network). For example, the outer boundary of the tool held by a worker, the contact area between the forklift fork and the pallet; the action trajectory features obtain the displacement vector sequence of the target object between consecutive frames through multi-object tracking algorithms (such as SORT or DeepSORT). For example, the motion path coordinate sequence of a person moving from point A to point B. During the implementation process, the skeletal detection, contour extraction, and trajectory tracking algorithms can be executed in parallel for each frame of video data to generate a time-aligned multi-modal feature set, providing a standardized input for subsequent model processing.

[0026] Step S130: Input the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, and perform cross-modal correlation fusion on the skeletal key point features, environmental associated object contour features, and action trajectory features through the multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector.

[0027] The pre-trained abnormal behavior recognition model is, for example, a deep neural network trained based on a large-scale annotated data set (such as a video library containing factory personnel smoking, making phone calls, crossing boundaries, illegal operations, etc.), and its architecture includes feature encoding, cross-modal fusion, and classification output modules. The multimodal feature encoding layer is the core component in the model responsible for processing heterogeneous features. It uses a multi-head attention mechanism or a graph convolutional network to achieve semantic associations between different modal features, such as mapping skeletal key point features to node attributes in the graph structure, environmental associated object contour features as edge connection weights, and motion trajectory features for dynamically adjusting the spatiotemporal dependencies of the graph. Cross-modal association fusion is, for example, a process of nonlinear transformation and weight assignment of different modal features in a unified feature space, such as dynamically adjusting the contribution of skeletal key point features and motion trajectory features through a gating mechanism, or using cross-attention to calculate the potential impact of environmental associated object contour features on skeletal motion patterns. The target behavior feature vector is a fused high-dimensional dense vector, whose dimension is related to the model design (such as 512 or 1024 dimensions). Each element in the vector encodes the abstract semantics of the target object behavior in the multimodal feature joint space. For example, the activation value of a specific dimension in the vector reflects whether the range of motion of the person exceeds the safety threshold, or whether the spatial relationship between the object contour and the skeletal position meets the standard operation requirements. During implementation, the system inputs the multi-frame behavior feature set into the model in chronological order. The multimodal feature encoding layer performs feature fusion frame by frame and aggregates the temporal context information, and finally outputs the target behavior feature vector representing the overall behavior pattern.

[0028] Step S140: performing similarity matching between the target behavior feature vector and a preset abnormal behavior feature library to determine whether the behavior type of the target object belongs to a predefined abnormal behavior category.

[0029] The preset abnormal behavior feature library is a database that stores the prototype vectors of known abnormal behavior categories. Each category corresponds to a predefined factory safety hazard behavior (such as entering a high-risk area without a safety helmet, crossing a conveyor belt illegally, smoking in a non-smoking area, etc.), and its prototype vector is obtained through clustering algorithms or deep feature statistics of classification models. Similarity matching, for example, calculates the distance measurement value (such as cosine similarity, Euclidean distance or Mahalanobis distance) between the target behavior feature vector and each prototype vector in the library, and selects the highest similarity value as the classification basis. The process of determining the behavior type includes two stages: threshold comparison and category determination: if the highest similarity exceeds the preset threshold (such as 0.85), the target object behavior is determined to belong to the corresponding abnormal category; otherwise, it is classified as normal behavior. For example, when the cosine similarity between the target behavior feature vector and the prototype vector of the "climbing high-risk equipment" category reaches 0.92, it is determined that the current behavior has a safety risk. In implementation, the abnormal behavior feature library needs to be updated regularly to include new abnormal patterns. The update method includes incremental learning of new category prototypes or retraining the model to generate an updated feature space.

[0030] Step S150: If it is determined to belong to the abnormal behavior category, trigger the corresponding warning signal according to the behavior type, and send the warning signal and the associated video frames to the target terminal device for real-time alarm display.

[0031] The warning signal is a standardized alarm instruction bound to the abnormal behavior category, including the alarm level (e.g., a first-level alarm represents immediate intervention, and a second-level alarm requires manual review), the alarm location (monitoring device number and plant coordinates), and the recommended disposal measures (e.g., notify security personnel or cut off the power supply of the device). The associated video frames are, for example, video clips containing key time points before and after the occurrence of the abnormal behavior (such as consecutive frames from 5 seconds before the start of the abnormal behavior to 3 seconds after the end), which are used to assist in manual review and event backtracking. The target terminal device is, for example, the monitoring large screen in the central control room of the plant, the mobile patrol terminal held by security personnel, or the smartphone application of the associated responsible person. The real-time alarm display needs to follow a low-latency transmission protocol (such as WebRTC or RTMP) to ensure that the alarm information reaches the terminal within milliseconds. In specific implementation, after triggering the warning signal, alarm log recording, associated video frame storage, and multi-terminal broadcast push are automatically executed, and at the same time, linkage control is started according to preset rules (such as turning on the sound and light alarm or locking the access control of the relevant area). For example, when "a person breaks into the dangerous machinery operation area" is recognized, the system immediately sends an alarm message containing the location map and real-time video stream to the nearest patrol terminal in this area and activates the on-site warning light to remind people to evacuate.

[0032] Through the above steps S110 - S150, the embodiment of the present invention obtains the real - time video stream data of the target factory area and divides it into multiple frames of behavior feature sets. By integrating the multi - modal features of skeletal key points, environmental object contours, and action trajectories, a target behavior feature vector that can comprehensively represent the behavior pattern is generated. Then, through the matching and warning mechanism of the abnormal behavior feature library, accurate and efficient abnormal behavior recognition and response are achieved. The details of the human body posture are captured through the skeletal key - point features, the interaction scenarios of dangerous items are located through the environmental - related object contour features, and the displacement rules are analyzed through the action - trajectory features. The cross - modal correlation and fusion of the three can break through the limitations of single - vision features, construct a complete feature portrait of abnormal behaviors from three dimensions: spatial interaction, object association, and motion pattern, and significantly improve the discrimination accuracy of complex abnormal behaviors such as illegal smoking and dangerous operations. In addition, by segmenting in the time dimension to retain the behavior continuity and pre - processing in the spatial dimension to extract key features with high information density, both redundant data interference is avoided and the dynamic evolution law of the behavior sequence is ensured to be completely captured, enabling the abnormal behavior recognition model to quickly lock the risk time period and spatial area and improve the real - time response speed to sudden abnormal behaviors. Further, based on the feature - matching results, different warning signals are dynamically triggered, and the feature library and model parameters are continuously optimized in combination with the feedback data of the terminal device, forming a closed - loop control mechanism of "recognition - warning - verification - iteration". This not only reduces the interference of false alarms on the production process but also can intervene in advance at the initial stage of illegal behaviors through the prediction mechanism, greatly reducing the incidence of factory - area safety accidents. The skeletal key - point features can accurately identify illegal postures (such as holding dangerous items), the environmental object contour features can associate specific equipment illegal operations (such as unauthorized power - connected equipment identification), and the action - trajectory features can detect abnormal movement paths (such as intrusion into dangerous areas). The synergistic effect of the three ensures that various risk behaviors can still be stably identified in complex scenarios such as light changes and occlusion interference. Through the dynamic update strategy of the benchmark feature vector of the abnormal behavior feature library, it can adapt to changes in different production - line layouts and equipment configurations, avoid the lagging defects of traditional static rule libraries, ensure that the recognition model continuously adapts to the evolution of the factory - area environment, and form a long - term available intelligent security system.

[0033] As an implementation manner, in step S120, the continuous video - stream data is segmented in the time dimension and pre - processed in the spatial dimension to obtain a multi - frame behavior feature set corresponding to the dynamic behavior sequence, including: Step S121: The continuous video - stream data is divided into multiple video segments according to a preset time interval, and each video segment contains a fixed number of consecutive video frames.

[0034] The preset time interval is, for example, the length of the video segmentation time preset according to the monitoring requirements of the target factory area, and its value is determined according to the time resolution requirements for the behavior recognition of the target object. For example, in a scenario where rapid action changes need to be captured (such as detecting arm - raising and smoking), it is set to 0.1 seconds, while in the monitoring of low - frequency long - cycle behaviors (such as a person staying in a restricted area for a long time), it is extended to 0.5 seconds (corresponding to 2 frames per second). The fixed number of consecutive video frames is, for example, the upper limit of the number of frames contained in each video segment. For example, in a configuration with a time interval of 0.04 seconds, a single video segment contains 30 frames of data, corresponding to 1.2 seconds of continuous monitoring footage, to ensure that a complete action cycle of the target object can be covered (such as the operation process of an operator from picking up a part to installing it completely). In a specific implementation, the system buffers and manages the input data through the video stream processing module, and triggers the frame cutting operation according to the preset time interval, generating a sequence of video segments that are time - continuous and non - overlapping.

[0035] Step S122: For each video frame, call the pre - trained skeletal key - point detection network to extract the skeletal key - point features of the target object. The skeletal key - point features include the three - dimensional coordinates of each joint point and the motion speed parameters between adjacent joint points.

[0036] The pre - trained skeletal key - point detection network is, for example, a deep - learning model trained based on a large - scale human body structure and pose dataset (such as the COCO key - point dataset). Its architecture can adopt the OpenPose framework based on a convolutional neural network or the ViTPose model based on Transformer, and has the ability to recover the three - dimensional joint coordinates of the target object from a monocular or binocular video stream. The three - dimensional coordinates of each joint point are represented by the normalized spatial position data output by the model. For example, in human pose estimation, the head joint point coordinates are (x head , y head , z head ), where z head is derived through binocular camera parallax calculation or a monocular depth estimation network. The motion speed parameters between adjacent joint points are, for example, the instantaneous velocity vectors calculated based on the displacement of joint points between consecutive video frames. For example, through the difference between the elbow joint coordinates (x 1 , y 1 ,z 1 ) and (x 2 , y 2 , z 2 ) in two adjacent frames (time difference Δt) divided by Δt, obtaining the three - dimensional velocity components (v x , v y , v z) In the personnel monitoring scenario of a liquor production plant, the system tracks the shoulder joints and wrist joints of the operating personnel in the video frame. When it detects that the speed of the shoulder joints suddenly increases and the height of the wrist joints drops sharply, it can initially determine that there is an abnormal behavior tendency of throwing dangerous items. At this time, the skeletal key point features will include the quantification index of abnormal accelerated motion.

[0037] Step S123: Synchronously call the object contour segmentation network to perform contour recognition on the objects in the video frame that have spatial interaction with the target object, and extract the environmental associated object contour features. The environmental associated object contour features include object shape descriptors and relative distance parameters from the target object.

[0038] The object contour segmentation network is, for example, a neural network model optimized based on semantic segmentation tasks (such as Mask R-CNN or U-Net). It identifies the object regions in the video frame that have physical contact or spatial proximity relationships with the target object through pixel-level classification, such as the objects held by personnel and the goods loaded on transportation vehicles. The object shape descriptor uses mathematical methods such as Fourier descriptors or Hu invariant moments to parameterize the contour polygon. For example, after discretizing the tool contour into 128 boundary points, its Fourier transform coefficients are calculated to form a 256-dimensional feature vector for distinguishing different types of objects (such as wrenches and hammers). The relative distance parameter is obtained by calculating the Euclidean distance between the joint points of the target object (such as the center point of the human hand) and the centroid of the object contour. If a binocular camera or depth sensor is used, the three-dimensional space distance can be directly obtained; in a monocular video scenario, it needs to be estimated through a perspective projection model and the known object size. For example, when monitoring a forklift transporting a pallet in a warehousing and logistics area, the system identifies the pallet contour and calculates its distance from the tip of the forklift's fork. When the distance exceeds the safety threshold (such as 2 meters) and the fork is not in the lifted state, the environmental associated object contour features will include the geometric relationship data of the abnormal separation state, providing a basis for subsequent determination of illegal operations.

[0039] Among them, as an implementation manner, step S123, calling the object contour segmentation network to perform contour recognition on the objects in the video frame that have spatial interaction with the target object and extract the environmental associated object contour features, may include the following steps: Step S1231: Call the object contour segmentation network to perform object region localization on the video frame, and generate the initial bounding box coordinates and confidence parameters of the target object.

[0040] The object contour segmentation network is, for example, a deep learning model trained based on the Faster R-CNN or YOLOv5 architecture. It roughly locates the target object in the input video frame through the Region Proposal Network (RPN) or the single-shot detection mechanism, and outputs the rectangular bounding box coordinates and the corresponding class confidence scores. The initial bounding box coordinates represent the upper left and lower right coordinates of the target object in the image in a normalized form. For example, in a video frame with a resolution of 1920×1080, the initial bounding box coordinates of the forklift target are (x_min = 0.35, y_min = 0.42, x_max = 0.68, y_max = 0.79), corresponding to a rectangular area with pixel coordinates from (672,453) to (1305, 853). The confidence parameter is the probability estimation value of the model for the existence of the target object within the bounding box. For example, when the forklift detection confidence reaches 0.92, it is determined that the bounding box is valid. In the scenario of personnel monitoring in a liquor production factory, the system calls the object contour segmentation network for the video frame of the operator, generates multiple initial bounding boxes, and the bounding box with the highest confidence (such as 0.95) corresponds to the operator's torso area, providing a benchmark for subsequent expansion of the detection area.

[0041] Step S1232: Expand the detection area by a preset ratio outward based on the initial bounding box coordinates to generate a candidate object search area, and the candidate object search area covers the spatial range within a preset radius around the target object.

[0042] The preset ratio is, for example, an expansion coefficient dynamically adjusted according to the target object category. For example, for a human target, it is expanded by 30% and 20% respectively along the width and height directions of the initial bounding box to cover the surrounding areas that the arms may reach. The calculation formula for the candidate object search area is: the width of the expanded bounding box = the original width × (1 + 2×horizontal expansion ratio), the height = the original height × (1 + 2×vertical expansion ratio), and the center point remains unchanged.

[0043] Step S1233: Call the candidate object detection network, traverse the pixels in the candidate object search area by sliding window, and identify a set of candidate bounding boxes of all potential interacting objects. The set of candidate bounding boxes includes the position coordinates of each candidate object and the object category prediction result.

[0044] The candidate object detection network can, for example, adopt a lightweight detection model based on MobileNetv3, which performs dense scanning within the candidate object search area at a stride of 8 pixels through a sliding window mechanism, and outputs the class probability distribution of objects within each window and the refined bounding box coordinates. The set of candidate bounding boxes removes redundant detection boxes with an overlap rate exceeding a threshold (such as 0.5) through the non-maximum suppression (NMS) algorithm, and retains the top K (such as K = 10) candidate objects with the highest confidence levels. For example, in a video frame of a warehousing and logistics area, candidate bounding boxes of a pallet (confidence level 0.88), a shelf column (confidence level 0.75), and a safety warning sign (confidence level 0.63) are detected within the candidate object search area, and their position coordinates are stored in the set in a normalized form. The object class prediction result is a multi-class probability vector output by the softmax function. For example, the probability value of the pallet class is 0.88, the shelf column is 0.75, the safety warning sign is 0.63, and the probability of the remaining classes is less than 0.1.

[0045] Step S1234: Based on the object class prediction result, match and filter it with a preset list of interactive object classes, and remove the candidate objects in the set of candidate bounding boxes that do not belong to the list to obtain a set of target interactive objects.

[0046] The preset list of interactive object classes is, for example, a list of object class names that have an operational association with the target object. For example, for the target of an operator, the list includes categories such as "safety helmet", "handheld tool", "chemical container", etc.; for the target of a forklift, the list includes categories such as "pallet", "shelf", "loading and unloading ramp", etc. The matching and filtering process traverses the set of candidate bounding boxes and retains the candidate objects whose class prediction results belong to the list and whose confidence levels exceed the threshold (such as 0.6). In the inspection scenario of an oil refinery, the candidate bounding boxes output by the candidate object detection network include "valve" (0.82), "pipe" (0.78), and "fire hydrant" (0.55). Among them, the fire hydrant is excluded because it is not included in the list of interactive object classes and its confidence level is lower than the threshold. The set of target interactive objects only retains the detection results of the valve and the pipe to ensure that subsequent processing focuses on relevant objects.

[0047] Step S1235: Input the image regions corresponding to the candidate objects in the set of target interactive objects into a pre-trained contour segmentation network to generate precise contour masks for each candidate object. The precise contour mask is a binary pixel matrix that describes the closed boundary of the object shape.

[0048] The pre-trained contour segmentation network adopts a U-Net architecture. Its encoder part extracts image features through residual convolutional layers, and its decoder part restores pixel-level contour details through upsampling and skip connections. The precise contour mask is a binary matrix with the same resolution as the input image region, where foreground pixels (value 1) represent the occupied area of the object, and background pixels (value 0) represent the non-object area.

[0049] Step S1236: Extract the object shape descriptor based on the precise contour mask. The object shape descriptor is encoded by a sequence of vertex coordinates of the contour polygon and curvature change features, and the curvature change features are generated by calculating the angular change rate of the lines connecting adjacent vertices.

[0050] The sequence of vertex coordinates of the contour polygon is obtained, for example, by simplifying the mask boundary into a polygon using the Douglas-Peucker algorithm. For example, the original contour containing 500 boundary points is compressed into 50 key vertices, and each vertex is represented by an (x i , y i ) coordinate pair. The calculation method of the curvature change feature is as follows: traverse the vertex sequence, calculate the included angle θ i-1 formed by three adjacent vertices P i , P i+1 , and then perform a first-order difference on the θ i sequence to obtain the curvature change rate Δθ i i . In the analysis of the wrench contour, the curvature change rate of the handle part approaches zero, while the wrench opening part shows periodic high curvature changes. This feature can effectively distinguish different tool categories. The object shape descriptor is finally encoded as a concatenated vector of normalized vertex coordinates and curvature change rates, such as a 200-dimensional vector (50 vertices × 2 coordinates + 50 curvature values), for subsequent similarity matching.

[0051] Step S1237: Obtain the coordinates of the torso center point in the skeletal key point features of the target object, calculate the centroid coordinates of the precise contour mask, and generate the relative distance parameter based on the three-dimensional space difference between the torso center point coordinates and the centroid coordinates.

[0052] The coordinates of the torso center point are, for example, the average of the three-dimensional coordinates of the key points in the thoracic or pelvic region of the skeletal key points of the target object. For example, the center point of the thoracic cavity of a human target is the spatial average position of the midpoints between the left and right shoulder joint points and the hip joint points. The centroid coordinates are obtained by calculating the average of the three-dimensional coordinates (including depth information) of all foreground pixels in the precise contour mask, and if a monocular camera is used, it is estimated through the proportional scaling assumption. The relative distance parameter includes the horizontal distance d x y = √((x t −x c ) 2 +(y​t −y c ) 2 ) and the vertical distance d z =|z t −z c |, where (t) represents the center point of the torso and (c) represents the center of mass. In a warehousing scenario, when the horizontal distance between the center point of the forklift torso and the center of mass of the pallet exceeds 3 meters, the relative distance parameter will trigger an abnormal separation warning, indicating a possible risk of cargo detachment.

[0053] Step S1238: Combine the object shape descriptor and the relative distance parameter of the same candidate object to form the environmental associated object contour feature.

[0054] The combination operation can be achieved through feature vector splicing or attention-weighted fusion. For example, connect a 200-dimensional object shape descriptor with a 2-dimensional relative distance parameter (d xy , d z ) to form a 202-dimensional feature vector, and then map it to a 256-dimensional unified feature space through a fully connected layer. This feature combination mechanism ensures that the system simultaneously perceives the geometric attributes and spatial relationships of objects, significantly improving the recognition accuracy of abnormal interaction behaviors.

[0055] Step S124: Based on the temporal relationship of the video segment, track the position change of the target object in adjacent video frames to generate the action trajectory feature, where the action trajectory feature includes a displacement direction sequence and an acceleration change curve.

[0056] The temporal relationship is, for example, the strict time order between frames within the video segment. The system establishes an index through frame numbers or timestamps to ensure the causal consistency of the trajectory tracking algorithm in the time dimension. The position change tracking uses a multi-object tracking algorithm (such as DeepSORT), predicts the position of the target object in the next frame through a Kalman filter, and performs data association with the actual detection results to eliminate tracking losses caused by occlusion or lighting changes. The displacement direction sequence records the movement direction of the target object between consecutive frames in the form of a vector sequence. For example, in the XY plane coordinate system, arrange the displacement vectors of 25 frames per second in chronological order to form a sequence containing 25 groups of (dx, dy) data, representing the temporal evolution pattern of the movement direction. The acceleration change curve is obtained by calculating the second-order difference of the displacement vector. For example, within a time interval of Δt, according to the velocity change amount Δ vx and Δ vyCalculate the tangential acceleration and the normal acceleration, and plot a curve graph showing the variation of the acceleration amplitude over time. In the scenario of factory vehicle monitoring, when it is detected that the acceleration curve of a certain transport vehicle in a straight channel shows high-frequency oscillation (for example, the acceleration fluctuates violently between ±2 m / s² within 0.5 seconds), the action trajectory feature can indicate that the vehicle has the risk of abnormal emergency braking or loss of control, triggering the subsequent analysis process.

[0057] Among them, as an implementation manner, in step S124, based on the timing relationship of the video clip, track the position change of the target object in adjacent video frames to generate the action trajectory feature, including: Step S1241: Extract the spatial coordinates of the torso center point of the target object from each video frame of the video clip, and arrange them in the order of the video frame acquisition time to generate an initial position coordinate sequence.

[0058] The spatial coordinates of the torso center point can be obtained, for example, from the three-dimensional coordinate data output by a bone key point detection network. For example, in the case of a human target, take the spatial average coordinate of the midpoints of the left and right shoulder joints and the hip joints as the torso center point; for a forklift target, take the geometric center point of the vehicle chassis. The initial position coordinate sequence is a coordinate array arranged in ascending order of timestamps. For example, in a 30-frame video clip, a sequence containing 30 groups of (x, y, z) data is formed, with a frame interval of 0.033 seconds. In the scenario of factory vehicle out-of-bounds detection, this sequence can accurately reflect the trajectory trend of the vehicle moving from the compliant driving area to the restricted area, providing basic data for trajectory analysis.

[0059] Step S1242: Calculate the difference between the spatial coordinates of the torso center point of two adjacent video frames in the initial position coordinate sequence to obtain the position change difference of the target object between adjacent video frames. The position change difference includes the horizontal direction offset and the vertical direction offset.

[0060] The difference calculation uses vector subtraction, that is, ΔP i = P i+1 − P i , where ΔP i includes the horizontal offset (Δx, Δy) and the vertical offset Δz. For example, if the torso center point of the i-th frame is (10.2 m, 5.5 m, 1.0 m) and the (i + 1)-th frame is (10.5 m, 5.3 m, 1.0 m), then the position change difference is (0.3 m, -0.2 m, 0.0 m). In the personnel monitoring in the three-dimensional storage area, the horizontal offset reflects the planar movement trend, and the vertical offset can be used to detect three-dimensional space abnormal behaviors such as climbing the shelves. When Δz exceeds 0.2 m for multiple consecutive frames, it indicates the potential risk of illegal climbing.

[0061] Step S1243: Determine the moving direction angle of the target object between each pair of adjacent video frames according to the proportional relationship between the horizontal direction offset and the vertical direction offset of the position change difference, and concatenate all the moving direction angles in chronological order to generate the displacement direction sequence.

[0062] The moving direction angle is calculated by the arctangent function: θ i = arctan2(Δy i , Δx i ), and the result is expressed in radians and normalized to the interval [0, 2π). For example, the angle corresponding to the horizontal offset Δx = 0.3m and Δy = -0.2m is 5.7596 radians (330 degrees), indicating movement in the northwest direction. The displacement direction sequence is an array containing N - 1 angle values (N is the number of video frames).

[0063] Step S1244: Calculate the instantaneous moving speed of the target object between each pair of adjacent video frames based on the ratio of the absolute value of the position change difference to the time interval between adjacent video frames, and arrange all the instantaneous moving speeds in chronological order to generate the speed change sequence.

[0064] The instantaneous moving speed v i = √(Δx i ² + Δy i ² + Δz i ²) / Δt, where Δt is the frame interval time (such as 0.033 seconds). For example, when Δx = 0.3m, Δy = -0.2m, and Δz = 0.1m, the instantaneous speed is √(0.09 + 0.04 + 0.01) / 0.033 ≈ 3.85m / s. The speed change sequence is smoothed by a sliding window to eliminate noise. For example, a moving average filter with a window size of 5 is used. In the inspection scenario of a liquor factory, it can effectively distinguish the speed patterns of normal walking (1.2m / s) and running (3m / s) of personnel, providing a quantitative basis for abnormal behavior determination.

[0065] Step S1245: Perform a time window sliding calculation on the difference between two adjacent instantaneous moving speeds in the speed change sequence to obtain the speed change difference of the target object within consecutive time windows, and fit all the speed change differences in chronological order to generate the acceleration change curve.

[0066] For example, the speed change difference Δv i = v i+1 -v i , and after a moving average process with a window size of 3, the smoothed acceleration sequence a i = (Δv i-1 +Δv i +Δv i+1) / 3Δt. The acceleration change curve fits the discrete acceleration points through cubic spline interpolation. In the logistics vehicle monitoring, this curve can clearly show the dynamic process of the vehicle from uniform motion (a≈0) to emergency braking (a < -3m / s²). When the abnormal acceleration peak exceeds the threshold, an early warning is triggered.

[0067] Step S1246: Bind the features of each movement direction angle in the displacement direction sequence with the rate change difference at the corresponding time node in the acceleration change curve to form fused trajectory data aligned with the time axis, and encode the fused trajectory data into the action trajectory features.

[0068] Feature binding is achieved, for example, by creating a timestamp index. For example, the i-th movement direction angle θ i is combined with the i-th acceleration value a i to form a data pair (θ i , a i ), constituting N-2 dimensional fused trajectory data. The encoding process uses a recurrent autoencoder to compress the time series data into a 128-dimensional latent space vector. In arm movement monitoring, this vector can characterize the smoothness, direction stability, and acceleration compliance of the movement trajectory. When the matching degree of the latent space vector with the normal operation mode library is lower than the threshold, it is determined as an abnormal trajectory feature, and a device shutdown and maintenance instruction is triggered. This encoding mechanism effectively fuses spatio-temporal motion features and significantly improves the recognition ability of complex trajectory patterns.

[0069] Step S125: Combine the bone key point features, environment-related object contour features, and action trajectory features corresponding to the same video frame to form the multi-frame behavior feature set.

[0070] The combination operation is, for example, a process of converting heterogeneous feature data into a unified format and performing spatial or channel dimension splicing. For example, the bone key point features (256-dimensional vector), environment-related object contour features (128-dimensional vector), and action trajectory features (64-dimensional vector) are concatenated in the feature dimension to form a 448-dimensional combined feature vector; or they are respectively mapped to a high-dimensional space and then weighted and fused through an attention mechanism. The multi-frame behavior feature set finally shows a three-dimensional tensor structure, and its dimensions are the number of frames, the number of feature types, and the feature dimension number. For example, for a video clip containing 30 frames, if 3 types of features are extracted for each frame and each type of feature dimension is 512, the set shape is 30×3×512.

[0071] As an implementation manner, before the step S122 of calling a pre-trained bone key point detection network to extract the bone key point features of the target object, the method provided by the embodiments of the present invention further includes a step of training the bone key point detection network, which may specifically include: Step S101: Obtain a historical monitoring video dataset in the factory area scene, and perform manual annotation on each video frame in the historical monitoring video dataset. The annotation content includes the joint positions and joint movement labels of the target object.

[0072] The historical monitoring video dataset is, for example, a collection of long-term video data collected from monitoring devices actually deployed in the target factory area, covering different lighting conditions, equipment operating states, and personnel operation scenarios. For example, 3600 hours of video collected continuously for three months in the fermentation workshop of a liquor production factory, including various scenarios such as personnel operations, vehicle transportation, and personnel patrols. In the manual annotation process, a professional annotation tool (such as LabelMe or CVAT) is used to mark the joints of the target object in the video frame frame by frame. The joint positions are recorded in the form of pixel coordinates or three-dimensional space coordinates, depending on the type of sensor (monocular / binocular camera or RGB-D device with depth information). The joint movement label is the classification result of the joint movement state annotated based on time series analysis. For example, the rotating joint of the arm is annotated as "uniform rotation" or "sudden stop and acceleration", and the human knee joint is annotated as "bending" or "extension". In the monitoring scenario of workers in a liquor production factory, the annotator marks the coordinates of 17 human body joints (including shoulders, elbows, wrists, hips, knees, etc.) of the workers in the video frame, and annotates the joint movement labels based on the displacement between consecutive frames, such as "horizontal swing of the upper limb" or "climbing action of the lower limb", to form structured training data.

[0073] Step S102: Construct an initial key point detection network, where the initial key point detection network includes a feature pyramid module and a three-dimensional coordinate regression module.

[0074] The feature pyramid module adopts a multi-scale feature fusion architecture (such as HRNet or FPN), and extracts local detailed features and global semantic features of the video frame through parallel convolution paths. For example, at an input resolution of 512×512 pixels, feature maps of three scales of 1 / 4, 1 / 8, and 1 / 16 are generated, which respectively capture the microscopic texture of the fingers, the contour of the human torso, and the background information of the factory area environment. The three-dimensional coordinate regression module consists of a fully connected layer and a coordinate transformation unit. Its input is the multi-scale feature vector output by the feature pyramid, and the output is the three-dimensional space coordinates (x, y, z) and motion speed parameters (v x , v y ,v z ). In the personnel detection scenario in a three-dimensional warehouse, the module calculates the depth information through the binocular camera parallax, maps the two-dimensional image coordinates to the three-dimensional coordinates in the warehouse coordinate system, and realizes the three-dimensional reconstruction of the position of the forklift driver's hand. In the network initialization stage, the weights of the convolutional layer are assigned using the He normal distribution, and the ImageNet pre-trained parameters are loaded to accelerate convergence.

[0075] Step S103: Input the video frames in the historical monitoring video dataset into the initial key point detection network, extract multi-scale spatial features through the feature pyramid module, and predict the three-dimensional coordinates and motion speed parameters of each joint point through the three-dimensional coordinate regression module.

[0076] Before the video frame input, preprocessing including normalization is required, such as scaling the resolution to the network input size (e.g., 512×512), normalizing the pixel values to the range [-1,1], and data augmentation operations (random rotation of ±15 degrees, brightness adjustment of ±20%). The feature pyramid module performs multi-level convolution and pooling operations on the preprocessed image. For example, in the HRNet architecture, high-dimensional information of feature maps at different scales is exchanged through branch networks to generate feature vectors that fuse microscopic finger textures and macroscopic arm postures. After receiving the feature vectors, the three-dimensional coordinate regression module maps them to the coordinates and speed parameters of the joint points through fully connected layers. For example, the predicted three-dimensional coordinates of the arm joints are (2.3m, 4.7m, 1.5m), and the motion speed is (0.2m / s, -0.1m / s, 0.05m / s).

[0077] Step S104: Calculate the Euclidean distance loss value between the predicted coordinates and the labeled coordinates, and combine the difference loss value between the motion speed parameters and the joint point motion labels to generate the total training loss value.

[0078] The Euclidean distance loss value L coord The calculation formula is: , where N is the total number of joint points, is the predicted coordinate, is the labeled coordinate. The difference loss value uses the smooth L1 loss function to calculate the deviation between the predicted speed and the labeled motion label. For example, when the joint point motion label is "accelerating", if the predicted speed direction is opposite to the labeled direction, a high loss value will be generated. The total training loss value L total =αL coord +βL velocity , where α = 0.7 and β = 0.3 are hyperparameters determined through grid search.

[0079] Step S105: Perform backpropagation optimization on the initial key point detection network based on the total training loss value until the total training loss value converges to a preset threshold to obtain the pre-trained skeletal key point detection network.

[0080] Backpropagation optimization uses the Adam optimizer, with the initial learning rate set to 3e−4, and the cosine annealing strategy is implemented to adjust the learning rate based on the validation set loss value. During the training process, 32 video frames are input in each batch. After calculating the loss through forward propagation, the network weight parameters are updated by the chain rule of differentiation. The convergence condition is set such that the fluctuation range of the total training loss value for 10 consecutive training epochs is less than 1e−4, and the validation set accuracy reaches over 98%. In the training of the smoking behavior dataset, after 120 epochs of training, the total training loss value drops from the initial 3.2 to 0.08 (the preset threshold is 0.1). At this time, the average prediction error of the network for the coordinates of the arm joint points stabilizes at 2.3 cm, and the accuracy of the speed direction matching reaches 96.5%, meeting the accuracy requirements for real-time monitoring in the factory area. After the trained network is deployed to the edge computing device, it can perform high-precision joint point detection on the target objects in the real-time video stream, providing reliable feature inputs for abnormal behavior recognition.

[0081] As an implementation manner, in step S130, the multi-frame behavior feature set is input into the pre-trained abnormal behavior recognition model. Through the multi-modal feature encoding layer in the abnormal behavior recognition model, cross-modal correlation fusion is performed on the bone key point features, environment-related object contour features, and action trajectory features to generate a target behavior feature vector, which may specifically include: Step S131: Input the bone key point features into the first feature encoding sub-network in the multi-modal feature encoding layer to generate a first encoded feature vector. The first feature encoding sub-network uses a temporal convolutional structure to capture the joint movement pattern.

[0082] The temporal convolutional structure is, for example, a stack of neural network layers where one-dimensional convolutional kernels slide along the time axis to perform feature extraction. It expands the receptive field by setting different dilation coefficients (such as 1, 2, 4) to capture long-period joint movement patterns. The input of the first feature encoding sub-network is a three-dimensional tensor composed of multi-frame bone key point features, with dimensions of time step × number of joint points × coordinate dimension (such as 30×17×3), and the output is the first encoded feature vector (such as 256-dimensional) after multiple layers of dilated convolution and non-linear activation.

[0083] Step S132: Input the environment-related object contour features into the second feature encoding sub-network in the multi-modal feature encoding layer to generate a second encoded feature vector. The second feature encoding sub-network uses an attention mechanism to focus on the object contour with the highest interaction frequency with the target object.

[0084] The attention mechanism is, for example, a computational module that dynamically assigns feature importance through a learnable weight matrix. Its query vector comes from the skeletal key-point features of the target object, and the key-value vector comes from the contour features of the environment-related objects. The attention weights are generated by calculating the similarity between the query and the key. After the second feature encoding sub-network performs position encoding on the single-frame object contour features, it is input into the multi-head self-attention layer to select the object contour that is closest to the target object in terms of spatial distance and has the longest contact time (such as the contour of the pallet continuously contacted by the forklift fork), and suppress the interference of irrelevant background objects. In the personnel monitoring in the stereoscopic warehousing area, when the operator maintains a distance less than 0.5 meters from the contour of the shelf column for 5 consecutive frames, the weight value corresponding to the shelf column in the attention weight matrix is increased to 0.92, while the weight of the distant tool cart is reduced to 0.05, so that the second encoded feature vector centrally encodes the shape and distance parameters of the high-risk interaction objects. The network aggregates the temporal attention weights through a gated recurrent unit (GRU) to further strengthen the feature expression of the high-frequency interaction objects.

[0085] Step S133: Input the action trajectory features into the third feature encoding sub-network in the multi-modal feature encoding layer to generate a third encoded feature vector. The third feature encoding sub-network uses a recurrent neural network structure to analyze the continuity of the displacement direction sequence.

[0086] The recurrent neural network structure selects a bidirectional long short-term memory network (Bi-LSTM), with 128 hidden layer units. The forward and backward LSTM units are used to capture the forward evolution law and the reverse causal relationship of the displacement direction sequence respectively. The input of the third feature encoding sub-network is the displacement direction angle sequence in the action trajectory features (such as 29 angle values corresponding to 30 frames), and the output is the third encoded feature vector (such as 128-dimensional) that fuses the temporal context information. In the scenario of vehicle out-of-bounds detection in the factory area, the Bi-LSTM network models the trajectory sequence of the forklift moving from the compliant area to the restricted area. When it is detected that the displacement direction angles of 8 consecutive frames continuously deviate towards the coordinate orientation of the restricted area (such as gradually shifting from 90 degrees to 135 degrees), the activation value of the dimension representing direction consistency in the third encoded feature vector exceeds the threshold, indicating a potential violation path. The network introduces a temporal attention mechanism in the output layer to assign higher weights to the features at the turning point of the trajectory, improving the sensitivity to sudden turning behaviors.

[0087] Step S134: Concatenate the first encoded feature vector, the second encoded feature vector, and the third encoded feature vector, and perform dimensionality reduction and fusion through a fully connected layer to obtain the target behavior feature vector.

[0088] The concatenation operation connects the first encoded feature vector (256 dimensions), the second encoded feature vector (256 dimensions), and the third encoded feature vector (128 dimensions) into a 640-dimensional composite vector along the feature dimension, which is then input into a two-layer fully connected network (with dimensions of 512 and 256, respectively). The ReLU activation function and the Dropout layer (ratio 0.3) are used to prevent overfitting, and finally a 256-dimensional target behavior feature vector is output. During the dimensionality reduction process, the fully connected layer performs nonlinear transformations on cross-modal features through weight matrices, such as enhancing the implicit association between joint motion and object interaction, suppressing redundant information, and maximizing the class separability of the target behavior feature vector in the latent space.

[0089] As an implementation method, the training process of the abnormal behavior recognition model includes the following steps: Step S201: Collect historical abnormal behavior video samples and normal behavior video samples in the factory area, extract the multi-frame behavior feature set for each video sample, and mark the corresponding abnormal behavior category label.

[0090] The historical abnormal behavior video samples cover the safety incidents that occurred in the target factory in the past three years, including 16 types of abnormal behaviors such as people climbing shelves, smoking, and using electronic devices in violation of regulations. Each type of sample has no less than 200 segments, and the duration is 5-30 seconds; the normal behavior video samples are randomly selected from daily monitoring, including standard operating procedures, compliance inspection routes, and routine equipment operations. The multi-frame behavior feature set extraction process follows the preprocessing process of steps S120-S124 to generate a structured data set containing skeleton key point features, environment-related object contour features, and motion trajectory features. Abnormal behavior category labels are encoded using One-hot encoding. For example, "entering a high-risk area without a safety helmet" corresponds to the label vector [1,0,…,0], and "smoking" corresponds to [0,1,…,0].

[0091] Step S202: constructing an initial abnormal behavior recognition model, wherein the initial abnormal behavior recognition model includes the multimodal feature encoding layer and the classifier layer.

[0092] The multimodal feature encoding layer consists of the first, second, and third feature encoding subnetworks and feature fusion modules defined in steps S131-S134. The classifier layer uses a two-layer fully connected network (256→128→16) plus a Softmax function to output the probability distribution of 16 types of abnormal behaviors. In the model initialization stage, the convolution kernel weights are initialized using the Xavier normal distribution, and the LSTM unit forget gate bias is set to 1.0 to alleviate the gradient disappearance. The model input dimension is adapted to, for example, 30 frames × 3 modal features, and the output layer is designed for the classification task of 14 types of abnormal behaviors and 1 normal category, and the dimensions of the fully connected layer are adjusted to match the actual needs.

[0093] Step S203: Input the multi-frame behavior feature set into the initial abnormal behavior recognition model, output the target behavior feature vector through the multi-modal feature encoding layer, and calculate the matching probabilities between the target behavior feature vector and the feature templates corresponding to each abnormal behavior category through the classifier layer.

[0094] The matching probability calculation uses cosine similarity measurement. Each category in the classifier layer corresponds to a learnable feature template vector (for example, 16 categories correspond to 16 256-dimensional vectors). After calculating the dot product between the target behavior feature vector and each template vector, it is normalized to a probability value by Softmax. In the behavior recognition of stereoscopic warehousing personnel, when the input feature represents the behavior of "throwing goods", the cosine similarity between the target behavior feature vector and the "throwing goods" template reaches 0.91, and the Softmax output probability is 0.87, which is significantly higher than other categories. During the model training stage, the parameters of the template vector and the feature encoding layer are optimized through backpropagation, so that the feature vectors of the same class samples are closely clustered in the latent space.

[0095] Step S204: Update the parameters of the initial abnormal behavior recognition model based on the cross-entropy loss value between the matching probability and the abnormal behavior category label.

[0096] The calculation formula of the cross-entropy loss value is , where y c is the one-hot encoding of the true label, and p c is the category probability output by Softmax. The parameter update uses the Adam optimizer, the initial learning rate is set to 1e−4, and it decays by 0.5 times every 10 training epochs. An early stopping mechanism is introduced during the training process. When the validation set loss does not decrease for 5 consecutive epochs, the training is terminated to prevent overfitting.

[0097] Step S205: Introduce an adversarial training strategy, add noise perturbation data to the multi-frame behavior feature set, calculate the classification consistency loss value between the perturbed feature set and the original feature set, and perform joint optimization in combination with the cross-entropy loss value until the model converges.

[0098] The adversarial training strategy can generate adversarial samples through the fast gradient sign method (FGSM). The noise perturbation intensity ε is set to 0.1, and perturbations are added along the loss gradient direction in the feature space. The classification consistency loss value L c onsist measures the KL divergence of the output probability distributions of the original feature set and the perturbed feature set. The joint optimization objective function is L t otal = L c e + λL consist (λ = 0.5). In the data training of the warehouse fire emergency drill, adversarial samples simulate the offset of the bone key point coordinates caused by camera jitter, and the model minimizes L c onsist to improve the robustness to noise, so that when frame loss or sudden light change occurs during transmission in the real scenario, the fluctuation range of the classification probability is controlled within ±0.05. The training termination condition is set that the fluctuation amplitude of the total training loss value is less than 1e−5 for 20 consecutive epochs, and the accuracy of the validation set reaches more than 98.5%. Finally, the model is deployed to the edge server of the factory area to realize real-time abnormal behavior recognition.

[0099] As an implementation manner, in step S140, based on the similarity matching between the target behavior feature vector and the preset abnormal behavior feature library, determining whether the behavior type of the target object belongs to the predefined abnormal behavior category includes: Step S141: Read all the reference feature vectors of the predefined abnormal behavior categories from the abnormal behavior feature library, and the reference feature vectors are obtained by centering the features extracted from the historical abnormal behavior data through a clustering algorithm.

[0100] The abnormal behavior feature library is a distributed database storing the reference feature vectors of all known abnormal behavior categories in the target factory area. Its reference feature vectors are generated by centering the multi-frame behavior feature sets extracted from the historical abnormal behavior case videos through a hierarchical clustering algorithm. For example, in the scenario of a liquor production factory, there are 12 types of abnormal behaviors such as operators entering high-risk areas without wearing safety helmets, smoking, and making or answering phone calls. Each type contains 3 - 5 clustering groups to cope with different manifestation forms of the same abnormal behavior. Centering processing means calculating the mean value of the feature dimensions for the multi-frame behavior feature sets within the same clustering group. For example, the reference feature vector of a certain group of the raising hand behavior is a 256-dimensional vector, and the arithmetic mean value of 500 historical feature samples within the group is taken for each dimension. The reading operation is implemented through the parallel query interface of the distributed storage node. When an abnormal detection is triggered in the monitoring center of the liquor production factory, all the reference feature vectors are loaded from the storage node to the memory cache to form a retrieval queue containing category labels, priority weights, and clustering group numbers, ensuring the real-time matching efficiency.

[0101] Step S142: Calculate the cosine similarity between the target behavior feature vector and each reference feature vector, and select the abnormal behavior categories corresponding to the top K reference feature vectors with the highest similarity as the candidate categories.

[0102] The cosine similarity is calculated using the ratio formula of the dot product of vectors to the product of the vector lengths: sim = (A·B) / (||A||×||B||), where A is the target behavior feature vector and B is the reference feature vector. In the monitoring scenario of a liquor production plant, the similarity between the target behavior feature vector and the reference feature vector of the "smoking" category reaches 0.93, and the similarity with the "illegal power connection" category is 0.62. When the system sets K = 3, the top three candidate categories are selected. The dynamic adjustment threshold is set according to the historical false alarm rate. For example, when the highest similarity of the candidate categories is lower than 0.85, the threshold is decreased by 5% to expand the retrieval range. After the candidate category list is generated, the system attaches the associated clustering group number and severity level to the matching result for subsequent verification process calls.

[0103] Step S143: Perform a time dimension alignment check on the reference feature vectors of the candidate categories to determine whether the change trend of the target behavior feature vector within consecutive time windows is consistent with the behavior pattern of the candidate categories.

[0104] The time dimension alignment check is implemented through the dynamic time warping (DTW) algorithm. The time series change curve of the target behavior feature vector is non-linearly aligned with the typical pattern curve of the reference feature vector of the candidate category, and the path cost is calculated to evaluate the trend consistency. For example, in the detection of personnel smoking behavior in a liquor production plant, the change curve of the hand movement parameters of the target feature vector within consecutive 10 frames needs to maintain a DTW path cost less than 0.15 with the standard smoking pattern curve of the reference feature vector. If it is detected that when an operator enters a high-risk area, the difference in the rising slope of the height change trend of the skeletal key points exceeds 20% compared with the reference feature vector of the "climbing the shelf" category, the time dimension check is determined to fail. A sliding window mechanism is introduced during the check, and local alignment is performed every 5 frames to ensure that sudden abnormal behaviors (such as a rapid run within 0.5 seconds) are not misjudged as normal fluctuations.

[0105] Step S144: If the check passes, determine the abnormal behavior category with the highest similarity in the candidate categories as the behavior type of the target object; otherwise, mark the behavior type as an unknown abnormal category and trigger the manual review process.

[0106] The verification pass condition needs to simultaneously meet the similarity threshold and the DTW path cost threshold. For example, the similarity ≥ 0.85 and DTW ≤ 0.2. In the scenario of personnel monitoring in a liquor production factory, when the similarity between the target behavior feature vector and the "illegal smoking" benchmark feature vector is 0.91 and DTW = 0.18, the system automatically marks this behavior as "illegal smoking" and triggers an alarm. If the similarity between a certain arm movement feature vector and the "smoking" benchmark feature vector reaches 0.88 but DTW = 0.25 (exceeding the threshold of 0.2), then this behavior is classified as an unknown abnormal category, and an alarm work order containing video clips, feature vectors, and matching details is sent to the safety officer's handheld terminal synchronously, requiring manual confirmation whether it is a new abnormal type. After the manual review passes, the multi-frame behavior feature set of the new abnormal type will be imported into the training set, triggering the incremental update process of the feature library.

[0107] As an implementation method, the abnormal behavior feature library can be constructed through the following steps: Step S301: Collect video of abnormal behavior cases that have occurred within the target factory area. The abnormal behavior cases include illegal smoking, illegal use of electronic devices, and intrusion into dangerous areas.

[0108] The video of abnormal behavior cases is sourced from the event video archives of the target factory area's security system, covering all safety accident records manually confirmed in the past three years. The video of the illegal smoking case needs to include the complete sequence of ignition, holding the cigarette, and smoking actions, with a duration of 10 - 30 seconds; the video of the intrusion into dangerous areas case needs to record the entire process video from the trigger of the out-of-bounds alarm to the evacuation of personnel. The video acquisition device uses an explosion-proof infrared camera to ensure that the video clarity in a high-risk environment reaches the 1080p / 25fps standard, and additional depth sensor data is attached for three-dimensional behavior reconstruction.

[0109] Step S302: Segment each video of abnormal behavior cases, extract the multi-frame behavior feature set within each time period, and label the corresponding abnormal behavior category and severity level.

[0110] The segmentation process cuts the continuous video into independent segments according to the start and end timestamps of the behavior occurrence. For example, a 2-minute illegal smoking video is segmented into three sub-segments: the ignition stage (0 - 5 seconds), the smoking stage (6 - 45 seconds), and the cigarette butt discarding stage (46 - 50 seconds). The process of extracting the multi-frame behavior feature set calls the preprocessing process of steps S120 - S124 to generate feature data including skeletal key points, object contours, and action trajectories. The severity level is divided into three levels according to the severity of the behavior consequences: level one (immediate danger, such as intrusion into a high-voltage area), level two (potential danger, such as not wearing a safety helmet), and level three (violation but low risk, such as using a mobile phone). In the case of a stereoscopic warehouse, smoking is marked as level two, and the over-high and tilted stacking of goods is marked as level three. The system adjusts the alarm response speed according to the level.

[0111] Step S303: Use the hierarchical clustering algorithm to group the multi-frame behavior feature sets of the same abnormal behavior category, and calculate the feature mean of each group as the benchmark feature vector.

[0112] The hierarchical clustering algorithm adopts the Ward variance minimization method, and sets the clustering termination condition as the intra-class distance being less than 0.1 and the inter-class distance being greater than 0.3. The benchmark feature vector of each group is obtained by calculating the mean of each dimension of the samples within the group. The clustering results are manually reviewed to confirm the rationality of the grouping, and the abnormal groups with too high feature dispersion (such as groups with sample standard deviation > 0.15 within the group) are removed.

[0113] Step S304: Assign a priority weight to each benchmark feature vector according to the severity level, and store the benchmark feature vector and its priority weight in the abnormal behavior feature library.

[0114] For example, the priority weight calculation is W = 1 + 0.5×(3 - L), where L is the severity level (levels 1 - 3). The weight of the benchmark feature vector for level 1 abnormal is 2.0, for level 2 is 1.5, and for level 3 is 1.0. When storing, a columnar database is used to sort in descending order of weight, and the high-weight categories are preferentially retrieved in the real-time matching process. For example, in the feature library of a liquor production factory, the weight of the "smoking" category is set to 2.0, and the benchmark feature vectors of its 3 clustering groups are stored in the memory hot zone to ensure that the matching latency is less than 10 ms. The distributed storage nodes adopt a three-replica mechanism to prevent data loss and achieve fast positioning through the consistent hashing algorithm.

[0115] As an implementation, in step S141, read all the benchmark feature vectors of the predefined abnormal behavior categories from the abnormal behavior feature library. The benchmark feature vectors are obtained by centralizing the features extracted from the historical abnormal behavior data through a clustering algorithm, and may include the following steps: Step S1411: Obtain the historical abnormal behavior case video data within the target factory area. The historical abnormal behavior case video data contains video stream segments labeled with abnormal behavior categories and severity levels.

[0116] The historical abnormal behavior case video data is extracted from the Oracle database of the factory security system, and includes video files, metadata tables (recording device numbers, timestamps, geographical coordinates), and annotation information tables. Invalid videos with a resolution lower than 720p and more than 30% key frame loss are removed through the data cleaning process to ensure that the input quality meets the feature extraction requirements. The video stream segments are stored in the H.265 encoding format, and the FFmpeg library accelerated by GPU is called during decoding to improve the processing efficiency.

[0117] Step S1412: Decode the historical abnormal behavior case video data frame by frame and perform noise filtering processing to generate a set of preprocessed video frames.

[0118] Frame-by-frame decoding can adopt the hardware-accelerated Video4Linux framework to convert the video stream into an independent frame sequence in YUV420p format. Noise filtering can include Gaussian filtering (kernel size 5×5, σ = 1.5) to eliminate high-frequency noise, median filtering (kernel size 3×3) to suppress impulse noise, and the CLAHE algorithm to enhance low-contrast regions.

[0119] Step S1413: Call the skeletal key point detection network and the object contour segmentation network for each video frame in the set of preprocessed video frames, extract the skeletal key point features, environment-related object contour features, and action trajectory features, and form a historical multi-frame behavior feature set.

[0120] The skeletal key point detection network can adopt the HRNet-W48 model with an input size of 512×512, and output the three-dimensional coordinates and confidence levels of 17 joint points of the human body. The object contour segmentation network is implemented based on Mask R-CNN to generate pixel-level object masks and class labels. The action trajectory features are extracted by the FairMOT tracking algorithm to obtain the displacement vectors between consecutive frames. The single-frame processing takes 35 ms. The extracted skeletal features include the key point coordinates of the driver's hands and waist, the environmental features include the Fourier descriptors of the tray contour, and the trajectory features record the displacement direction angle of the forklift per second. The historical multi-frame behavior feature set is stored in an Apache Parquet columnar file in chronological order, with a compression rate increased by 40%, accelerating the data reading during clustering analysis.

[0121] Step S1414: Group the historical multi-frame behavior feature set according to the abnormal behavior categories, and apply the hierarchical clustering algorithm to the historical multi-frame behavior feature sets of the same abnormal behavior category to generate multiple clustering groups, where each clustering group contains a subset of multi-frame behavior features with a feature similarity higher than a preset threshold.

[0122] During the grouping process, the feature set is divided into different directories according to the labeled abnormal category ID. For example, the "illegal smoking" category contains 1,200 feature samples. Hierarchical clustering is implemented using AgglomerativeClustering in the scikit-learn library. The distance metric selects the Manhattan distance to adapt to high-dimensional sparse features, and the dendrogram cutting height is set to 0.25.

[0123] Step S1415: Calculate the mean value of the subset of multi-frame behavior features within each clustering group, and use the arithmetic mean value of each feature dimension as the benchmark feature vector of the clustering group.

[0124] The mean value calculation can use the nanmean function of the Numpy library to handle missing values. For example, in the 256-dimensional feature vector of the smoking behavior group A, the mean value of the 128th dimension (related to torque) is 0.76, and the standard deviation is 0.05. After the benchmark feature vector is generated, L2 normalization is performed to ensure the effectiveness of the cosine similarity calculation. In the case of illegal climbing and power connection, the mean value of the benchmark feature vector in the height change dimension is +0.3 m / s, accurately reflecting the continuous upward movement trend. Statistical information such as the number of additional samples and the maximum outlier distance of the benchmark vector for each clustering group is provided for subsequent priority calculation.

[0125] Step S1416: Associatively store the benchmark feature vector with the corresponding abnormal behavior category, clustering group number, and severity level in the distributed storage nodes of the abnormal behavior feature library.

[0126] For distributed storage, for example, the Cassandra database is used. The wide table structure is designed to include fields such as feature vector (blob type), category ID (int), grouping number (varchar), weight (float), etc. The primary key is designed as (category ID, grouping number) to support fast retrieval by category.

[0127] Step S1417: When the target behavior feature vector is received, load all the benchmark feature vectors of abnormal behavior categories from the distributed storage nodes, and form a priority retrieval queue in descending order according to the severity level.

[0128] The loading process can use multi-threaded parallel reading, and each thread is responsible for the range scan of a category ID. For example, in the monitoring center, it takes 150 ms to load 120 benchmark feature vectors, and the first-level abnormal categories are arranged in the front when forming the priority queue. The queue memory structure is implemented using a max heap, and the weight value is used as the heap key value to ensure that high-priority features participate in the matching first. The retrieval queue is attached with the most recent access timestamp, and the LRU strategy is used to eliminate low-priority features that have not been used for more than 30 days to keep the memory occupancy within 8 GB.

[0129] Step S1418: Traverse the benchmark feature vectors in the priority retrieval queue, and sequentially perform similarity matching between the target behavior feature vector and each benchmark feature vector in the queue until the preset maximum number of matches is reached or the matching similarity exceeds the dynamically adjusted threshold.

[0130] The traversal algorithm adopts an early termination strategy. When the similarity is detected to exceed 0.95, the result is immediately returned; otherwise, the matching continues until the queue traversal is completed or the 100 - match upper limit is reached. The dynamic adjustment threshold changes automatically according to the real - time system load. When the CPU utilization rate is higher than 80%, the threshold is increased by 5% to reduce the amount of calculation. For example, when the matching similarity between the target feature vector and the 15th reference feature vector (smoking category) reaches 0.96, the system terminates the matching and returns the result, and the entire round of matching takes 22 ms. The matching log records detailed process data for subsequent algorithm optimization and misdiagnosis analysis.

[0131] As an implementation manner, in step S150, if it is determined to belong to the abnormal behavior category, a corresponding warning signal is triggered according to the behavior type, which may specifically include: Step S151: When the behavior type is the first - type abnormal behavior, a first - level warning signal is generated, the audible and visual alarm in the target factory area is controlled to start, and at the same time, an emergency warning message including a behavior screenshot and location information is sent to the terminal device of the security management personnel.

[0132] The first - type abnormal behavior is, for example, a predefined abnormal behavior category with immediate danger, such as an operator smoking, illegally using electronic devices, illegally climbing to connect electricity, etc. The generation logic of the first - level warning signal is associated with the matching result of the reference feature vector marked with the highest - priority weight in the abnormal behavior feature library. A pulse - modulation signal can be sent to the audible and visual alarm deployed in the target area through the factory Internet of Things control bus to trigger it to start with a 120 - decibel beeping frequency and a red strobe light mode, covering the audible and visual warning within a radius of 50 meters. The construction of the emergency warning message includes the current timestamp, the monitoring device number, the geographical coordinates, and a 5 - second key frame sequence of the behavior intercepted from the video stream (JPEG format with a compression rate of 80%). It is pushed to the explosion - proof intelligent terminal device equipped by the security management personnel through the MQTT protocol to trigger its vibration and screen high - brightness display. For example, in the monitoring scenario of a liquor factory, when it is recognized that a person is smoking in the inflammable and explosive area, the explosion - proof audible and visual alarm at the top of the area is immediately activated, and an alarm message including the facial features of the person (after blurring) and the plane location map of the device area is sent to the handheld terminals of multiple nearest security officers, and at the same time, the incident location is marked with a red - circle flash on the large - screen map in the central control room.

[0133] Step S152: When the behavior type is the second - type abnormal behavior, a second - level warning signal is generated and a behavior video clip and recommended handling measures are sent to the terminal device.

[0134] The second type of abnormal behavior is, for example, a predefined abnormal behavior category with potential risks but without the need for immediate intervention. For example, the operator does not wear a safety helmet correctly or an unlit cigarette holding action is detected in the no-smoking area. The generation of the secondary warning signal can trigger an audible and visual alarm mode of a yellow warning light and an intermittent buzzer of 80 decibels. The warning message transmits a 10-second H.265-encoded behavior video clip (resolution 1920×1080, frame rate 25fps) to the mobile terminal of the relevant responsible person through the factory area's 5G private network. The video clip is attached with structured metadata including the start time of the abnormal behavior, the duration, and the associated device status parameters.

[0135] Step S153: When the behavior type is an unknown abnormal category, generate a tertiary warning signal and upload the associated video stream to the cloud analysis platform, and start a multi-expert collaborative annotation mechanism to update the abnormal behavior feature library.

[0136] The unknown abnormal category is, for example, a behavior type whose matching similarity with all benchmark feature vectors in the abnormal behavior feature library is lower than a dynamic threshold (such as 0.75) and the time dimension verification fails. For example, a behavior pattern that appears for the first time or a complex violation operation by an untrained person. For example, the tertiary warning signal activates a blue rotating light and a low-frequency buzzer of 60 decibels to avoid excessive interference with normal operations. At the same time, upload the complete video stream (HEVC-encoded, with depth information) of 30 seconds before and after the abnormal behavior and a multi-frame behavior feature set to the distributed object storage bucket of the cloud analysis platform. A multi-expert collaborative annotation mechanism can be adopted. Create an annotation task through the RESTful API interface, divide the video stream into segments at 5-second intervals and randomly allocate them to the terminal devices of at least multiple certified experts. Independently annotate the abnormal behavior type, severity level, and key feature dimensions through a dedicated annotation tool.

[0137] As an implementation, in the step S153, starting the multi-expert collaborative annotation mechanism to update the abnormal behavior feature library may include the following sub-steps: Step S1531: Divide the video stream data corresponding to the unknown abnormal category into multiple annotation tasks to be annotated and allocate them to multiple expert terminals for independent annotation.

[0138] The video stream segmentation adopts an adaptive cutting algorithm based on behavior integrity, and preferentially cuts at the start frame of the action (such as the mutation point of the arm speed) and the end frame (such as the action stop timestamp) to ensure that each annotation task to be annotated contains a complete abnormal behavior cycle. The task allocation strategy adopts a load balancing algorithm, considering the current workload of the expert's tasks to be processed, the matching degree of the professional field, and the historical annotation accuracy (such as abnormal situations related to equipment are preferentially allocated to maintenance experts). Each task is allocated to at least 3 experts for back-to-back annotation.

[0139] Step S1532: Receive the annotation results returned by each expert terminal, where the annotation results include behavior type suggestions and feature correction parameters.

[0140] The annotation results can be encapsulated in JSON format, including abnormal behavior types (supporting custom input), confidence scores (levels 1-5), suggestions for correcting key feature dimensions, and severity level assessments. The system verifies the identity of experts through digital signatures and performs a hash check on the submission timestamp to prevent tampering.

[0141] Step S1533: Use the majority voting algorithm to perform consistency verification on the annotation results. If the behavior types suggested by more than a preset proportion of experts are the same, add the behavior types to the new category in the abnormal behavior feature library.

[0142] For example, the majority voting algorithm can set a threshold that requires at least 66.7% of the experts to reach a consensus. When there is a tie, a similarity comparison of the feature correction parameters is triggered. The process of adding a new category to the library includes generating a new benchmark feature vector (taking the mean of the feature suggestions of the experts), setting an initial priority weight (according to the severity level), and updating the fully connected dimension of the classifier layer.

[0143] Step S1534: If the annotation results do not reach a consensus, initiate a secondary annotation process until an annotation result that meets the consistency condition is generated.

[0144] For example, the secondary annotation process expands the expert pool to more people. The system performs a significance analysis on the controversial feature dimensions and automatically generates a comparison view (such as overlaying and displaying the tremor frequency ranges marked by different experts) to assist in decision-making.

[0145] As an independently implementable embodiment, the method provided by the embodiments of the present invention further includes the step of performing incremental learning on the abnormal behavior recognition model, which specifically may include: Step S160: Regularly collect the early warning processing records and manual review results fed back by the target terminal device, and screen out the target case data where the model recognition is incorrect or the confidence level is lower than the threshold.

[0146] The target terminal device is, for example, a mobile inspection terminal equipped for factory area safety management personnel, an operation console in the central control room, and a cloud alarm management platform. The early warning processing records fed back by it are stored in the form of structured logs, including fields such as alarm trigger time, operator number, disposal measures, and review conclusion. The manual review results are entered through the annotation interface of the terminal device. The engineer checks the correct category on the review interface and submits the corrected label. The screening conditions for cases where the model recognition is incorrect include: cases where the manual review determines a false alarm or missed alarm after the early warning signal is triggered, cases where the probability output by the classifier is lower than the dynamic confidence threshold (such as 0.65), and cases where the same abnormal category is matched three times in a row but the actual behavior does not match.

[0147] Step S170: Re-extract features from the video stream in the target case data to generate an incremental training dataset, and add corrected behavior type labels to each data sample.

[0148] The feature re-extraction process calls the same preprocessing pipeline as in the initial training phase, including calling a skeletal key point detection network to extract a 30-frame-per-second joint three-dimensional coordinate sequence, an object contour segmentation network to generate a pixel-level mask, and an action trajectory tracking algorithm to calculate displacement vectors. The construction of the incremental training dataset requires time dimension expansion of the original video stream. For example, the video stream from 10 seconds before the behavior occurs to 5 seconds after the fault is resolved in the target case is re-segmented into 15-second segments containing the complete behavior context, and a 256-dimensional multimodal feature vector is extracted from each segment. The corrected behavior type labels are updated according to the results of manual review. In the scenario of pipeline leakage detection in a liquor production plant, for 10 target cases where the original model misjudged "flange seal failure" as "pressure gauge failure", the system re-extracts features from the infrared thermal imaging video 2 minutes before and after the leakage occurs to generate an incremental dataset containing temperature gradient changes, steam diffusion profiles, and emergency actions of maintenance personnel. Each sample is labeled with the new "seal failure" category and assigned a unique hash identifier.

[0149] Step S180: Freeze the parameters of the multimodal feature encoding layer in the abnormal behavior recognition model, and only fine-tune the weights of the classifier layer.

[0150] Parameter freezing is achieved by setting the requires_grad attribute of the feature encoding layer to False, ensuring that the weights of the temporal convolutional network, attention mechanism module, and recurrent neural network components remain unchanged during incremental training. The fine-tuning training of the classifier layer uses an SGD optimizer with a small learning rate (such as 1e-5), and only updates the weight matrix and bias terms of the fully connected layer to prevent damage to the learned cross-modal feature association patterns.

[0151] Step S190: During the fine-tuning training process, use the elastic weight consolidation algorithm to retain the important parameters of the original classifier layer, avoiding catastrophic forgetting of the learned abnormal behavior categories.

[0152] The execution logic of the elastic weight consolidation algorithm is to calculate the Fisher information matrix of each parameter in the original classifier layer on the historical training data, identify the weight channels crucial for the discrimination of the mastered categories, and impose regularization constraints to protect the numerical change range of these channels during parameter update. In specific implementation, the algorithm calculates the importance score Ω i for each weight parameter w i = 1 / (F i + ε), where F iis the Fisher information, ε = 1e-8 to prevent division-by-zero errors, and λ∑Ω is added to the loss function. i (w i − w i old ) 2 regularization term (λ = 0.5).

[0153] Based on the foregoing embodiments, an abnormal behavior recognition and monitoring device is provided in an embodiment of the present invention. Each unit included in the device, as well as each module included in each unit, can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0154] Figure 2 is a schematic structural diagram of the composition of an abnormal behavior recognition and monitoring device provided in an embodiment of the present invention. As Figure 2 shown, the abnormal behavior recognition and monitoring device 200 includes: A data acquisition module 210, configured to acquire continuous video stream data collected in real time by a plurality of monitoring devices in a target factory area, where the continuous video stream data includes a dynamic behavior sequence of at least one target object; A feature extraction module 220, configured to perform time dimension segmentation and space dimension preprocessing on the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence, where the multi-frame behavior feature set includes the skeletal key point features, environmental associated object contour features, and action trajectory features of the target object in each video frame; A feature fusion module 230, configured to input the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, and perform cross-modal correlation fusion on the skeletal key point features, environmental associated object contour features, and action trajectory features through a multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector; A behavior recognition module 240, configured to perform similarity matching based on the target behavior feature vector and a preset abnormal behavior feature library to determine whether the behavior type of the target object belongs to a predefined abnormal behavior category; An abnormal warning module 250, configured to, if it is determined that it belongs to the abnormal behavior category, trigger a corresponding warning signal according to the behavior type, and send the warning signal and the associated video frame to a target terminal device for real-time alarm display.

[0155] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided by the embodiments of the present invention can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding.

[0156] It should be noted that in the embodiments of the present invention, if the above-mentioned AI-based abnormal behavior recognition and monitoring method in the factory area is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present invention are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0157] The embodiments of the present invention provide a computer system, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0158] The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0159] The embodiments of the present invention provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0160] An embodiment of the present invention provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. This computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0161] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present invention, please refer to the descriptions of the method embodiments of the present invention for understanding.

[0162] Figure 3 The following is a schematic diagram of the hardware entity of a computer system provided by an embodiment of the present invention. As Figure 3 shown, the hardware entity of the computer system 1000 includes a processor 1001 and a memory 1002. Among them, the memory 1002 stores a computer program that can run on the processor 1001. When the processor 1001 executes the program, the steps in the method of any of the above embodiments are implemented.

[0163] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be electrical, mechanical, or other forms.

[0164] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0165] In addition, each functional unit in the embodiments of the present invention may all be integrated into one processing unit, or each unit may be separately regarded as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated unit may be implemented in the form of hardware, or in the form of a hardware plus a software functional unit.

[0166] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.

[0167] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related art, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks and other various media that can store program codes.

[0168] The above is only the implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A method for identifying and monitoring abnormal behavior in a factory based on AI, characterized in that: The method comprises: Acquire continuous video stream data collected in real time by multiple monitoring devices in a target factory area, wherein the continuous video stream data includes a dynamic behavior sequence of at least one target object; The continuous video stream data is segmented in time dimension and preprocessed in space dimension to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence, wherein the multi-frame behavior feature set includes skeleton key point features, environment-related object contour features, and motion trajectory features of the target object in each video frame; Inputting the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, cross-modal correlation fusion of the skeleton key point features, environment-related object contour features and motion trajectory features is performed through a multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector; Based on the similarity matching between the target behavior feature vector and a preset abnormal behavior feature library, determining whether the behavior type of the target object belongs to a predefined abnormal behavior category; If it is determined to belong to the abnormal behavior category, a corresponding warning signal is triggered according to the behavior type, and the warning signal and the associated video frame are sent to the target terminal device for real-time alarm display.

2. The method according to claim 1, characterized in that The step of performing time dimension segmentation and space dimension preprocessing on the continuous video stream data to obtain a multi-frame behavior feature set corresponding to the dynamic behavior sequence includes: Dividing the continuous video stream data into a plurality of video segments according to a preset time interval, each video segment comprising a fixed number of continuous video frames; For each video frame, a pre-trained skeleton key point detection network is called to extract the skeleton key point features of the target object, wherein the skeleton key point features include the three-dimensional coordinates of each joint point and the motion speed parameters between adjacent joint points; Synchronously calling an object contour segmentation network to perform contour recognition on an object that has spatial interaction with the target object in the video frame, and extracting contour features of the environment-related object, wherein the contour features of the environment-related object include an object shape descriptor and a relative distance parameter to the target object; Based on the time sequence relationship of the video clips, the position change of the target object in adjacent video frames is tracked to generate the motion trajectory feature, wherein the motion trajectory feature includes a displacement direction sequence and an acceleration change curve; The skeleton key point features, environment-related object contour features and motion trajectory features corresponding to the same video frame are combined to form the multi-frame behavior feature set.

3. The method according to claim 2, characterized in that Before calling the pre-trained skeleton key point detection network to extract the skeleton key point features of the target object, the method further includes the step of training the skeleton key point detection network: Obtain a historical surveillance video data set in a factory scene, and manually annotate each video frame in the historical surveillance video data set, wherein the annotated content includes the joint point position and joint point motion label of the target object; Constructing an initial key point detection network, wherein the initial key point detection network includes a feature pyramid module and a three-dimensional coordinate regression module; Input the video frames in the historical surveillance video data set into the initial key point detection network, extract multi-scale spatial features through the feature pyramid module, and predict the three-dimensional coordinates and motion speed parameters of each joint point through the three-dimensional coordinate regression module; Calculating the Euclidean distance loss value between the predicted coordinates and the annotated coordinates, and combining the difference loss value between the motion speed parameter and the joint point motion label to generate a total training loss value; The initial key point detection network is back-propagated and optimized based on the total training loss value until the total training loss value converges to a preset threshold value, thereby obtaining the pre-trained skeleton key point detection network.

4. The method according to claim 1, characterized in that: The step of inputting the multi-frame behavior feature set into a pre-trained abnormal behavior recognition model, and cross-modally associating and fusing the skeleton key point features, the environment-related object contour features, and the motion trajectory features through a multi-modal feature encoding layer in the abnormal behavior recognition model to generate a target behavior feature vector includes: Inputting the skeleton key point features into the first feature encoding sub-network in the multimodal feature encoding layer to generate a first encoding feature vector, wherein the first feature encoding sub-network uses a temporal convolution structure to capture the joint point motion pattern; Inputting the contour feature of the environment-associated object into a second feature encoding subnetwork in the multimodal feature encoding layer to generate a second encoding feature vector, wherein the second feature encoding subnetwork uses an attention mechanism to focus on the contour of the object with the highest interaction frequency with the target object; Inputting the motion trajectory feature into a third feature encoding subnetwork in the multimodal feature encoding layer to generate a third encoding feature vector, wherein the third feature encoding subnetwork uses a recurrent neural network structure to analyze the continuity of a displacement direction sequence; The first encoded feature vector, the second encoded feature vector and the third encoded feature vector are concatenated, and dimensionality reduction and fusion are performed through a fully connected layer to obtain the target behavior feature vector.

5. The method according to claim 4, characterized in that The training process of the abnormal behavior recognition model includes: Collect historical abnormal behavior video samples and normal behavior video samples in the factory area, extract the multi-frame behavior feature set for each video sample, and mark the corresponding abnormal behavior category label; Constructing an initial abnormal behavior recognition model, wherein the initial abnormal behavior recognition model includes the multimodal feature encoding layer and the classifier layer; Inputting the multi-frame behavior feature set into the initial abnormal behavior recognition model, outputting the target behavior feature vector through the multimodal feature encoding layer, and calculating the matching probability between the target behavior feature vector and the feature template corresponding to each abnormal behavior category through the classifier layer; Based on the cross entropy loss value between the matching probability and the abnormal behavior category label, updating the parameters of the initial abnormal behavior recognition model; An adversarial training strategy is introduced to add noise perturbation data to the multi-frame behavior feature set, and the classification consistency loss value of the perturbed feature set and the original feature set is calculated, and the cross entropy loss value is combined for joint optimization until the model converges.

6. The method according to claim 1, characterized in that The determining whether the behavior type of the target object belongs to a predefined abnormal behavior category based on similarity matching between the target behavior feature vector and a preset abnormal behavior feature library includes: Reading benchmark feature vectors of all predefined abnormal behavior categories from the abnormal behavior feature library, wherein the benchmark feature vectors are obtained by centralizing features extracted from historical abnormal behavior data using a clustering algorithm; Calculate the cosine similarity between the target behavior feature vector and each benchmark feature vector, and select the abnormal behavior categories corresponding to the first K benchmark feature vectors with the highest similarity as candidate categories; Performing a time dimension alignment check on the benchmark feature vector of the candidate category to determine whether the change trend of the target behavior feature vector in a continuous time window is consistent with the behavior pattern of the candidate category; If the verification passes, the abnormal behavior category with the highest similarity among the candidate categories is determined as the behavior type of the target object; otherwise, the behavior type is marked as an unknown abnormal category and a manual review process is triggered.

7. The method according to claim 6, characterized in that The method for constructing the abnormal behavior feature library includes: Collect videos of abnormal behavior cases that have occurred in the target factory area, including illegal smoking, illegal use of electronic devices, and intrusion into dangerous areas; Each abnormal behavior case video is segmented, and a multi-frame behavior feature set within each time period is extracted, and the corresponding abnormal behavior category and severity level are marked; A hierarchical clustering algorithm is used to group the multi-frame behavior feature sets of the same abnormal behavior category, and the feature mean of each group is calculated as the reference feature vector; A priority weight is assigned to each reference feature vector according to the severity level, and the reference feature vector and its priority weight are stored in the abnormal behavior feature library.

8. The method according to claim 1, characterized in that If it is determined that the abnormal behavior category is involved, triggering a corresponding warning signal according to the behavior type includes: When the behavior type is the first type of abnormal behavior, a first-level warning signal is generated and the sound and light alarm in the target factory area is controlled to be activated, and an emergency warning message containing a screenshot of the behavior and location information is sent to the terminal device of the security manager; When the behavior type is the second type of abnormal behavior, a second-level warning signal is generated and a behavior video clip and suggested disposal measures are sent to the terminal device; When the behavior type is an unknown abnormal category, a third-level warning signal is generated and the associated video stream is uploaded to the cloud analysis platform; The video stream data corresponding to the unknown anomaly category is divided into a plurality of tasks to be labeled, and the tasks are assigned to a plurality of expert terminals for independent labeling; Receiving the annotation results returned by each expert terminal, wherein the annotation results include behavior type suggestions and feature correction parameters; A majority voting algorithm is used to perform consistency check on the annotation results. If more than a preset proportion of experts suggest the same behavior type, the behavior type is added to the newly added category of the abnormal behavior feature library; If the annotation results do not reach consistency, a secondary annotation process is initiated until an annotation result that meets the consistency condition is generated.

9. The method according to claim 1, characterized in that: The method further comprises the step of performing incremental learning on the abnormal behavior recognition model: Regularly collect the warning processing records and manual review results fed back by the target terminal devices, and filter out target case data with model recognition errors or confidence levels below a threshold; Re-extracting features of the video stream in the target case data to generate an incremental training data set, and adding a corrected behavior type label to each data sample; Freeze the parameters of the multimodal feature encoding layer in the abnormal behavior recognition model, and only fine-tune the weights of the classifier layer; During the fine-tuning training process, the elastic weight merging algorithm is used to retain the important parameters of the original classifier layer.

10. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN115205943A

  • Human body action recognition method and device, storage medium and vehicle

    CN115578720A

  • Energy station abnormal behavior early warning method based on multi-source image recognition and deep learning

    CN116758475A

  • Double-flow human body behavior recognition method and system based on spatio-temporal information

    CN118430063A

  • Human body abnormal action recognition method and device, computer equipment and medium

    CN119028011A

Cited By

  • Construction site real-time monitoring system and method based on unmanned aerial vehicle technology

    CN120298977A

  • Real-time monitoring and abnormity early warning shooting training safety system and method

    CN120298979A

  • Cigarette station personnel violation identification method and system based on deep learning

    CN120318916A

  • Public transport passenger flow monitoring method and system based on human body detection

    CN120451915A

  • Tunnel facility abnormity judgment method and system based on environmental feature fusion

    CN120510578A