Dynamic behavior compliance verification method and device based on AI multi-mode recognition, electronic equipment and storage medium
By using AI multi-modal recognition technology, combined with OpenPose, CNN-LSTM and DeepLabv3+ networks, the system can identify staff leaving their posts or sleeping on duty and abnormal safety passages in power business halls. This solves the problems of large blind spots and high identification difficulty in existing supervision, and enables real-time risk detection and efficient compliance verification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-03-13
AI Technical Summary
The existing supervision of power service halls suffers from problems such as large blind spots, high difficulty in identification, and poor real-time performance. The traditional combination of video surveillance and manual inspection cannot achieve comprehensive supervision of the service process. Furthermore, the existing AI-based new cameras are costly and lack multi-mode recognition strategies, which cannot meet the unique service compliance supervision needs of the power industry.
A dynamic behavior compliance verification method based on AI multi-modal recognition is adopted. Through the identification models of personnel leaving their posts and sleeping on duty and the identification models of abnormal safety passages, the OpenPose algorithm, CNN-LSTM multi-classification model and DeepLabv3+ network are used in combination with the YOLO algorithm to perform real-time video data analysis, identify personnel leaving their posts and sleeping on duty and abnormal safety passages, and generate corresponding alarms.
It enables precise verification of the service process in the business hall, timely detection of security risks, improved operational efficiency, reduced manual intervention, and enhanced efficiency and accuracy of compliance checks, adapting to the special scenarios of power business halls.
Smart Images

Figure CN121661378A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of behavior verification technology, specifically a dynamic behavior compliance verification method, apparatus, electronic device, and storage medium based on AI multi-modal recognition. Background Technology
[0002] Against the backdrop of the power industry's digital transformation and the increasing demand for intelligent management of service halls, the compliance supervision of service halls, as a core link in the operation and management of power companies, directly impacts service quality, customer satisfaction, and corporate image through its supervisory efficiency, risk identification capabilities, and response timeliness. However, current service hall supervision faces numerous technical challenges. Existing supervision methods largely rely on a combination of traditional video surveillance and manual inspections, lacking intelligent analysis tools for audio and video data from service sites. Manual monitoring suffers from large blind spots and high identification difficulty, and can only provide retrospective review, failing to provide real-time alerts and thus failing to meet the needs of refined service hall management. Furthermore, existing systems operate in isolation, creating a "digital divide" with intelligent management platforms, hindering data interoperability and unified management, resulting in low supervisory efficiency.
[0003] Currently, although new cameras with some AI recognition capabilities exist, they are poorly adapted to the specific environment of power service halls. Specifically: (1) The cost of intelligent transformation of traditional cameras is high, and the general AI recognition function is difficult to meet the service compliance supervision requirements unique to the power industry; (2) The existing application of “audio and video + AI” technology in power business halls lacks a multi-mode recognition strategy for service personnel behavior norms, abnormal customer behavior and security risks, and cannot achieve comprehensive supervision of the service process; (3) It is not possible to verify the compliance of personnel dynamic behavior from multiple perspectives such as behavior pattern recognition, emotion recognition, and behavior continuity recognition, so as to discover risks in a timely manner and achieve refined management of the business hall. Summary of the Invention
[0004] This invention provides a dynamic behavior compliance verification method, device, electronic device, and storage medium based on AI multi-mode recognition, which overcomes the shortcomings of the prior art. It can effectively solve the problems of large management blind spots, high recognition difficulty, and poor real-time performance in the existing traditional video surveillance combined with manual inspection, as well as the high cost of existing inspection methods based on AI recognition of new cameras.
[0005] One of the technical solutions of this invention is achieved through the following measures: a dynamic behavior compliance verification method based on AI multi-modal recognition, comprising: Acquire real-time video data; Real-time video data is input into the behavior compliance verification model to obtain the behavior compliance verification results. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. Real-time video data is input into the personnel leaving their post and sleeping on duty identification model to extract the corresponding posture spatiotemporal sequence features and compare them with normal behavior to identify personnel leaving their post and sleeping on duty. Real-time video data is input into the safety passage anomaly identification model to obtain the safety passage anomaly identification results. The safety passage anomaly identification model is constructed using object detection algorithms and DeepLabv3+ networks.
[0006] The following are further optimizations and / or improvements to the above-mentioned technical solution: The above-mentioned real-time video data is input into the personnel off-duty / sleeping detection model. The corresponding spatiotemporal sequence features of posture are extracted and compared with normal behavior to identify personnel off-duty / sleeping behavior, including: Acquire real-time video data and perform preprocessing, including converting it to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV. The preprocessed real-time video data is input into the pose feature estimation model to extract the coordinates of the human skeleton joints in each frame of the image and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. The spatiotemporal sequence features of the person's posture are input into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model with several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. Based on the posture classification results, the spatiotemporal sequence features of the corresponding normal on-duty behavior postures are retrieved from the pattern library. The dynamic time warping algorithm is used to perform similarity analysis between these features and the spatiotemporal sequence features of the personnel postures to obtain multiple similarity values. If the similarity value is lower than the similarity threshold for a set threshold period, it is determined that the person has been absent from their post and sleeping on duty.
[0007] The above also includes continuously collecting real-time video data, determining the initial alarm time when the behavior of leaving the post and sleeping on duty is determined, and determining the alarm cancellation time when the behavior of leaving the post and sleeping on duty is determined. The difference between the initial alarm time and the alarm cancellation time is the time of leaving the post. If the time of leaving the post exceeds the time of leaving the post threshold, an alarm for exceeding the time limit of leaving the post will be issued.
[0008] The above-mentioned input real-time video data is fed into the security channel anomaly detection model to obtain security channel anomaly detection results, including: YOLO is used to quickly scan each frame of real-time video data globally to obtain the target bounding box of suspicious objects in the real-time video data. Each frame of real-time video data is input into the segmentation mask acquisition model to obtain a segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization sub-model. The region segmentation sub-model is trained on the DeepLabv3+ network using several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization process includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter out false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The target bounding box is superimposed on the segmentation mask to analyze the relative positional relationship between the boundary of the target bounding box and the key semantic region; The security channel anomaly result of a certain frame image is determined based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; Repeat the above steps until the set threshold is met in the continuous abnormal state of the security passage, and generate a security passage abnormal alarm.
[0009] The aforementioned first anomaly identification rule includes: Intrusion into a restricted area: If a person is detected and their target bounding box is located in or significantly crosses the area, it is considered an anomaly; Unauthorized stay / piling: If a large abandoned object is detected and its target bounding box is completely located within the core area of the safe passage, and the stay time exceeds the threshold, it is considered abnormal.
[0010] Structural damage: If no specific object is detected, but the segmentation mask shows a large area missing from the channel body area, a severely distorted shape, or intrusion by a non-channel category, it is judged as an anomaly.
[0011] The second technical solution of the present invention is achieved through the following measures: a dynamic behavior compliance verification device based on AI multi-mode recognition, comprising: The video data acquisition unit acquires real-time video data. The compliance verification unit inputs real-time video data into the behavior compliance verification model to obtain the behavior compliance verification result. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. The personnel leaving their post and sleeping on duty identification model is input into the real-time video data, extracts the corresponding posture spatiotemporal sequence features, and compares them with normal behavior to identify personnel leaving their post and sleeping on duty. The safety passage anomaly identification model is input into the real-time video data to obtain the safety passage anomaly identification result. The safety passage anomaly identification model is constructed using object detection algorithms and DeepLabv3+ network.
[0012] The following are further optimizations and / or improvements to the above-mentioned technical solution: The aforementioned compliance verification unit includes: The personnel absence / sleep detection module inputs real-time video data into the personnel absence / sleep detection model, extracts the corresponding spatiotemporal sequence features of posture, and compares them with normal behavior to identify personnel absence / sleep behavior, including: Acquire real-time video data and perform preprocessing, including converting it to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV. The preprocessed real-time video data is input into the pose feature estimation model to extract the coordinates of the human skeleton joints in each frame of the image and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. The spatiotemporal sequence features of the person's posture are input into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model with several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. Based on the posture classification results, the spatiotemporal sequence features of the corresponding normal on-duty behavior postures are retrieved from the pattern library. The dynamic time warping algorithm is used to perform similarity analysis between these features and the spatiotemporal sequence features of the personnel postures to obtain multiple similarity values. If the similarity value is lower than the similarity threshold within a set threshold time period, it is determined that the person has been absent from their post and sleeping on duty. The security passage anomaly detection module takes real-time video data as input to the security passage anomaly detection model and obtains the security passage anomaly detection results, including: YOLO is used to quickly scan each frame of real-time video data globally to obtain the target bounding box of suspicious objects in the real-time video data. Each frame of real-time video data is input into the segmentation mask acquisition model to obtain a segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization sub-model. The region segmentation sub-model is trained on the DeepLabv3+ network using several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization process includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter out false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The target bounding box is superimposed on the segmentation mask to analyze the relative positional relationship between the boundary of the target bounding box and the key semantic region; The security channel anomaly result of a certain frame image is determined based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; Repeat the above steps until the set threshold is met in the continuous abnormal state of the security passage, and generate a security passage abnormal alarm.
[0013] The third technical solution of the present invention is achieved through the following measures: an electronic device, including a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the steps in the dynamic behavior compliance verification method based on AI multi-modal recognition.
[0014] The fourth technical solution of the present invention is achieved through the following measures: a storage medium storing a computer program that can be read by a computer, the computer program being configured to execute the steps in the dynamic behavior compliance verification method based on AI multi-mode recognition at runtime.
[0015] This invention requires no modification to existing cameras; it only needs to access real-time video data collected by existing cameras to verify behavioral compliance in two directions: personnel leaving their posts / sleeping on duty and safety passage anomaly detection. This enables accurate verification of behavioral compliance during service processes in the business hall, timely detection and handling of security risks, improved operational efficiency, reduced manual intervention, and enhanced efficiency, accuracy, and intelligence of compliance checks in the business hall. Specifically, the personnel leaving their posts / sleeping on duty detection incorporates the OpenPose algorithm, CNN-LSTM multi-classification model, and dynamic time warping algorithm to automatically identify personnel leaving their posts / sleeping on duty using real-time video data. The safety passage anomaly detection incorporates the YOLO algorithm and DeepLabv3+ network to accurately segment the safety passage area and detect target objects, thereby achieving accurate identification of safety passage anomalies. Attached Figure Description
[0016] Appendix Figure 1This is a schematic diagram of the dynamic behavior compliance verification method provided in an embodiment of the present invention.
[0017] Appendix Figure 2 This is a schematic diagram of the process for identifying personnel leaving their posts and sleeping on duty, as provided in an embodiment of the present invention.
[0018] Appendix Figure 3 This is a schematic diagram of the security channel anomaly identification method provided in an embodiment of the present invention.
[0019] Appendix Figure 4 This is a schematic diagram of the structure of the dynamic behavior compliance verification device provided in an embodiment of the present invention. Detailed Implementation
[0020] The present invention is not limited to the following embodiments, and the specific implementation can be determined according to the technical solution of the present invention and the actual situation.
[0021] Those skilled in the art will understand that, unless specifically stated otherwise, in the embodiments of the present invention, a "module" or "unit" refers to a computer program or part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0022] In addition, in the embodiments of the present invention, "multiple" refers to two or more, and "first" and "second" are used to distinguish descriptions and should not be construed as implying relative importance.
[0023] This invention provides a method, apparatus, electronic device, and storage medium for dynamic behavioral compliance verification based on AI multi-mode recognition. This AI multi-mode recognition-based dynamic behavioral compliance verification apparatus can be integrated into a computer device, which can be a server, a terminal, or other similar device; it can also be executed jointly by a terminal and a server. These examples should not be construed as limiting the invention.
[0024] The aforementioned terminals may include mobile phones, wearable smart devices, tablet computers, laptops, personal computers (PCs), and in-vehicle computers, etc., and this invention does not limit them. This invention also does not limit the number of terminal devices.
[0025] The aforementioned server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This invention does not limit these features.
[0026] For example, computer equipment acquires real-time video data; the real-time video data is input into a behavior compliance verification model to obtain the behavior compliance verification result. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. The real-time video data is input into the personnel leaving their post and sleeping on duty identification model to extract the corresponding posture spatiotemporal sequence features and compare them with normal behavior to identify personnel leaving their post and sleeping on duty. The real-time video data is input into the safety passage anomaly identification model to obtain the safety passage anomaly identification result. The safety passage anomaly identification model is constructed using an object detection algorithm and a DeepLabv3+ network.
[0027] Based on this, the technical solution of the present invention will be described and explained below with reference to several examples.
[0028] Example 1: As shown in the attached document Figure 1 As shown in the figure, this invention discloses a dynamic behavior compliance verification method based on AI multi-modal recognition, including: Step S110: Obtain real-time video data; The aforementioned acquisition of real-time video data can fully utilize the existing network cameras in the business hall. These network cameras, based on image sensor technology, continuously collect real-time video data within the business hall at a predetermined frame rate and resolution. The real-time video data is then denoised using a combination of median filtering and bilateral filtering algorithms to effectively remove noise points from the image. It is then encapsulated using a standard encoding format and transmitted stably via the TCP / IP protocol for subsequent dynamic behavior compliance verification.
[0029] Step S120: Input real-time video data into the behavior compliance verification model to obtain the behavior compliance verification result. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. Input real-time video data into the personnel leaving their post and sleeping on duty identification model, extract the corresponding posture spatiotemporal sequence features, and compare them with normal behavior to identify personnel leaving their post and sleeping on duty behavior. Input real-time video data into the safety passage anomaly identification model to obtain the safety passage anomaly identification result. The safety passage anomaly identification model is constructed using an object detection algorithm and a DeepLabv3+ network.
[0030] This invention discloses a dynamic behavior compliance verification method based on AI multi-modal recognition. It verifies behavior compliance from two directions: identifying personnel leaving their posts or sleeping on duty and identifying abnormalities in security passages. This enables accurate verification of behavior compliance during business hall services, timely detection and handling of security risks, improved operational efficiency, reduced manual intervention, and enhanced efficiency, accuracy, and intelligence of business hall compliance checks.
[0031] Example 2: As shown in the attached document Figure 2 As shown, this embodiment of the invention is a further optimization of the above embodiment, wherein real-time video data is input into the personnel leaving their post and sleeping on duty recognition model, the corresponding posture spatiotemporal sequence features are extracted, and compared with normal behavior to identify personnel leaving their post and sleeping on duty behavior, including: Step S210: Acquire real-time video data and perform preprocessing, including converting to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV.
[0032] Step S220: Input the preprocessed real-time video data into the pose feature estimation model, extract the coordinates of the human skeleton joints of each frame image, and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. OpenPose, a mainstream real-time multi-person pose estimation framework, employs a two-stream convolutional neural network architecture. After extracting image features through a VGG19 network, it branches into two parallel processing streams: one generates Part Confidence Maps (PCMs), and the other generates Part Affinity Fields (PAFs). For video sequences, frame-by-frame processing is required, and PCM coordinates are aggregated along the time axis to form spatiotemporal sequence data. This data is then used to train the OpenPose human pose estimation model, resulting in a pose feature estimation model. This model extracts the coordinates of the human skeleton joints in each frame and aggregates all these coordinates along the time axis to form the spatiotemporal sequence features of the human pose. The OpenPose human pose estimation model supports multi-PCM detection, enabling simultaneous detection of 130 PCMs across the body, hands, and face. It also uses PAFs to encode limb connections, addressing occlusion and complex background issues.
[0033] Step S230: Input the spatiotemporal sequence features of the person's posture into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model using several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. The construction process of the above pose classification model includes: (1) The OpenPose human pose estimation model is used as the pre-training framework; (2) Define abnormal behaviors, such as sleeping on duty, using a mobile phone, smoking, etc.; (3) Dataset collection and annotation; Historical video surveillance data from multiple service halls was collected, including normal and abnormal behavior scenarios under different time periods, angles, and lighting conditions. The VATIC tool was used to annotate the start and end keyframes of the same user behavior. The OpenPose human pose estimation model was used to process the video, obtaining the spatiotemporal sequence features of the personnel's poses for each historical video surveillance data point. The annotation tool directly read these spatiotemporal sequence features and displayed them as a skeleton animation. Annotators then performed behavior annotation based on the skeleton animation, resulting in several samples. Each sample includes the spatiotemporal sequence features of the personnel's poses and the corresponding pose classification result. Here, the pose classification result is the behavior annotation, and the behavior includes normal behavior, abnormal behavior, and others.
[0034] (4) A pose classification model is generated by training a CNN-LSTM multi-classification model; Sample format: Defined as a tensor of (T, N, 3), where T is the time step (number of frames), N is the number of skeletal joints of the person, and 3 represents (x, y, confidence) of each skeletal joint; Spatial feature extraction (CNN): Treat the human skeleton joint points (N points) in each frame as an "image", but in fact they are two-dimensional coordinate points. Use 1x1 convolution or fully connected layers to extract spatial features. That is, flatten the coordinates of the human skeleton joint points in each frame into a vector of length 2*N (ignoring confidence) or 3*N, and then extract spatial features through fully connected layers. Temporal Feature Extraction (LSTM): The spatial features (a vector) of each frame are input into the LSTM in temporal order, and the LSTM captures the temporal dynamics; Output layer: Connect a fully connected layer and a softmax layer to the output of the LSTM (either the output of the last time step or the output of the time distribution) to output the probability of each behavior category, and obtain the pose classification result based on the probability.
[0035] Step S240: Based on the posture classification result, call the spatiotemporal sequence features of the corresponding normal on-duty behavior posture in the pattern library, and use the dynamic time warping algorithm to perform similarity analysis with the spatiotemporal sequence features of the personnel posture to obtain multiple similarity values (i.e., a similarity value will be generated for each frame image). Step S250: If the similarity value is lower than the similarity threshold for a set threshold time period, then it is determined that the behavior of leaving the post and sleeping on duty has occurred. That is, if the similarity value is lower than the similarity threshold for a period of time that is greater than or equal to the set threshold time period, then it is determined that the behavior of leaving the post and sleeping on duty has occurred.
[0036] Furthermore, real-time video data is continuously collected. When the behavior of leaving the post and sleeping on duty is determined, the initial alarm time is determined. When the behavior of leaving the post and sleeping on duty is determined, the alarm cancellation time is determined. The difference between the initial alarm time and the alarm cancellation time is the time of leaving the post. If the time of leaving the post exceeds the time of leaving the post, an alarm for exceeding the time limit of leaving the post is issued.
[0037] Example 3: As shown in the attached document Figure 3 As shown, this embodiment of the invention is a further optimization of the above embodiment, wherein inputting real-time video data into the secure channel anomaly identification model to obtain secure channel anomaly identification results includes: Step S310: Use YOLO to quickly scan each frame of the real-time video data to obtain the target bounding box of the suspicious target object in the real-time video data. After detecting suspicious objects within the security passage, the system outputs the target bounding box, category label, and confidence level. Suspicious objects include people / animals (appearing in restricted areas or lingering), large abandoned objects / obstacles (boxes, goods, equipment), etc.
[0038] Step S320: Input each frame of real-time video data into the segmentation mask acquisition model to obtain the segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization processing sub-model. The region segmentation sub-model is obtained by training the DeepLabv3+ network with several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization processing sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization processing includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The purpose of the aforementioned region segmentation sub-model is to segment the secure channel region in each frame of the image, perform pixel-level segmentation on the same frame of the image, accurately understand the physical structure and semantic region of the secure channel, including separating the background / non-channel region, the channel body region, etc., the semantic category label of each pixel, and generate a segmentation mask.
[0039] The training process of the region segmentation sub-model includes: (a) Data collection Images of the target scene area under different lighting conditions (day / night), occlusions (objects, pedestrians), and viewing angles (overhead / eye level) were collected, and the ADE20K public dataset was used to assist in training.
[0040] (II) Model Training Transfer learning: Initialize weights: Load the COCO pre-trained model, freeze the early layers of the Backbone, and fine-tune the higher layers and the ASPP module; Weighted cross-entropy loss: To address the issue of a small proportion of the safe passage region (background weight 0.3, safe passage weight 0.7), Dice Loss is chosen to enhance boundary learning.
[0041] Optimizer: Adam (initial learning rate 1e-4), learning rate strategy: cosine annealing.
[0042] Step S330: Overlay the target bounding box onto the segmentation mask and analyze the relative positional relationship (intersection, containment, proximity) between the boundary of the target bounding box and key semantic regions (such as the boundary of the danger zone and the boundary of the safety passage). Step S340: Determine the abnormal result of the safe channel of a certain frame image based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; The aforementioned first anomaly identification rule includes: Intrusion into a restricted area: If a person is detected and their target bounding box is located in or significantly crosses the area, it is considered an anomaly; Unauthorized stay / piling: If a large abandoned object is detected and its target bounding box is completely located within the core area of the safe passage, and the stay time exceeds the threshold, it is considered abnormal.
[0043] Structural damage: If no specific object is detected, but the segmentation mask shows a large area missing from the channel body area, a severely distorted shape, or intrusion by a non-channel category, it is judged as an anomaly.
[0044] Step S350: Repeat the above steps. If the abnormal state of the security channel continues to meet the set threshold, generate a security channel abnormal alarm.
[0045] Example 4: As shown in the appendix Figure 4 As shown in the figure, an embodiment of the present invention discloses a dynamic behavior compliance verification device based on AI multi-mode recognition, comprising: The video data acquisition unit acquires real-time video data. The compliance verification unit inputs real-time video data into the behavior compliance verification model to obtain the behavior compliance verification result. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. The personnel leaving their post and sleeping on duty identification model is input into the real-time video data, extracts the corresponding posture spatiotemporal sequence features, and compares them with normal behavior to identify personnel leaving their post and sleeping on duty. The safety passage anomaly identification model is input into the real-time video data to obtain the safety passage anomaly identification result. The safety passage anomaly identification model is constructed using object detection algorithms and DeepLabv3+ network.
[0046] The device can be an AI-BOX equipped with a high-performance multi-core CPU and a GPU with powerful parallel computing capabilities, and deployed inside the business hall.
[0047] The compliance verification unit includes: The personnel absence / sleep detection module inputs real-time video data into the personnel absence / sleep detection model, extracts the corresponding spatiotemporal sequence features of posture, and compares them with normal behavior to identify personnel absence / sleep behavior, including: Acquire real-time video data and perform preprocessing, including converting it to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV. The preprocessed real-time video data is input into the pose feature estimation model to extract the coordinates of the human skeleton joints in each frame of the image and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. The spatiotemporal sequence features of the person's posture are input into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model with several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. Based on the posture classification results, the spatiotemporal sequence features of the corresponding normal on-duty behavior postures are retrieved from the pattern library. The dynamic time warping algorithm is used to perform similarity analysis between these features and the spatiotemporal sequence features of the personnel postures to obtain multiple similarity values. If the similarity value is lower than the similarity threshold within a set threshold time period, it is determined that the person has been absent from their post and sleeping on duty. The security passage anomaly detection module takes real-time video data as input to the security passage anomaly detection model and obtains the security passage anomaly detection results, including: YOLO is used to quickly scan each frame of real-time video data globally to obtain the target bounding box of suspicious objects in the real-time video data. Each frame of real-time video data is input into the segmentation mask acquisition model to obtain a segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization sub-model. The region segmentation sub-model is trained on the DeepLabv3+ network using several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization process includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter out false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The target bounding box is superimposed on the segmentation mask to analyze the relative positional relationship between the boundary of the target bounding box and the key semantic region; The security channel anomaly result of a certain frame image is determined based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; Repeat the above steps until the set threshold is met in the continuous abnormal state of the security passage, and generate a security passage abnormal alarm.
[0048] Example 5: This embodiment of the invention discloses a storage medium storing a computer program that can be read by a computer. The computer program is configured to execute a dynamic behavior compliance verification method based on AI multi-mode recognition at runtime.
[0049] The aforementioned storage media may include, but are not limited to, USB flash drives, read-only memory, portable hard drives, magnetic disks, optical disks, and other media capable of storing computer programs.
[0050] Example 6: This embodiment of the invention discloses an electronic device, including a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement a dynamic behavior compliance verification method based on AI multi-modal recognition.
[0051] The processor described above can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. It can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The memory can include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, portable hard drives, magnetic disks, or optical disks.
[0052] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0053] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0055] The above content is only a specific embodiment of the present invention, which has strong adaptability and implementation effect. However, the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered within the protection scope of the present invention. Therefore, equivalent changes made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. A dynamic behavior compliance verification method based on AI multi-modal recognition, characterized in that, include: Acquire real-time video data; Real-time video data is input into the behavior compliance verification model to obtain the behavior compliance verification results. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. Real-time video data is input into the personnel leaving their post and sleeping on duty identification model to extract the corresponding posture spatiotemporal sequence features and compare them with normal behavior to identify personnel leaving their post and sleeping on duty. Real-time video data is input into the safety passage anomaly identification model to obtain the safety passage anomaly identification results. The safety passage anomaly identification model is constructed using object detection algorithms and DeepLabv3+ networks.
2. The dynamic behavior compliance verification method based on AI multi-modal recognition according to claim 1, characterized in that, The process involves inputting real-time video data into the personnel absence / sleep detection model, extracting corresponding spatiotemporal posture sequence features, and comparing them with normal behavior to identify personnel absence / sleep behavior, including: Acquire real-time video data and perform preprocessing, including converting it to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV. The preprocessed real-time video data is input into the pose feature estimation model to extract the coordinates of the human skeleton joints in each frame of the image and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. The spatiotemporal sequence features of the person's posture are input into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model with several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. Based on the posture classification results, the spatiotemporal sequence features of the corresponding normal on-duty behavior postures are retrieved from the pattern library. The dynamic time warping algorithm is used to perform similarity analysis between these features and the spatiotemporal sequence features of the personnel postures to obtain multiple similarity values. If the similarity value is lower than the similarity threshold for a set threshold period, it is determined that the person has been absent from their post and sleeping on duty.
3. The dynamic behavior compliance verification method based on AI multi-modal recognition according to claim 2, characterized in that, It also includes continuously collecting real-time video data, determining the initial alarm time when the behavior of leaving the post and sleeping on duty is determined, and determining the alarm cancellation time when the behavior of leaving the post and sleeping on duty is determined. The difference between the initial alarm time and the alarm cancellation time is the time of leaving the post. If the time of leaving the post exceeds the time of leaving the post, an alarm for exceeding the time limit of leaving the post will be issued.
4. The dynamic behavior compliance verification method based on AI multi-modal recognition according to claim 1, 2, or 3, characterized in that, The process of inputting real-time video data into the secure channel anomaly detection model to obtain secure channel anomaly detection results includes: YOLO is used to quickly scan each frame of real-time video data globally to obtain the target bounding box of suspicious objects in the real-time video data. Each frame of real-time video data is input into the segmentation mask acquisition model to obtain a segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization sub-model. The region segmentation sub-model is trained on the DeepLabv3+ network using several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization process includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter out false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The target bounding box is superimposed on the segmentation mask to analyze the relative positional relationship between the boundary of the target bounding box and the key semantic region; The security channel anomaly result of a certain frame image is determined based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; Repeat the above steps until the set threshold is met in the continuous abnormal state of the security passage, and generate a security passage abnormal alarm.
5. The dynamic behavior compliance verification method based on AI multi-modal recognition according to claim 4, characterized in that, The first anomaly identification rule includes: Intrusion into a restricted area: If a person is detected and their target bounding box is located in or significantly crosses the area, it is considered an anomaly; Unauthorized stay / piling: If a large abandoned object is detected and its target bounding box is completely located within the core area of the safe passage, and the stay time exceeds the threshold, it is considered abnormal; Structural damage: If no specific object is detected, but the segmentation mask shows a large area missing from the channel body area, a severely distorted shape, or intrusion by a non-channel category, it is judged as an anomaly.
6. A dynamic behavior compliance verification device based on AI multi-modal recognition, applying the method described in any one of claims 1 to 5, characterized in that, include: The video data acquisition unit acquires real-time video data. The compliance verification unit inputs real-time video data into the behavior compliance verification model to obtain the behavior compliance verification result. The behavior compliance verification model includes a personnel leaving their post and sleeping on duty identification model and a safety passage anomaly identification model. The personnel leaving their post and sleeping on duty identification model is input into the real-time video data, extracts the corresponding posture spatiotemporal sequence features, and compares them with normal behavior to identify personnel leaving their post and sleeping on duty. The safety passage anomaly identification model is input into the real-time video data to obtain the safety passage anomaly identification result. The safety passage anomaly identification model is constructed using object detection algorithms and DeepLabv3+ network.
7. The dynamic behavior compliance verification device based on AI multi-mode recognition according to claim 6, characterized in that, The compliance verification unit includes: The personnel absence / sleep detection module inputs real-time video data into the personnel absence / sleep detection model, extracts the corresponding spatiotemporal sequence features of posture, and compares them with normal behavior to identify personnel absence / sleep behavior, including: Acquire real-time video data and perform preprocessing, including converting it to H264 encoding, converting the resolution of each frame to 640*640 resolution, and converting the color space from RGB to YUV. The preprocessed real-time video data is input into the pose feature estimation model to extract the coordinates of the human skeleton joints in each frame of the image and construct the spatiotemporal sequence features of the human pose. The pose feature estimation model is obtained by training the OpenPose human pose estimation model with several samples. Each sample includes a human pose map and the corresponding coordinates of the human skeleton joints. The spatiotemporal sequence features of the person's posture are input into the posture classification model to obtain the corresponding posture classification result. The posture classification model is obtained by training a CNN-LSTM multi-classification model with several samples. Each sample includes the spatiotemporal sequence features of the person's posture and the corresponding posture classification result. Based on the posture classification results, the spatiotemporal sequence features of the corresponding normal on-duty behavior postures are retrieved from the pattern library. The dynamic time warping algorithm is used to perform similarity analysis between these features and the spatiotemporal sequence features of the personnel postures to obtain multiple similarity values. If the similarity value is lower than the similarity threshold within a set threshold time period, it is determined that the person has been absent from their post and sleeping on duty. The security passage anomaly detection module takes real-time video data as input to the security passage anomaly detection model and obtains the security passage anomaly detection results, including: YOLO is used to quickly scan each frame of real-time video data globally to obtain the target bounding box of suspicious objects in the real-time video data. Each frame of real-time video data is input into the segmentation mask acquisition model to obtain a segmentation mask. The segmentation mask acquisition model includes a region segmentation sub-model and an optimization sub-model. The region segmentation sub-model is trained on the DeepLabv3+ network using several samples. Each sample includes an image and the corresponding target scene region after segmentation. The optimization sub-model optimizes the segmentation results of the region segmentation sub-model. The optimization process includes using closing operations to fill small holes, opening operations to smooth boundaries, using conditional random fields to refine edges, analyzing connected components to filter out false detection regions with too small an area, and outputting the motion region ROI to calculate the absolute difference between the current frame and the background. The target bounding box is superimposed on the segmentation mask to analyze the relative positional relationship between the boundary of the target bounding box and the key semantic region; The security channel anomaly result of a certain frame image is determined based on the first anomaly recognition rule, wherein the first anomaly recognition rule is set using the relative positional relationship between the boundary of the target bounding box and the key semantic region; Repeat the above steps until the set threshold is met in the continuous abnormal state of the security passage, and generate a security passage abnormal alarm.
8. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the steps of the method as claimed in any one of claims 1 to 5.
9. A storage medium, characterized in that, The storage medium stores a computer program that can be read by a computer, the computer program being configured to execute the steps of the method as described in any one of claims 1 to 5 when it is run.