Abnormal behavior identification method based on multiple types of local tags

By adopting a fusion method of multiple local tags and pose recognition modules in behavior recognition technology, the problem of low recognition accuracy when multiple behaviors occur simultaneously in the prior art is solved, and higher recognition accuracy and reliability are achieved.

CN120148099AInactive Publication Date: 2025-06-13SHAANXI PUBLIC INFORMATION IND CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510091675.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing behavior recognition technologies are difficult to effectively deal with the situation where multiple behaviors occur simultaneously, resulting in low recognition accuracy and affecting the reliability of the identification results.

Method used

The abnormal behavior recognition method based on multiple categories of local tags is adopted. The posture recognition module detects the key points of the human body and generates a skeleton map, fuses the original image with the posture bone points, and uses the multi-label classification module to classify the actions of different parts of the human body, and comprehensively determines the final behavior recognition result through the rule judgment module.

Benefits of technology

It improves the ability to distinguish complex behaviors, enhances the accuracy and reliability of identification results, and can effectively deal with the behavior recognition problem of incomplete targets with local characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148099A_ABST
    Figure CN120148099A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent behavior recognition, and particularly discloses an abnormal behavior recognition method based on multi-class local tags, and the method comprises the steps: S1, receiving video stream input through a front-end input and feedback module, displaying a recognition result, providing a user interaction interface, and classifying and recognizing a plurality of behaviors through a multi-tag classification module at the same time, the method has the advantages that the recognition capability of complex behaviors is improved, the flexibility is high, the posture recognition module provides human body posture data with more accurate features, the basic accuracy of behavior recognition is improved, the posture features are more remarkable due to the generation of the skeleton diagram, meanwhile, the posture features are fully fused with the original diagram, and the action features are enhanced; according to the method, the rule judgment module can integrate multi-label results, the accuracy of abnormal behavior detection is improved, and meanwhile, the behavior recognition problem of incomplete targets with incomplete local features is solved through posture matching, so that the recognition of the abnormal behaviors of the target individuals is realized, and the accuracy of recognition results is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent behavior recognition, and specifically to an abnormal behavior recognition method based on multiple types of local labels. Background Technique

[0002] With the rapid development of artificial intelligence in computer vision technology, behavior recognition has become an important research field. Behavior recognition not only has broad application prospects in smart home security systems, but also plays an important role in fields such as public safety, medical monitoring, and intelligent retail.

[0003] Behavior recognition technology mainly analyzes video streams or image sequences to detect and recognize various human behaviors and actions. Traditional recognition technologies mainly include methods based on template matching, methods based on machine learning, and methods based on deep learning. Among them, template matching technology identifies behaviors by comparing with predefined behavior templates, while machine learning-based methods achieve behavior recognition through feature extraction and classifier training. Deep learning is based on deep learning methods such as convolutional neural networks and long short-term memory networks for behavior recognition. Deep learning methods can automatically learn features and be trained on large-scale data to achieve the recognition of individual behaviors.

[0004] Several common recognition methods in the prior art mostly adopt single-label classification, that is, each sample corresponds to one behavior category. However, behaviors in the real world are often multi-level and diverse. Single-label classification methods can only assign one label to the behavior at each moment and cannot effectively handle the situation where multiple behaviors occur simultaneously, resulting in relatively low accuracy of behavior recognition and affecting the reliability of recognition results. Summary of the Invention

[0005] The purpose of the present invention is to provide an abnormal behavior recognition method based on multiple types of local labels to solve the following technical problems:

[0006] How to achieve simultaneous processing of multiple labels to improve the ability to distinguish complex behaviors.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] An abnormal behavior recognition method based on multiple types of local labels, the method includes:

[0009] S1: Receive video stream input through the front-end input and feedback module, display the recognition result, provide a user interaction interface, and support the user to input new rules or label categories to adjust the system behavior;

[0010] S2: The pose recognition module uses a pose estimation model to detect the human key points in the video stream, estimates the positions of the human skeleton points, and generates a skeleton graph synthesized by these points from these skeleton points, and fuses the detected original image with the mapped pose skeleton points;

[0011] S3: The multi-label classification module performs multi-label classification on the actions of different parts of the human body, processes the input data, and outputs multiple labels, respectively corresponding to the actions of different parts;

[0012] S4: The rule determination module makes a comprehensive determination according to the labels output by the multi-label classification module according to the predefined rules, outputs the final behavior recognition result, and at the same time performs pose judgment on some incomplete recognition individuals through pose matching;

[0013] S5: The training fine-tuning and rule formulation module performs custom training and fine-tuning on the multi-label classification model, and performs iterative updates based on new data.

[0014] Further, the working process of the front-end input and feedback module in S1 includes:

[0015] S101: The user selects corresponding operations in the front-end system, including information such as model selection, rule selection, frame extraction frequency, alarm category, etc.;

[0016] S102: The video access front-end forms pictures through frame extraction and then accesses the multi-class local pose estimation and classification recognition system at the back-end;

[0017] S103: Display the category label data feedback by the back-end detection system and perform alarm and statistics related operations according to the previous settings.

[0018] Further, the working process of the pose recognition module in S2 includes:

[0019] S201: Connect the input picture data to the pose recognition algorithm, generate pose recognition skeleton point predictions, and correct the generated positions of the skeleton points;

[0020] S202: Connect the skeleton points through the set parameters and generation strategies, set identification information such as the colors of the points corresponding to the parts, and finally generate a skeleton graph synthesized by these points;

[0021] S203: After simply processing the skeleton graph, fuse the recognized individuals within the cropped detection frame with the skeleton graph, and transmit the fused result to the multi-label classifier.

[0022] Further, the process of correcting the generated positions of the skeleton points in S203 includes:

[0023] Through the formula Calculate to obtain the correction coefficient ρ of the generated position of the bone points;

[0024] Among them, sg is the height of the individual to be recognized, sg 01 is the preset individual height, tx is the body size of the recognized individual, tx 01 is the preset individual body size, θ is the error influence coefficient, which is set by empirical fitting, Yy is the shadow area in the picture data, μ is the influence coefficient of the shadow area position, which is obtained based on the influence of the shadow area at different positions in the picture on the picture quality through testing, zmj is the total area of the picture, f x is the adjustment coefficient comparison table function, according to The influence of the numerical value of the calculation result on the generated position of the bone points is obtained based on testing.

[0025] Furthermore, the working process of the multi-label classification module in S3 includes:

[0026] S301: After performing scale normalization on the fused picture or feature map generated in the previous stage according to the human body ratio, connect it to the multi-label classification algorithm group;

[0027] S302: The multi-label classification algorithm group performs multi-label recognition on the input data and stores the recognition results;

[0028] S303: Summarize the multi-label classification recognition results and transmit them to the rule determination module.

[0029] Furthermore, the working process of the rule determination module in S4 includes:

[0030] S401: Screen and filter the multi-label output results, and select the incomplete target individuals whose local parts cannot recognize the posture due to occlusion and put them into the library to be evaluated;

[0031] S402: According to the rule setting, directly perform rule recognition on the screened complete recognition result pictures, and judge whether the target individual has abnormal behavior according to the recognition results;

[0032] S403: Perform posture matching on the incomplete individuals in the library to be evaluated, use the known parts for matching and estimate the possible behaviors according to the matching accuracy;

[0033] S404: Summarize all the results and feedback them to the front-end platform.

[0034] Furthermore, the posture matching process in S403 includes:

[0035] Through the formula Calculate to obtain the behavior score Score of the i-th bone point in the incomplete individual i ;

[0036] and calculate the average value r of the total behavior scores of all skeleton points in the incomplete individual through the formula ;

[0037] After that, by comparing the average value r of the total behavior scores of all skeleton points in the incomplete individual with the preset threshold range [r 1 , r 2 ;

[0038] If r ≤ r 1 , it is determined that the incomplete individual has no abnormal behavior or has mild abnormal behavior and no treatment is required;

[0039] If r ∈ [r 1 , r 2 , it is determined that the incomplete individual has moderate abnormal behavior and a reminder is made;

[0040] If r ≥ r 2 , it is determined that the incomplete individual has highly abnormal behavior and a warning is made;

[0041] where i is any skeleton point recognized in the incomplete individual, n is the total number of skeleton points recognized in the incomplete individual, and are the coordinates of the i-th skeleton point of the incomplete target individual and the coordinates of the i-th skeleton point in the matching library respectively, A i and B i are vectors connecting the coordinates of the i-th skeleton point, · represents the dot product, ||A i || and ||B i || are the moduli of the two groups of vectors respectively.

[0042] Furthermore, the process of the training fine-tuning and rule-making module in S5 includes:

[0043] S501: Input training data, set training parameters, select a model and an alarm category;

[0044] S502: Train and test the specified model according to the selected requirements;

[0045] S503: Design a decision rule and test the rule using the prediction results of the new model;

[0046] S504: Evaluate and save the results, and feedback the training results to the front end, and at the same time put the new model on the line.

[0047] Advantages of the present invention:

[0048] (1) The present invention classifies and identifies multiple behaviors simultaneously through a multi-label classification module, improving the ability to distinguish complex behaviors, and being highly flexible. The pose recognition module provides more accurate human pose data, enhancing the basic accuracy of behavior recognition. The generation of the skeleton diagram makes the pose features more prominent, and at the same time, fully integrating it with the original image strengthens the action features, enabling the rule determination module to comprehensively consider the multi-label results, improving the accuracy of abnormal behavior detection. Meanwhile, pose matching makes up for the problem of behavior recognition for incomplete targets with insufficient local features, thus realizing the recognition of abnormal behaviors of target individuals and ensuring the accuracy of the recognition results.

[0049] (2) The present invention detects the human key points in the video stream through a pose estimation model, estimates the positions of the human skeleton points, and generates a skeleton diagram composed of these points through these skeleton points. The original image detected is fused with the mapped pose skeleton points. Since it provides more accurate human pose data, the basic accuracy of behavior recognition can be improved, making the generation of the skeleton diagram make the pose features more prominent. At the same time, fully integrating it with the original image strengthens the action features, making it more convenient for downstream tasks to learn and having good adaptability to various scenarios and pose changes.

[0050] (3) The present invention conducts multi-label classification on the actions of different parts of the human body and uses a single multi-label classification or an algorithm group composed of multiple single-label classifications to achieve refined recognition of complex behaviors. Then, the input data is processed to output multiple labels, each corresponding to the actions of different parts, thereby improving the recognition accuracy and fineness, and being able to classify and identify multiple behaviors simultaneously, thus improving the ability to distinguish complex behaviors, and being highly flexible, and can be customized and extended according to requirements and the actual situation of the data.

[0051] (4) The present invention outputs the final behavior recognition result by comprehensively judging according to the labels output by the multi-label classification module and in accordance with predefined rules, and can simultaneously perform pose judgment on some incomplete recognition individuals through pose matching, thus providing a flexible rule formulation and adjustment function, adapting to different application requirements, and being able to comprehensively consider the multi-label results, improving the accuracy of abnormal behavior detection, and making up for the problem of behavior recognition for incomplete targets with insufficient local features, improving the accuracy and reliability of the recognition results.

[0052] (5) The present invention calculates the average value r of the total behavior scores of all skeleton points in the incomplete individual and compares it with the preset threshold range [r 1 , r 2Compare. By setting like this, according to the comparison result, it can be judged whether there is abnormal behavior in the incomplete individual, and different warning methods can be formulated according to different judgment results, so as to realize the identification of abnormal behavior of the target individual based on local tags, and the accuracy of the abnormal behavior identification result can be guaranteed. Brief Description of the Drawings

[0053] The present invention will be further described below with reference to the accompanying drawings.

[0054] Figure 1 is a flowchart of a method for identifying abnormal behavior based on multiple types of local tags in the present invention;

[0055] Figure 2 is a flowchart of the working process of the front-end input and feedback module in the present invention;

[0056] Figure 3 is a flowchart of the working process of the posture recognition module in the present invention;

[0057] Figure 4 is a flowchart of image fusion in the present invention;

[0058] Figure 5 is a flowchart of feature fusion in the present invention;

[0059] Figure 6 is a flowchart of the working process of the multi-label classification module in the present invention;

[0060] Figure 7 is a schematic block diagram of binary classification action recognition in the present invention;

[0061] Figure 8 is a flowchart of the working process of the rule determination module in the present invention;

[0062] Figure 9 is a flowchart of the process of the training fine-tuning and rule formulation module in the present invention.

[0063] Reference Signs: Detailed Embodiments

[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0065] Please refer to Figure 1 As shown, in one embodiment, the present application provides a method for identifying abnormal behavior based on multiple types of local tags, and the method includes:

[0066] S1: Receive video stream input through the front - end input and feedback module, display the recognition result, provide a user interaction interface, and support the user to input new rules or label categories to adjust the system behavior;

[0067] S2: Detect the human body key points in the video stream through the pose recognition module using a pose estimation model, estimate the positions of the human body bone points, generate a skeleton graph synthesized by these points through these bone points, and fuse the detected original image with the mapped pose bone points;

[0068] S3: Perform multi - label classification on the actions of different parts of the human body through the multi - label classification module, process the input data, and output multiple labels corresponding to the actions of different parts respectively;

[0069] S4: According to the labels output by the multi - label classification module, the rule determination module makes a comprehensive determination according to the predefined rules, outputs the final behavior recognition result, and at the same time makes a pose judgment on some incomplete recognition individuals through pose matching;

[0070] S5: Customize the training and fine - tuning of the multi - label classification model through the training fine - tuning and rule - making module, and perform iterative updates based on new data;

[0071] Through the above technical solutions, this embodiment provides an abnormal behavior recognition method based on multi - class local labels. First, receive video stream input through the front - end input and feedback module, display the recognition result, provide a user interaction interface, and support the user to input new rules or label categories to adjust the system behavior. Then, detect the human body key points in the video stream through the pose recognition module using a pose estimation model, estimate the positions of the human body bone points, and generate a skeleton graph synthesized by these points through these bone points, and fuse the detected original image with the mapped pose bone points. Subsequently, perform multi - label classification on the actions of different parts of the human body through the multi - label classification module, process the input data, and output multiple labels corresponding to the actions of different parts respectively. Then, according to the labels output by the multi - label classification module, the rule determination module makes a comprehensive determination according to the predefined rules, outputs the final behavior recognition result, and at the same time makes a pose judgment on some incomplete recognition individuals through pose matching. Finally, the training fine - tuning and rule - making module can customize the training and fine - tuning of the multi - label classification model and perform iterative updates based on new data;

[0072] By setting it this way, the multi-label classification module can classify and identify multiple behaviors simultaneously, improving the ability to distinguish complex behaviors, and it has strong flexibility. The pose recognition module provides more accurate human pose data, improving the basic accuracy of behavior recognition. The generation of the skeleton diagram makes the pose features more prominent, and at the same time, fully integrating it with the original image strengthens the action features, enabling the rule determination module to comprehensively consider the multi-label results, improving the accuracy of abnormal behavior detection. At the same time, pose matching makes up for the problem of behavior recognition of incomplete target individuals with incomplete local features, thus realizing the recognition of abnormal behaviors of target individuals and ensuring the accuracy of the recognition results.

[0073] Please refer to Figure 2 As shown, the working process of the front-end input and feedback module in S1 includes:

[0074] S101: The user selects corresponding operations in the front-end system, including information such as model selection, rule selection, frame extraction frequency, alarm category, etc.;

[0075] S102: The video accesses the front-end, forms pictures through frame extraction, and then accesses the multi-class local pose estimation and classification recognition system at the back-end;

[0076] S103: Display the category label data fed back by the back-end detection system and perform alarm and statistical related operations according to the previous settings;

[0077] Through the above technical solution, this embodiment provides the working process of the front-end input and feedback module in S1. First, the user selects corresponding operations in the front-end system, including information such as model selection, rule selection, frame extraction frequency, alarm category, etc., and after the video accesses the front-end and forms pictures through frame extraction, it accesses the multi-class local pose estimation and classification recognition system at the back-end. Finally, display the category label data fed back by the back-end detection system and perform alarm and statistical related operations according to the previous settings;

[0078] The front-end input and feedback module is responsible for receiving video stream input, displaying recognition results, and providing a user interaction interface, enabling the module to feedback the recognition results to the user in real time, and supporting the user to input new rules or label categories to adjust the system behavior, thus providing an intuitive operation interface for the user, facilitating operation and monitoring, displaying the recognition results in real time, ensuring that the user can obtain information in a timely manner, supporting user input and adjustment, and improving the flexibility and operability of the system.

[0079] Please refer to Figure 3 As shown, the working process of the pose recognition module in S2 includes:

[0080] S201: Connect the input picture data to the pose recognition algorithm, generate pose recognition skeleton point predictions, and correct the generated positions of the skeleton points;

[0081] S202: Connect the skeletal points according to the set parameters and generation strategy, set identification information such as the color of the points corresponding to the parts, and finally generate a skeleton graph synthesized by these points;

[0082] S203: After simply processing the skeleton graph, fuse the recognized individual within the cropped detection box with the skeleton graph, and input the fused result into the multi-label classifier;

[0083] Through the above technical solution, this embodiment provides the working process of the pose recognition module in S2. First, connect the input image data to the pose recognition algorithm to generate pose recognition skeletal point predictions and correct the generated positions of the skeletal points. Then, connect the skeletal points according to the set parameters and generation strategy, and set identification information such as the color of the points corresponding to the parts. Finally, generate a skeleton graph synthesized by these points. Finally, after simply processing the skeleton graph, fuse the recognized individual within the cropped detection box with the skeleton graph, and input the fused result into the multi-label classifier;

[0084] Among them, there are two types of fusion methods in this process: direct image fusion and feature fusion. The direct image fusion method directly superimposes the image and the skeleton graph. This method is equivalent to directly regarding the points and lines in the skeleton as part of the body of the target individual to strengthen its action features. In the next multi-label classification, the skeleton points in these enhanced images will be regarded as part of the target individual and learned by the classifier. The principle is as Figure 4 shown;

[0085] The feature fusion method first inputs the original image and the skeleton graph into a feature extraction network respectively for feature extraction at the same scale, and then enhances the original pose features by means of feature map superposition or channel superposition. The generated feature map will be directly used as the input data for the next multi-label classification. The principle is as Figure 5 shown;

[0086] Detect the human key points in the video stream through the pose estimation model, estimate the positions of the human skeletal points, and generate a skeleton graph synthesized by these points. Fuse the original detected image with the mapped pose skeletal points. Since more accurate human pose data with features are provided, the basic accuracy of behavior recognition can be improved, making the generation of the skeleton graph make the pose features more prominent. At the same time, fully fusing it with the original image strengthens the action features, making it more convenient for downstream tasks to learn and having good adaptability to various scenarios and pose changes.

[0087] The process of correcting the generated positions of the skeletal points in S203 includes:

[0088] Through the formula Calculate to obtain the correction coefficient ρ of the generated position of the bone points;

[0089] Among them, sg is the height of the individual to be recognized, sg 01 is the preset individual height, tx is the body size of the recognized individual, tx 01 is the preset individual body size, θ is the error influence coefficient, which is set by empirical fitting, Yy is the shadow area in the picture data, μ is the shadow area position influence coefficient, which is obtained based on the influence of the shadow area at different positions in the picture on the picture quality through testing, zmj is the total area of the picture, f x is the adjustment coefficient comparison table function, according to the influence of the numerical value of the calculation result on the generated position of the bone points is obtained through testing;

[0090] Through the above technical solution, this embodiment provides the correction coefficient ρ of the generated position of the bone points, which can be calculated by the formula By combining parameters such as the height and body type of the individual to be recognized, and the quality parameters of the picture information, this embodiment can calculate the correction coefficient of the generated position of the bone points. This data reflects the difference between the individual to be recognized and the preset individual, so as to provide a correction parameter for ensuring the generated position of the bone points of the recognized individual in the subsequent process and ensuring the accuracy of the generated position of the bone points.

[0091] Please refer to Figure 6 As shown, the working process of the multi-label classification module in S3 includes:

[0092] S301: Normalize the scale of the fused picture or feature map generated in the previous stage according to the human body ratio and then connect it to the multi-label classification algorithm group;

[0093] S302: The multi-label classification algorithm group performs multi-label recognition on the input data and stores the recognition results;

[0094] S303: Summarize the multi-label classification recognition results and transmit them to the rule determination module;

[0095] Through the above technical solution, this embodiment provides the working process of the multi-label classification module. First, normalize the scale of the fused picture or feature map generated in the previous stage according to the human body ratio and then connect it to the multi-label classification algorithm group. Then, the multi-label classification algorithm group performs multi-label recognition on the input data and stores the recognition results. Finally, summarize the multi-label classification recognition results and transmit them to the rule determination module. Among them, the multi-label classification algorithm group in this process can be composed of a single multi-label classification algorithm or multiple single-label algorithms with different model complexities. When the recognition difficulty of different parts in the input data is different, considering other situations such as efficiency, the specific selection can be adjusted by itself to adapt to the training and inference of the data set. For example Figure 7as shown;

[0096] Taking binary classification action recognition, such as whether the body twists, as an example, if the judgment conditions are not clear enough, it may affect the recognition. In addition to setting strict data annotation category judgment norms in data annotation to strengthen the characteristics of body twisting for the classifier to learn, the bone point features fused in the input feature map are also more conducive to the classifier to learn the significant features of actions. In addition, in exam monitoring, the action amplitude of cheaters may be extremely small. Considering that body twisting is a continuous process, the classifier may only give a twist judgment when the action is obvious. Here, the binary classifier can be optimized to output a probability value between [0, 1] representing the probability of body twisting according to the twisting amplitude instead of the previous 01 binary classification. On this basis, the rule judgment in the next stage will also be modified to judge whether the action occurs, and instead, it will be a threshold judgment of the probability size. The above situation depends on the requirements;

[0097] By setting like this, through multi-label classification of the actions of different parts of the human body and using a single multi-label classification or an algorithm group composed of multiple single-label classifications, the refined recognition of complex behaviors can be realized. After processing the input data, multiple labels can be output, corresponding to the actions of different parts respectively, thereby improving the recognition accuracy and fineness, and being able to classify and recognize multiple behaviors simultaneously, thus improving the discrimination ability of complex behaviors, and having strong flexibility, which can be customized and extended according to the requirements and the actual situation of the data.

[0098] Please refer to Figure 7 as shown, the working process of the rule judgment module in S4 includes:

[0099] S401: Screen and filter the multi-label output results, select the incomplete target individuals whose local parts cannot recognize the posture due to occlusion, and put them into the library to be evaluated;

[0100] S402: According to the rule setting, directly perform rule recognition on the screened complete recognition result pictures, and judge whether the target individual has abnormal behaviors according to the recognition results;

[0101] S403: Match the postures of the incomplete individuals in the library to be evaluated, use the known parts for matching, and estimate the possible behaviors according to the matching accuracy;

[0102] S404: Summarize all the results and feedback them to the front-end platform;

[0103] Through the above technical solution, this embodiment provides the working process of the rule determination module in S4. First, the multi-label output results are screened and filtered, and the incomplete target individuals whose local parts cannot identify the posture due to occlusion are selected and put into the library to be evaluated. Then, according to the rule setting, the complete recognized result pictures after screening are directly subjected to rule recognition, and whether the target individual has abnormal behavior is judged according to the recognition result. Subsequently, the incomplete individuals in the library to be evaluated are subjected to posture matching, and the known parts are used for matching and the possible behavior is estimated according to the matching accuracy. Finally, all the results can be summarized and fed back to the front-end platform;

[0104] This process is mainly divided into two parallel parts: performing rule determination on the complete recognition results and performing posture matching on the incomplete results. The main process of rule determination is to judge the final behavior by judging the action relationship and angle relationship between different parts. Posture matching is to match and score the recognized local labels of the incomplete target with each behavior in the behavior library, and finally select the behavior with the highest average score;

[0105] By setting like this, this module can output the final behavior recognition result by comprehensively judging according to the labels output by the multi-label classification module and in accordance with the predefined rules, and can simultaneously judge the postures of some incomplete recognized individuals through posture matching, thus providing a flexible rule formulation and adjustment function, adapting to different application requirements, and being able to synthesize multi-label results, improving the accuracy of abnormal behavior detection, and making up for the behavior recognition problem of incomplete target individuals with incomplete local features, improving the accuracy and reliability of the recognition result.

[0106] The posture matching process in S403 includes:

[0107] Through the formula Calculate the behavior score Score of the i-th bone point in the incomplete individual i ;

[0108] And through the formula Calculate the average value r of the total behavior scores of all bone points in the incomplete individual;

[0109] After that, the average value r of the total behavior scores of all bone points in the incomplete individual is compared with the preset threshold interval [r 1 , r 2 ;

[0110] If r ≤ r 1 , it is judged that the incomplete individual has no abnormal behavior or has mild abnormal behavior and no treatment is done;

[0111] If r ∈ [r 1 , r 2, determine that the incomplete individual has moderately abnormal behavior and give a reminder;

[0112] If r ≥ r 2 , determine that the incomplete individual has highly abnormal behavior and give a warning;

[0113] Among them, i is any bone point recognized in the incomplete individual, and n is the total number of bone points recognized in the incomplete individual. and are respectively the coordinates of the i-th bone point of the incomplete target individual and the coordinates of the i-th bone point in the matching library. A i and B i are vectors connecting the coordinates of the i-th bone point, · represents the dot product, ||A i || and ||B i || are the moduli of the two groups of vectors respectively;

[0114] Through the above technical solution, this embodiment provides the average value r of the total behavior scores of all bone points in the incomplete individual, and by comparing the average value r of the total behavior scores of all bone points in the incomplete individual with the preset threshold range [r 1 , r 2 , by setting like this, according to the comparison result, it can be judged whether the incomplete individual has abnormal behavior, and different warning methods can be formulated according to different judgment results, so as to realize the identification of the abnormal behavior of the target individual based on local labels, and the accuracy of the abnormal behavior identification result can be guaranteed.

[0115] Please refer to Figure 8 as shown, the process of training, fine-tuning and rule formulation module in S5 includes:

[0116] S501: Input training data, set training parameters, select a model and an alarm category;

[0117] S502: Train and test the specified model according to the selected requirements;

[0118] S503: Design a judgment rule and test the rule using the prediction results of the new model;

[0119] S504: Evaluate and save the results, feedback the training results to the front end, and at the same time put the new model online;

[0120] Through the above technical solution, this embodiment provides the process of training, fine-tuning and rule formulation in S5. First, training data is input, training parameters are set, a model and an alarm category are selected. Then, the specified model is trained and tested according to the selected requirements. Subsequently, a decision rule is designed and the rule is tested using the prediction results of the new model. Finally, the results are evaluated and saved, and the training results are fed back to the front end. At the same time, the new model is launched. The training, fine-tuning and rule formulation module provides the ability to continuously optimize the model, improves the recognition performance, supports custom training, adapts to different application scenarios, has flexible rule formulation, and can be adjusted and optimized according to actual requirements.

[0121] The above has described in detail an embodiment of the present invention, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of application of the present invention shall still fall within the scope covered by the patent of the present invention.

Claims

1. A method for identifying abnormal behavior based on multiple local labels, characterized in that: The method comprises: S1: Receives video stream input through the front-end input and feedback module, displays recognition results, provides a user interaction interface, and supports users to enter new rules or label categories to adjust system behavior; S2: The posture recognition module detects the key points of the human body in the video stream through the posture estimation model, estimates the position of the human skeleton points, and generates a skeleton graph synthesized by these skeleton points, and fuses the detected original image with the mapped posture skeleton points; S3: Perform multi-label classification on the movements of different parts of the human body through a multi-label classification module, process the input data, and output multiple labels corresponding to the movements of different parts; S4: The rule judgment module performs comprehensive judgment according to the labels output by the multi-label classification module according to predefined rules, outputs the final behavior recognition result, and performs posture judgment on partially incompletely recognized individuals through posture matching; S5: Customize training and fine-tuning of multi-label classification models through training, fine-tuning and rule-making modules, and iteratively update based on new data.

2. According to claim 1, the abnormal behavior recognition method based on multiple local labels is characterized in that: The working process of the front-end input and feedback module in S1 includes: S101: The user selects corresponding operations in the front-end system, including model selection, rule selection, frame extraction frequency, alarm category and other information; S102: The video access front end extracts frames to form an image, which is then connected to a multi-class local posture estimation and classification recognition system at the back end; S103: Display the category label data fed back by the backend detection system and perform alarm and statistics related operations according to the previous settings.

3. The abnormal behavior recognition method based on multiple local labels according to claim 1 is characterized in that: The working process of the gesture recognition module in S2 includes: S201: Connecting the input image data to the posture recognition algorithm, generating posture recognition skeleton point prediction, and correcting the skeleton point generation position; S202: Connecting the skeleton points through the set parameters and generation strategy, and setting the color and other identification information of the corresponding part points, and finally generating a skeleton graph synthesized by these points; S203: After a simple processing of the skeleton image, the identified individuals in the cropped detection frame are fused with the skeleton image, and the fused result is passed to the multi-label classifier.

4. The abnormal behavior recognition method based on multiple local labels according to claim 3 is characterized in that: The process of correcting the generated position of the skeleton point in S203 includes: By formula Calculate and obtain the correction coefficient ρ of the bone point generation position; Among them, sg is the height of the individual to be identified, sg 01 is the preset individual height, tx is the size of the identified individual, tx 01 is the preset individual body size, θ is the error influence coefficient, which is set according to empirical fitting, Yy is the shadow area in the image data, μ is the shadow area position influence coefficient, which is obtained based on the test according to the influence of the shadow area at different positions in the image on the image quality, zmj is the total area of ​​the image, and fx is the adjustment coefficient comparison table function, according to The influence of the calculated value on the position of bone point generation is obtained based on testing.

5. The abnormal behavior recognition method based on multiple local labels according to claim 1 is characterized in that: The working process of the multi-label classification module in S3 includes: S301: The fused image or feature map generated in the previous stage is scale-normalized according to the human body proportions and then connected to the multi-label classification algorithm group; S302: The multi-label classification algorithm group performs multi-label recognition on the input data and stores the recognition results; S303: Summarize the multi-label classification and recognition results and pass them into the rule determination module.

6. The abnormal behavior recognition method based on multiple local labels according to claim 1 is characterized in that: The working process of the rule determination module in S4 includes: S401: Filter the multi-label output results, select the incomplete target individuals whose posture cannot be recognized due to occlusion, and put them into the evaluation database; S402: According to the rule setting, directly perform rule recognition on the complete recognition result image after screening, and judge whether the target individual has abnormal behavior according to the recognition result; S403: performing posture matching on the incomplete individuals in the database to be evaluated, using known parts for matching and estimating possible behaviors according to matching accuracy; S404: Summarize all the results and feed them back to the front-end platform.

7. The abnormal behavior recognition method based on multiple local labels according to claim 6 is characterized in that: The posture matching process in S403 includes: By formula Calculate the behavior score of the i-th bone point in the incomplete individual i ; And through the formula Calculate the average total behavioral score r of all skeletal points in the mutilated individuals; Then, the average value r of the total behavior score of all skeletal points in the incomplete individual is compared with the preset threshold interval [r1, r2] of the total behavior score; If r≤r1, the disabled individual is judged to have no abnormal behavior or mild abnormal behavior, and no treatment is taken; If r∈[r1, r2], the disabled individual is judged to have moderate abnormal behavior and a reminder is given; If r ≥ r2, the disabled individual is judged to have highly abnormal behavior and a warning is issued; Among them, i is any bone point identified in the incomplete individual, n is the total number of bone points identified in the incomplete individual, and are the coordinates of the i-th bone point of the incomplete target individual and the coordinates of the i-th bone point in the matching library, respectively. i and B i is the vector connecting the coordinates of the i-th bone point, · represents the dot product, ||A i || and ||B i || respectively to the magnitude of the two sets of vectors.

8. The abnormal behavior recognition method based on multiple local labels according to claim 1 is characterized in that: The process of training the fine-tuning and rule-making module in S5 includes: S501: input training data, set training parameters, select model and alarm category; S502: Train and test the specified model according to the selected requirements; S503: Designing judgment rules and testing the rules using the prediction results of the new model; S504: Evaluate and save the results, feed back the training results to the front end, and launch the new model.

Citation Information

Cited By

  • Community abnormal behavior identification method and device based on video analysis

    CN121982817A