Abnormal behavior early warning method and device based on intention-behavior fusion
By adopting the early warning method of intention-behavior fusion in video surveillance, the behavioral characteristics and intention characteristics in video are extracted and analyzed, and the problem of difficulty in capturing hidden abnormal intentions is solved in the existing technology, and a more accurate and timely abnormal behavior warning is achieved.
Patent Information
- Application Number
- CN202510049222.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has shortcomings in detecting and early warning of potential abnormal intentions, and it is difficult to capture more concealed abnormal intentions in time and action, making it difficult to ensure the accuracy and timeliness of early warnings.
An abnormal behavior warning method based on intention-behavior fusion is adopted. By obtaining the original video of the target person, the behavior feature vector is extracted and the behavior feature sequence is generated, and the fusion feature is calculated in combination with the weights to obtain the intention feature. Then, input the intent feature into the intent prediction model to generate the abnormal intent prediction results, and determine the hidden intent positioning results based on the results to generate abnormal behavior warning information.
It significantly improves the accuracy of abnormal behavior detection and early warning capabilities, can effectively integrate multiple behavioral characteristics, enhance the prediction ability of intention prediction models, provide clearer abnormal behavior paths and prediction basis, and improves security personnel's understanding and response capabilities for potential threats.
Smart Images

Figure CN120047997A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of social security technologies, and particularly to an abnormal behavior early warning method and device based on intention-behavior fusion. Background Art
[0002] In modern society, public safety and personal property protection have increasingly become important concerns. Although relevant monitoring technologies can record and identify obvious abnormal behaviors, there are still significant deficiencies in detecting hidden abnormal intentions.
[0003] In related technologies, traditional temporal action localization methods mainly focus on identifying and classifying significant actions from videos and marking the start and end times of these actions. Datasets such as THUMOS14, ActivityNet, and Kinetics are usually used. The wide application of these datasets has promoted the development of temporal action localization technologies, giving rise to advanced action localization models such as Actionformer and TriDet. Although these models perform well in localizing significant actions, in the fields of intelligent video surveillance and theft prevention, some methods and systems have proposed predictive solutions, such as methods and systems for preventing theft based on video recognition and detecting movement trajectories, through moving target extraction, pedestrian recognition, trajectory detection, and behavior judgment, as well as methods and systems for detecting theft behaviors such as climbing stairs and climbing through windows based on video surveillance.
[0004] However, although the temporal action localization methods in related technologies have achieved certain results in some aspects, they still mainly focus on the recognition of explicit behaviors, and it is difficult to capture those more concealed and potential abnormal intentions in terms of time and actions, and localize and early warn them. The application scope and effectiveness are relatively limited, and it is also difficult to ensure the accuracy and timeliness of early warnings, which urgently need to be solved. Summary of the Invention
[0005] This application provides an abnormal behavior early warning method and device based on intention-behavior fusion to solve the problems that although the temporal action localization methods in related technologies have achieved certain results in some aspects, they still mainly focus on the recognition of explicit behaviors, it is difficult to capture those more concealed and potential abnormal intentions in terms of time and actions, and localize and early warn them, the application scope and effectiveness are relatively limited, and it is also difficult to ensure the accuracy and timeliness of early warnings.
[0006] The first aspect embodiment of this application provides an abnormal behavior early warning method based on intention-behavior fusion, including the following steps: obtaining the original video of the target person, extracting the behavior feature vectors of each video frame of the target person in the original video, and generating a behavior feature sequence according to the behavior feature vectors; calculating the fusion feature of the behavior features based on the behavior feature sequence and the corresponding weights, and obtaining the intention feature of the target person according to the fusion feature; inputting the intention feature into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and determining a hidden intention localization result of the target person according to the abnormal prediction result, and generating an abnormal behavior early warning information of the target person based on the hidden intention localization result.
[0007] Optionally, in an embodiment of this application, the obtaining the original video of the target person and extracting the behavior feature vectors of each video frame of the target person in the original video includes: identifying suspicious video segments in the target hidden intention video that meet preset conditions; using the suspicious video segments to correct the target video feature extraction model to obtain a final video feature extraction model, and using the final video feature extraction model to extract the behavior feature vectors of each video frame of the target person in the original video.
[0008] Optionally, in an embodiment of this application, the determining the hidden intention localization result of the target person according to the abnormal prediction result includes: collecting the types of abnormal behaviors in the original video; assigning suspicion scores to at least one type of abnormal behavior in each frame of the original video; calculating the total suspicion score of each frame according to the types of abnormal behaviors and the suspicion scores, normalizing the total suspicion score to obtain a normalized total suspicion score, and determining the hidden intention localization result of the target person according to the normalized total suspicion score.
[0009] Optionally, in an embodiment of this application, the hidden intention localization result includes at least one of a start time prediction result, an end time prediction result, and an abnormal degree prediction result of the abnormal intention of the target person.
[0010] Optionally, in an embodiment of this application, the inputting the intention feature into a preset intention prediction model to analyze and generate an abnormal intention prediction result of the target person includes: analyzing the intention features at different time scales to capture short-term behavior change information of the target person within a first preset duration in the future and long-term behavior change information within a second preset duration in the future, where the second preset duration is greater than the first preset duration; generating an abnormal intention prediction result of the target person according to the short-term behavior change information and the long-term behavior change information.
[0011] The second aspect of the present application provides an abnormal behavior warning device based on intention-behavior fusion, including: an extraction module, configured to obtain the original video of the target person, extract the behavior feature vectors of each video frame of the target person in the original video, and generate a behavior feature sequence according to the behavior feature vectors; a fusion module, configured to calculate the fusion features of the behavior features based on the behavior feature sequence and the corresponding weights, and obtain the intention features of the target person according to the fusion features; a warning module, configured to input the intention features into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and determine a hidden intention positioning result of the target person according to the abnormal prediction result, and generate an abnormal behavior warning information of the target person based on the hidden intention positioning result.
[0012] Optionally, in an embodiment of the present application, the extraction module includes: an identification unit, configured to identify suspicious video segments in the target hidden intention video that meet preset conditions; an extraction unit, configured to use the suspicious video segments to correct the target video feature extraction model to obtain a final video feature extraction model, and use the final video feature extraction model to extract the behavior feature vectors of each video frame of the target person in the original video.
[0013] Optionally, in an embodiment of the present application, the warning module includes: a collection unit, configured to collect the types of abnormal behaviors in the original video; an assignment unit, configured to assign suspicion scores to at least one type of abnormal behavior in each frame of the original video; a calculation unit, configured to calculate the total suspicion score of each frame according to the types of abnormal behaviors and the suspicion scores, normalize the total suspicion score to obtain a normalized total suspicion score, and determine the hidden intention positioning result of the target person according to the normalized total suspicion score.
[0014] Optionally, in an embodiment of the present application, the hidden intention positioning result includes at least one of a start time prediction result, an end time prediction result, and an abnormal degree prediction result of the abnormal intention of the target person.
[0015] Optionally, in an embodiment of the present application, the warning module includes: an analysis unit, configured to analyze the intention features on different time scales to capture short-term behavior change information of the target person within a first preset duration in the future and long-term behavior change information within a second preset duration in the future, where the second preset duration is greater than the first preset duration; a generation unit, configured to generate an abnormal intention prediction result of the target person according to the short-term behavior change information and the long-term behavior change information.
[0016] A third - aspect embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the abnormal behavior early - warning method based on intention - behavior fusion as described in the above - mentioned embodiments.
[0017] A fourth - aspect embodiment of the present application provides a computer - readable storage medium storing a computer program, which when executed by a processor implements the above - mentioned abnormal behavior early - warning method based on intention - behavior fusion.
[0018] A fifth - aspect embodiment of the present application provides a computer program product including a computer program, which when executed is used to implement the above - mentioned abnormal behavior early - warning method based on intention - behavior fusion.
[0019] Embodiments of the present application can extract behavior features from the original video of the target person, perform weighted calculation to obtain the intention features of the target person, and finally obtain the hidden - intention positioning result of the target person, which is convenient for timely generating abnormal behavior early - warning information for the target person. Thus, the concept of time - based intention positioning is realized. By analyzing past behavior sequences to predict future abnormal behaviors and hidden behavior intentions, the accuracy of abnormal behavior detection and the ability of early warning are significantly improved. Further, the present application can also effectively fuse various behavior features to enhance the prediction ability of the intention prediction model; at the same time, the present application also has high interpretability and operability. Through multi - level behavior annotation and fine - grained behavior analysis, it can provide a clearer abnormal behavior path and prediction basis, enabling security personnel to more accurately understand and respond to potential threats. Thus, it solves the problems in the related - art time - action positioning methods. Although certain achievements have been made in some aspects, they still mainly focus on the recognition of explicit behaviors, it is difficult to capture those more concealed and potential abnormal intentions in terms of time and actions, and to locate and warn them, the application scope and effectiveness are relatively limited, and it is also difficult to guarantee the accuracy and timeliness of early warning.
[0020] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above - mentioned and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0022] Figure 1 It is a flowchart of an abnormal behavior early - warning method based on intention - behavior fusion according to an embodiment of the present application;
[0023] Figure 2Schematic diagram of the framework of the IAF module according to an embodiment of the present application;
[0024] Figure 3 Schematic diagram for comparison between traditional time-action localization and time-intention localization according to an embodiment of the present application;
[0025] Figure 4 Flowchart of the abnormal behavior warning method based on intention-action fusion according to an embodiment of the present application;
[0026] Figure 5 Schematic diagram of the model result of the hidden abnormal intention data set according to an embodiment of the present application;
[0027] Figure 6 Schematic diagram of the structure of the abnormal behavior warning device based on intention-behavior fusion according to an embodiment of the present application;
[0028] Figure 7 Schematic diagram of the structure of the electronic device according to an embodiment of the present application.
[0029] Reference numerals:
[0030] 10 - Abnormal behavior warning device based on intention-behavior fusion: 100 - Extraction module, 200 - Fusion module, and 300 - Warning module; 701 - Memory, 702 - Processor, and 703 - Communication interface. Detailed implementation manners
[0031] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0032] The following describes the abnormal behavior warning method, device, electronic device, and storage medium based on intention-behavior fusion according to the embodiments of the present application. Regarding the time-action localization method in the related technology mentioned in the above background technology, although certain achievements have been made in some aspects, it still mainly focuses on the recognition of explicit behaviors, and it is difficult to capture those more concealed and potential abnormal intentions in terms of time and actions, and localize and warn them. The application scope and effectiveness are relatively limited, and it is also difficult to ensure the accuracy and timeliness of the warning. The present application provides an abnormal behavior warning method based on intention-behavior fusion. In this method, the behavior features can be extracted from the original video of the target person and weighted calculation is performed to obtain the intention features of the target person, and finally the hidden intention localization result of the target person is obtained, which is convenient for generating abnormal behavior warning information for the target person in a timely manner. Thus, the concept of time-intention localization is realized. By analyzing the past behavior sequences, future abnormal behaviors and hidden behavior intentions can be predicted, significantly improving the accuracy of abnormal behavior detection and the ability of early warning. Further, the present application can also effectively fuse various behavior features to enhance the prediction ability of the intention prediction model; at the same time, the present application also has high interpretability and operability. Through multi-level behavior annotation and fine-grained behavior analysis, a clearer abnormal behavior path and prediction basis can be provided, enabling security personnel to more accurately understand and respond to potential threats. Thus, the problems in the related technology that although the time-action localization method has achieved certain results in some aspects, it still mainly focuses on the recognition of explicit behaviors, it is difficult to capture those more concealed and potential abnormal intentions in terms of time and actions, and localize and warn them, the application scope and effectiveness are relatively limited, and it is also difficult to ensure the accuracy and timeliness of the warning, etc. are solved.
[0033] Before explaining the abnormal behavior warning method based on intention-behavior fusion according to the embodiments of the present application, the structure of the abnormal behavior warning system based on intention-behavior fusion according to the embodiments of the present application will be described first.
[0034] The abnormal behavior warning system based on intention-behavior fusion according to the embodiments of the present application includes but is not limited to four parts: a final video feature extraction model, an IAF module, an intention prediction model, and a post-processing module.
[0035] Among them, the final video feature extraction model is mainly used to extract the behavior features of the target person from the original video;
[0036] The IAF module is mainly used to learn the weights of the corresponding behavior features according to the extracted behavior features, and perform weighted summation with the corresponding behavior features to obtain a weighted sum, so as to fuse with the intention features to obtain a fused feature;
[0037] The intention prediction model is mainly used to analyze the input fusion features and output the corresponding abnormal intention prediction results;
[0038] The post - processing module is mainly used to post - process the abnormal intention prediction results output by the intention prediction model and output the hidden intention location results including the time location information of the abnormal intention prediction results.
[0039] Specifically, Figure 1 FIG. is a flowchart of an abnormal behavior warning method based on intention - behavior fusion provided by an embodiment of the present application.
[0040] As Figure 1 shown, the abnormal behavior warning method based on intention - behavior fusion includes the following steps:
[0041] In step S101, the original video of the target person is obtained, and the behavior feature vectors of each video frame of the target person in the original video are extracted to generate a behavior feature sequence according to the behavior feature vectors.
[0042] It can be understood that the target person here refers to the person whose behavior is analyzed to determine whether abnormal behavior warning is required. For example, some people with abnormal behaviors in places with large crowds.
[0043] In some embodiments, the target persons are mostly those with slightly abnormal action behaviors compared with ordinary people. Before the target person actually makes a behavior with a relatively high degree of abnormality that requires warning, in order to avoid interfering with them, the present application can first obtain the original video of the target person so as to analyze them based on the original video of the target person.
[0044] Among them, the original video can be understood here as the historical video of the target person, which can be the historical video in the current scene in the past period of time, or the historical video of the target person in other scenes. Specifically, it can be selected or adjusted by those skilled in the art according to the actual situation. Here, it is only for illustrative purposes and not for specific limitation.
[0045] After obtaining the original video of the target person, the embodiment of the present application mainly but not limited to uses its behavior features as the main analysis part. After extracting the behavior feature vectors of each video frame of the target person in the original video, these extracted features are then organized into a feature sequence for subsequent intention prediction of the target person.
[0046] Next, a further description is given of the process of how to extract the behavior feature vectors of the target person in each video frame from the original video.
[0047] Optionally, in an embodiment of the present application, the original video of the target person is obtained, and the behavior feature vectors of each video frame of the target person in the original video are extracted, including: identifying suspicious video segments that meet preset conditions in the target hidden intention video; using the suspicious video segments to correct the target video feature extraction model to obtain a final video feature extraction model, so as to extract the behavior feature vectors of each video frame of the target person in the original video by using the final video feature extraction model.
[0048] It can be understood that the target hidden intention video here refers to some videos that can be extracted to obtain the behavior features required for training the final video feature extraction model. The target video feature extraction model here can be understood as an original video feature extraction model that has not been trained to extract the behavior features and other features of the people in the video. The preset conditions here refer to certain criteria or conditions that the suspicious video segments should meet. For example, there are suspicious people, and the suspicious people have abnormal intentions, etc.
[0049] In the actual execution process, when extracting the behavior feature vectors of each video frame of the target person in the original video, in order to be able to directly extract the behavior features required by the present application from the video by using the model, and to more comprehensively cover the possible abnormal behaviors of the target person, the present application can extract suspicious video segments based on a certain target hidden video, and then use the suspicious video segments to train and optimize the target video feature extraction model to obtain a final video feature extraction model that can be actually applied in the present application, and then use the final video feature extraction model to directly extract the behavior feature vectors of each video frame of the target person in the original video from the original video.
[0050] For example, the present application can have professional technicians in the field mark the suspicious video segments in some target hidden intention videos, and then integrate this annotation information. By taking the union and intersection of different annotation data, retain the segments considered suspicious by any annotator (forming annotation L) and the segments unanimously considered suspicious by all annotators (forming annotation S) to simulate different sensitivity levels. For example, assume that multiple annotators have annotated the same video segment, and record the annotation result of each annotator as L i (where i represents different annotators), then:
[0051]
[0052] Among them, L represents the set of segments considered suspicious by any annotator, and S represents the set of segments unanimously considered suspicious by all annotators. Such a method can effectively simulate different sensitivity levels to more comprehensively cover possible abnormal behaviors.
[0053] Thus, the suspicious video segments required for training the target video feature extraction model in the embodiments of the present application can be obtained, and the target video feature extraction model can be trained using these suspicious video segments to obtain the final video feature extraction model.
[0054] Subsequently, in the embodiments of the present application, the final video feature extraction model can be used to extract the behavior features of the target person from the original video. When extracting the behavior features of the target person from the original video, the embodiments of the present application can, but are not limited to, divide the behavior features of the target person into feature vectors for each video frame. These feature vectors include, but are not limited to, information with time and space dimensions. Further, the embodiments of the present application can also organize the extracted feature vectors into a feature sequence for subsequent dynamic feature fusion and intention prediction. Specifically, let the feature vector of video frame t be F t , then its expression can be as follows:
[0055] F t = Model(I t ),
[0056] where I t represents the image data of the t-th frame, and Model represents the final feature extraction model. These feature vectors include information with time and space dimensions. Organizing the extracted feature vectors into a feature sequence {F 1 , F 2 , …, F T} can be used for subsequent dynamic feature fusion and intention prediction.
[0057] Step S102: Calculate the fusion feature of the behavior features based on the behavior feature sequence and the corresponding weights, so as to obtain the intention feature of the target person according to the fusion feature;
[0058] In some other embodiments, considering that the intention of the target person will have a certain impact on the behavior of the target person, the present application can use the IAF module to learn the weights of different behaviors, and then perform dynamic weighted summation on the features of different behavior categories to calculate the dynamic feature result of the behavior features of the target person, that is, the fusion feature, so as to obtain the intention feature of the target person. For the convenience of calculation, the embodiments of the present application can convert the obtained intention feature into a vector form and input it into a certain model to predict the hidden intention of the target person.
[0059] Figure 2 is a schematic framework diagram of the IAF module in an embodiment of the present application. As Figure 2As shown in the figure, the low-level action features are input into the IAF (Intention-Action Fusion) module. The IAF module can perform affine coupling on the obtained behavior feature sequence, that is, perform weighted summation on the behavior feature sequence and the corresponding learnable weights to obtain a weighted sum, and fuse it with the original behavior features to obtain fused features, and perform normalization processing on the fused features. Finally, the intention features of the target person are obtained from the normalized fused features. Among them, the low-level action features here refer to the behavior features extracted by using the target video feature extraction model or the final video feature extraction model; the original features refer to the relevant features of the behavior obtained by further feature extraction processing of the behavior features.
[0060] Specifically, the embodiment of the present application can first calculate the weights of each behavior category of the target person, and use W k to represent the weight of the k-th type of behavior feature. It should be noted that these weights can be predefined, mainly obtained through learning in the training stage, and can be set or adjusted through training data or professional knowledge in the later stage. The present application does not make specific limitations. The setting of the weights reflects the importance of each behavior feature in predicting abnormal intentions. Among them, the weight W k in the embodiment of the present application can be determined by, but not limited to, the following methods:
[0061]
[0062] Among them, g(A i,k ) represents the score of the k-th type of behavior in the i-th sample, and N represents the total number of samples.
[0063] Next, the embodiment of the present application needs to perform weighted calculation on the feature values of the behavior features at each time point. For the behavior feature T b,k,t at each time point t, that is, at time point t, the feature value of the k-th type of behavior in the b-th video, calculate the weighted sum S b,t , and the formula can be represented by, but not limited to, the following:
[0064]
[0065] Among them, k represents the number of behavior categories.
[0066] Finally, the weighted sum S b,t is dynamically fused with the original features of the target person to obtain updated fused features, and the fused features are normalized, and the intention features of the target person are characterized by the normalized fused features, which can be represented in vector form by, but not limited to:
[0067] T′ b,l,t = T b,l,t + S b,1,t ,
[0068] Among them, l represents the intention level and is applicable to all time points t.
[0069] In the embodiments of the present application, an updated fusion feature can be generated through dynamic fusion of weighted sums, and the normalized fusion feature is used to characterize the intention feature of the target person. Thus, the intention feature in the embodiments of the present application includes the behavior feature and intention feature at each time point. Further, the embodiments of the present application can also combine the intention features to form an intention feature vector sequence, and perform intention prediction of the target person according to the intention feature vector sequence.
[0070] In addition, the embodiments of the present application also introduce a dynamic weight adjustment method based on the attention mechanism, which can automatically adjust the weight of each behavior category according to the content of the original video, context dynamics, and the importance of behavior features at different time points, thereby improving the accuracy of the obtained fusion feature. Among them, the attention mechanism in the embodiments of the present application can be but is not limited to being expressed as:
[0071] W k ′ = Attention(T b,k,t ),
[0072] where Attention represents the attention function, which can dynamically adjust the weight according to the current video content and behavior features.
[0073] The embodiments of the present application can combine time information and key behavior features, perform weighted analysis on multiple behavior subcategories in the video through dynamic feature fusion, dynamically adjust the weights of each behavior feature, significantly improve the sensitivity and interpretability of the model in identifying hidden abnormal intentions, thereby optimizing abnormal intention prediction. While paying attention to obvious actions, the behavior sequence intention and mental state of the target person are also analyzed, thereby enhancing the ability of the model to identify potential threats in complex environments, and improving the accuracy and timeliness of the early warning of the present application.
[0074] Step S103: Input the intention feature into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and determine a hidden intention localization result of the target person according to the abnormal prediction result, so as to generate an abnormal behavior early warning information of the target person based on the hidden intention localization result. The hidden intention localization result includes at least one of a start time prediction result, an end time prediction result, and an abnormal degree prediction result of the abnormal intention of the target person.
[0075] It can be understood that the preset intention video model here refers to an intention prediction model trained using the fusion feature, which can output some initial prediction results of the target person, such as whether the target person has an abnormal intention, etc.
[0076] As a possible implementation, after obtaining the intended feature sequence vector in the embodiments of the present application, it can be input into a certain intention prediction model, and the intention prediction model analyzes the fused feature vectors at each time point to obtain the abnormal intention prediction result of the target person.
[0077] For example, let the fused feature vector at the t-th moment be T′ t Then, the abnormal intention prediction result can be but is not limited to being expressed in the form of a formula as:
[0078] P t = Model(T t ′ ),
[0079] where Model represents the intention prediction model, and P t represents the abnormal intention prediction result at the t-th moment.
[0080] Among them, the abnormal intention prediction result only includes the degree of abnormal intention of the target person, and does not include time positioning information such as the start time and end time of the abnormal intention of the abnormal intention of the target person. Therefore, the embodiments of the present application can also use a post-processing module to further post-process the abnormal intention prediction result to obtain and output the hidden intention positioning result of the target person. Among them, the hidden intention positioning result includes but is not limited to the start time, end time, and abnormal degree (for example, normal, uncertain, suspicious, alarm) of the abnormal intention of the target person.
[0081] Finally, the embodiments of the present application can further analyze and process the video according to the predicted abnormal intention prediction result and the hidden intention positioning result to generate an abnormal behavior warning message for the target person. For example, warning on the video interface: "This target person is suspected of having an attack intention and attack behavior", etc., and can trigger relevant security systems according to this warning message.
[0082] Optionally, in an embodiment of the present application, determining the hidden intention positioning result of the target person according to the abnormal prediction result includes: collecting the types of abnormal behaviors of the target hidden intention video; assigning a suspicion score to at least one type of abnormal behavior in each frame of the target hidden intention video; calculating the total suspicion score of each frame according to the types of abnormal behaviors and the suspicion scores, normalizing the total suspicion score to obtain the normalized total suspicion score, so as to determine the hidden intention positioning result of the target person according to the normalized total suspicion score.
[0083] In the actual execution process, since the dataset in the initial stage only has a series of annotations based on the start and end times of suspicious behaviors in videos with target hidden intentions, the embodiments of the present application can first calculate the total suspicion score at this time point through the joint linear assignment method, based on the types, durations, frequencies, etc. of the suspicious behaviors that occur at this time point, and evaluate the accuracy of the hidden intention localization results output during the training process with this total suspicion score.
[0084] And in the actual application process, the embodiments of the present application can also output the total suspicion score as the label of the finally output hidden intention localization result. Among them, this total suspicion score is mainly but not limited to obtained through a series of processes on the suspicion scores assigned to each type of abnormal behavior in each frame of the original video.
[0085] For example, the embodiments of the present application can first collect the types of abnormal behaviors of the target person in the original video, and then assign a suspicion score to each type of abnormal behavior in each frame of the original video. Finally, combine the types of abnormal behaviors and the suspicion scores of each frame to calculate the total suspicion score of each frame and perform normalization processing, and finally obtain the normalized total suspicion score, thereby training and obtaining a certain intention prediction model. The specific process can be expressed as follows:
[0086] (1) Suspicion score assignment: First, label each behavior category (such as "suspicious gaze") and assign a suspicion score to each frame of this behavior category. Among them, this suspicion score can be calculated based on but not limited to the repeatability and duration of the behavior. For example, setting a scoring function for a behavior A as f(A), the suspicion score calculation formula can be expressed as:
[0087] f(A) = α·d(A) + β·r(A),
[0088] where d(A) represents the duration of behavior A, r(A) represents the number of repetitions of behavior A, and α and β are weight coefficients used to adjust the influence of duration and number of repetitions on the score.
[0089] (2) Calculate behavior accumulation: Perform frame-level behavior accumulation. For each frame, accumulate the suspicion scores of all behavior category labels to form a total suspicion score. Let the total suspicion score of the t-th frame be S t , then it can be expressed as:
[0090]
[0091] where K represents the number of behavior categories, and A k,t represents the occurrence situation of the k-th type of behavior in the t-th frame. These scores reflect the cumulative effect of all suspicious behaviors that occur in this frame.
[0092] (3) Fractional normalization: Perform Sigmoid normalization on the accumulated suspicion scores to make them fall between 0 and 1, ensuring the consistency and comparability of the scores. Let S norm,t be the normalized score, then it can be expressed as follows:
[0093]
[0094] where S t is the total suspicion score of the t-th frame.
[0095] Optionally, in an embodiment of the present application, input the intention feature into a preset intention prediction model for analysis and generate an abnormal intention prediction result of the target person, including: analyzing the intention features on different time scales to capture the short-term behavior change information of the target person within the first preset duration in the future and the long-term behavior change information within the second preset duration in the future, where the second preset duration is greater than the first preset duration; generating an abnormal intention prediction result of the target person according to the short-term behavior change information and the long-term behavior change information.
[0096] It can be understood that both the first preset duration and the second preset duration can be understood as the durations set in advance here. For example, 5 minutes and 10 minutes, etc., where the second preset duration is greater than the first preset duration.
[0097] In other embodiments, in order to improve the robustness of the abnormal intention localization result in the present application, the present application can also use a multi-scale analysis method to analyze the features on different time scales to capture the short-term and long-term behavior changes of the target person and improve the robustness. Specifically, the multi-scale analysis method can be expressed as:
[0098]
[0099] where Scale s represents the feature analysis function of the s-th scale, S represents the number of scales, and M t represents the feature vector after comprehensive multi-scale analysis. In addition, the system continuously updates and optimizes the model parameters during operation, performs online learning based on new data, and improves the system's adaptability to newly emerging abnormal behavior patterns. The online learning mechanism can be expressed as:
[0100]
[0101] where θ t represents the model parameter at the t-th moment, η represents the learning rate, L represents the loss function, and D t represents the new data at the t-th moment. Through online learning, the system can continuously optimize its detection and early warning capabilities for abnormal behaviors, improving the overall security and reliability.
[0102] Figure 3 A comparative schematic diagram of traditional time-action localization and time-intention localization for an embodiment of this application. As Figure 3 shown, compared with the traditional time-action localization method, the embodiment of this application can, based on the concept of time-intention localization, identify and locate potential hidden intentions in unclipped videos, reveal the hidden behavioral intentions of target personnel by analyzing the behavior sequence, thereby realizing not being limited to the current visible actions, but also being able to predict future abnormal behaviors by analyzing past behaviors and evaluate the intentions behind these behaviors, significantly improving the accuracy of abnormal behavior detection and the ability of early warning in this application.
[0103] The following specific embodiment is used to illustrate the abnormal behavior early warning method based on intention-behavior fusion in the embodiment of this application and verify its effectiveness.
[0104] Figure 4 A flowchart of the abnormal behavior early warning method based on intention-action fusion for an embodiment of this application. As Figure 4 shown:
[0105] Step S401, obtain the hidden intention video of the target personnel;
[0106] Step S402, by taking the union and intersection of different annotation data, retain the segments considered suspicious by any annotator (forming annotation L) and the segments unanimously considered suspicious by all annotators (forming annotation S) to simulate different sensitivity levels and obtain the initial behavioral characteristics at different sensitivity levels;
[0107] Step S403, assign a suspicion score to each frame of each behavior category label;
[0108] Step S404, for each frame, accumulate the suspicion scores of all behavior category labels to form a total suspicion score;
[0109] Step S405, normalize the accumulated suspicion scores so that they fall between 0 and 1, determine the reliability of the entire system with the normalized suspicion scores, and use them as the labels of the finally output abnormal intention prediction results;
[0110] Step S406, obtain the feature vectors of each video frame, these feature vectors include information in the time dimension and the space dimension, and organize the extracted feature vectors into a feature sequence for subsequent dynamic feature fusion and intention prediction;
[0111] Step S407, calculate the weights of each behavior category;
[0112] Step S408, calculate the weighted sum of the behavioral characteristics at each time point;
[0113] Step S409: Integrate the weighted sum with the initial behavioral features, and normalize the integrated features. Finally, the normalized integrated features represent the intention features of the target person, and an intention feature vector is obtained.
[0114] Step S410: Input the intention feature vector into a certain intention prediction model to obtain an intention prediction result.
[0115] In the embodiment of the present application, the mean average precision (mAP) is selected as the evaluation index for the experiment, and the definition of the evaluation index can be expressed as follows:
[0116]
[0117] where N i is the number of detections of class i, P(k) is the precision of the k-th detection, AP represents the average precision of each test instance, and C is the total number of test instances.
[0118] Figure 5 is a schematic diagram of the model results of the hidden abnormal intention dataset in an embodiment of the present application. According to Figure 5 the information, Table 1 lists the accuracy comparison (mAP@IoU) of the Actionformer model with and without the IAF module, and Table 2 lists the accuracy comparison (mAP@IoU) of the TriDet model with and without the IAF module, which can be expressed as follows:
[0119] Table 1
[0120]
[0121] Table 2
[0122]
[0123] Table 1 shows the performance of the Actionformer model and the present application under different annotation criteria. In terms of the mean average precision (mAP), the method for early warning of abnormal behaviors based on intention-behavior fusion integrating the IAF module in the embodiment of the present application is significantly improved compared with the method without using the IAF module. For example, in the context of Annotation 3, the average mAP of the Actionformer model is increased from 1.09% to 2.87%, and the average mAP of the TriDet model is increased from 1.10% to 3.90%, with increases of 163.3% and 254.5% respectively. This indicates that traditional behavior recognition technologies have limited effects in dealing with hidden abnormal behaviors, while the present application explores hidden intentions by detecting hidden information, significantly improving the performance of the model.
[0124] Table 2 shows the performance of the TriDet model under different annotation criteria. In the scenario of annotating S, the average mAP of the TriDet model increased from 0.65% to 1.33%, an increase of 104.6%. In the scenario of annotating L, the average mAP of the TriDet model increased from 2.92% to 4.96%, an increase of 69.9%. These results indicate that this application can effectively learn and fuse various behavior features, significantly improving the model's prediction ability in different scenarios.
[0125] According to the abnormal behavior warning method based on intention-behavior fusion proposed in the embodiments of this application, behavior features can be extracted from the original video of the target person and fused with the intention features of the target person, and finally the hidden intention positioning result of the target person can be obtained, facilitating the timely generation of abnormal behavior warning information for the target person. Thus, the concept of time intention positioning is realized. By analyzing past behavior sequences to predict future abnormal behaviors and hidden behavior intentions, the accuracy of abnormal behavior detection and the ability of early warning are significantly improved. Further, this application can also effectively fuse various behavior features and enhance the prediction ability of the intention prediction model; at the same time, this application also has high interpretability and operability. Through multi-level behavior annotation and fine-grained behavior analysis, it can provide a clearer abnormal behavior path and prediction basis, enabling security personnel to more accurately understand and respond to potential threats. Thus, it solves the problems in the related art that although the time action positioning method has achieved certain results in some aspects, it still mainly focuses on the recognition of explicit behaviors, is difficult to capture those more hidden and potential abnormal intentions in terms of time and actions, and locate and warn them, and the application scope and effectiveness are relatively limited, and it is also difficult to guarantee the accuracy and timeliness of the warning.
[0126] Next, an abnormal behavior warning device based on intention-behavior fusion proposed in the embodiments of this application will be described with reference to the accompanying drawings.
[0127] Figure 6 It is a schematic structural diagram of an abnormal behavior warning device based on intention-behavior fusion according to the embodiments of this application.
[0128] As Figure 6 shown, the abnormal behavior warning device 10 based on intention-behavior fusion includes: an extraction module 100, a fusion module 200, and a warning module 300.
[0129] Among them, the extraction module 100 is used to obtain the original video of the target person, extract the behavior feature vector of each video frame of the target person in the original video, and generate a behavior feature sequence according to the behavior feature vector.
[0130] The fusion module 200 is configured to calculate the fused features of the behavioral features based on the behavioral feature sequence and the corresponding weights, so as to obtain the intention features of the target person according to the fused features.
[0131] The warning module 300 is configured to input the intention features into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and determine a hidden intention location result of the target person according to the abnormal prediction result, so as to generate an abnormal behavior warning message of the target person based on the hidden intention location result.
[0132] Optionally, in an embodiment of the present application, the extraction module 100 includes: an identification unit and an extraction unit.
[0133] Among them, the identification unit is configured to identify suspicious video segments in the target hidden intention video that meet preset conditions.
[0134] The extraction unit is configured to extract the behavioral feature vectors of each video frame of the target person in the original video based on the suspicious video segments by using a target video feature extraction model.
[0135] Optionally, in an embodiment of the present application, the warning module 300 includes: a collection unit, an assignment unit, and a calculation unit.
[0136] Among them, the collection unit is configured to collect the types of abnormal behaviors in the original video.
[0137] The assignment unit is configured to assign a suspicion score to at least one type of abnormal behavior in each frame of the original video.
[0138] The calculation unit is configured to calculate the total suspicion score of each frame according to the types of abnormal behaviors and the suspicion scores, perform normalization processing on the total suspicion score to obtain the normalized total suspicion score, so as to determine the hidden intention location result of the target person according to the normalized total suspicion score.
[0139] Optionally, in an embodiment of the present application, the hidden intention location result includes at least one of a start time prediction result, an end time prediction result, and an abnormality degree prediction result of the abnormal intention of the target person.
[0140] Optionally, in an embodiment of the present application, the warning module 300 includes: an analysis unit and a generation unit.
[0141] Among them, the analysis unit is configured to analyze the intention features on different time scales to capture the short-term behavior change information of the target person within a first preset duration in the future and the long-term behavior change information within a second preset duration in the future, where the second preset duration is greater than the first preset duration.
[0142] A generating unit, configured to generate a prediction result of the abnormal intention of the target person according to the short-term behavior change information and the long-term behavior change information.
[0143] It should be noted that the foregoing explanation of the embodiment of the abnormal behavior early warning method based on intention-behavior fusion is also applicable to the abnormal behavior early warning device based on intention-behavior fusion of this embodiment, and will not be elaborated here.
[0144] According to the abnormal behavior early warning device based on intention-behavior fusion proposed in the embodiments of the present application, the behavior characteristics can be extracted from the original video of the target person and fused with the intention characteristics of the target person, and finally the hidden intention positioning result of the target person can be obtained, which is convenient for generating abnormal behavior early warning information for the target person in a timely manner. Thus, the concept of time-intention positioning is realized, and by analyzing the past behavior sequences, the future abnormal behaviors and hidden behavior intentions can be predicted, significantly improving the accuracy of abnormal behavior detection and the ability of early warning. Further, the present application can also effectively fuse various behavior characteristics to enhance the prediction ability of the intention prediction model; at the same time, the present application also has high interpretability and operability. Through multi-level behavior annotation and fine-grained behavior analysis, a clearer abnormal behavior path and prediction basis can be provided, enabling security personnel to more accurately understand and respond to potential threats. Thus, it solves the problems that although the time-action positioning method in the related technology has achieved certain results in some aspects, it still mainly focuses on the recognition of explicit behaviors, and it is difficult to capture those more concealed and potential abnormal intentions in terms of time and action, and to locate and warn them, the application scope and effectiveness are relatively limited, and it is also difficult to guarantee the accuracy and timeliness of early warning.
[0145] Figure 7 The structural schematic diagram of the electronic device provided by the embodiments of the present application. The electronic device may include:
[0146] A memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.
[0147] When the processor 702 executes the program, it implements the abnormal behavior early warning method based on intention-behavior fusion provided in the foregoing embodiments.
[0148] Further, the electronic device further includes:
[0149] A communication interface 703, configured for communication between the memory 701 and the processor 702.
[0150] The memory 701 is used for storing a computer program executable on the processor 702.
[0151] The memory 701 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0152] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0153] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a chip, the memory 701, the processor 702, and the communication interface 703 can communicate with each other through an internal interface.
[0154] The processor 702 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0155] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned abnormal behavior warning method based on intention-behavior fusion is implemented.
[0156] The embodiments of the present application also provide a computer program product, including a computer program, and the computer program can run computer instructions, and when the computer instructions are executed by a processor, the abnormal behavior warning method based on intention-behavior fusion provided by the embodiments of the present application is implemented.
[0157] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0158] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0159] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0160] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0161] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0162] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above-described embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0163] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0164] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An abnormal behavior early warning method based on intention-behavior fusion, characterized in that: The following steps are involved: Acquire an original video of a target person, extract a behavior feature vector of each video frame of the target person in the original video, and generate a behavior feature sequence according to the behavior feature vector; Based on the behavior feature sequence and the corresponding weights, a fusion feature of the behavior feature is calculated to obtain the intention feature of the target person according to the fusion feature; The intention feature is input into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and the hidden intention positioning result of the target person is determined according to the abnormal prediction result, so as to generate abnormal behavior warning information of the target person based on the hidden intention positioning result.
2. The method according to claim 1, characterized in that The step of obtaining an original video of a target person and extracting a behavior feature vector of each video frame of the target person in the original video includes: Identify suspicious video clips that meet preset conditions in the target hidden intention video; The target video feature extraction model is modified using the suspicious video clip to obtain a final video feature extraction model, so as to use the final video feature extraction model to extract the behavior feature vector of the target person in each video frame of the original video.
3. The method according to claim 1, characterized in that The determining the hidden intention positioning result of the target person according to the abnormal prediction result includes: Collecting the types of abnormal behaviors in the original video; assigning a suspicion score to at least one type of abnormal behavior in each frame of the original video; The total suspicion score of each frame is calculated according to the abnormal behavior type and the suspicion score, and the total suspicion score is normalized to obtain the normalized total suspicion score, so as to determine the hidden intention positioning result of the target person according to the normalized total suspicion score.
4. The method according to claim 1, characterized in that: The hidden intention positioning result includes at least one of the start time prediction result, the end time prediction result and the abnormal degree prediction result of the abnormal intention of the target person.
5. The method according to claim 1, characterized in that The step of inputting the intention feature into a preset intention prediction model for analysis and generating an abnormal intention prediction result of the target person includes: Analyze the intention features at different time scales to capture the short-term behavior change information of the target person within a first preset time period in the future and the long-term behavior change information within a second preset time period in the future, wherein the second preset time period is greater than the first preset time period; The abnormal intention prediction result of the target person is generated according to the short-term behavior change information and the long-term behavior change information.
6. An abnormal behavior early warning device based on intention-behavior fusion, characterized in that: include: An extraction module, used to obtain an original video of a target person, extract a behavior feature vector of each video frame of the target person in the original video, and generate a behavior feature sequence according to the behavior feature vector; A fusion module, used for calculating the fusion feature of the behavior feature based on the behavior feature sequence and the corresponding weight, so as to obtain the intention feature of the target person according to the fusion feature; An early warning module is used to input the intention feature into a preset intention prediction model to generate an abnormal intention prediction result of the target person, and determine the hidden intention positioning result of the target person according to the abnormal prediction result, so as to generate abnormal behavior early warning information of the target person based on the hidden intention positioning result.
7. The device according to claim 6, characterized in that The extraction module comprises: An identification unit, used to identify suspicious video clips that meet preset conditions in the target hidden intention video; The extraction unit is used to modify the target video feature extraction model using the suspicious video clip to obtain a final video feature extraction model, so as to use the final video feature extraction model to extract the behavior feature vector of the target person in each video frame of the original video.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the abnormal behavior early warning method based on intention-behavior fusion as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the abnormal behavior early warning method based on intention-behavior fusion as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, it is used to implement the abnormal behavior early warning method based on intention-behavior fusion as described in any one of claims 1 to 5.