Abnormal behavior recognition method and device, medium and program product
By acquiring multi-frame images and combining multiple rounds of question text, the problem of low accuracy in abnormal behavior recognition in traditional visual deep learning models is solved, and more efficient abnormal behavior recognition is achieved.
Patent Information
- Application Number
- CN202510518368.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, behavior recognition methods based on traditional visual deep learning models usually adopt frame-by-frame independent analysis strategies, which cannot accurately reflect the behavior of characters, resulting in a low accuracy of abnormal behavior recognition.
By acquiring the multi-frame image of the target area, the first recognition result is determined, and in the case of abnormal behavior, the second recognition is performed in combination with the current frame image and the preset multi-round question text, and the detailed features are deeply explored using the timing information between the multi-frame images and the multi-round question and answer sessions.
The accuracy of abnormal behavior recognition is improved, and the accuracy of abnormal behavior recognition is effectively improved through the timing information of multi-frame images and the multi-round question and answer mechanism.
Smart Images

Figure CN120452059A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video surveillance technology in artificial intelligence, and in particular to an abnormal behavior recognition method, device, medium and program product. Background Art
[0002] With the acceleration of urbanization, the dramatic increase in public traffic, and the rapid development of network technology, various abnormal behavior incidents (such as violent conflicts, theft, terrorist attacks, and online fraud) have become a serious threat to public safety, social stability, and individual rights. Traditional manual monitoring methods are no longer able to cope with the massive amount of data and real-time requirements, making it difficult to efficiently and accurately identify abnormal behavior.
[0003] In existing technologies, traditional visual deep learning models (such as the YOLO target detection model) are used to independently analyze video streams or image sequences frame by frame, and their deep feature extraction and classification capabilities are used to detect, segment or identify objects, scenes or actions in each frame of the image, and ultimately output static recognition results for each frame (such as the presence or absence of abnormal behavior).
[0004] However, in the existing technology, behavior recognition methods based on traditional visual deep learning models usually adopt a strategy of independent frame-by-frame analysis, extracting static features of single-frame images through deep networks, which cannot accurately reflect human behavior, resulting in a low accuracy rate in abnormal behavior recognition. Summary of the Invention
[0005] The embodiments of the present application provide an abnormal behavior identification method, device, medium, and program product to solve the problem of low accuracy in identifying abnormal behavior.
[0006] In a first aspect, an embodiment of the present application provides a method for identifying abnormal behavior, comprising:
[0007] Acquire multiple frames of images of the target area, wherein the multiple frames of images include a current frame image and at least one frame image before the current frame;
[0008] Determine a first recognition result according to the multiple frames of image;
[0009] When the first recognition result indicates the presence of abnormal behavior, a second recognition result is determined based on the current frame image and preset multiple rounds of question texts, wherein the multiple rounds of question texts include multiple questions for inquiring about the characteristics of the person in the current frame image.
[0010] As an optional implementation manner, determining the first recognition result according to the multiple frames of images includes:
[0011] splicing the multiple frames of image into a fused image in chronological order;
[0012] The first recognition result is determined according to the fused image.
[0013] As an optional implementation manner, determining the first recognition result according to the fused image includes:
[0014] Dividing the fused image into a plurality of non-overlapping image blocks;
[0015] Extracting a feature vector of each of the image blocks, and extracting a context feature vector of the multiple frames of images based on the feature vectors of each of the image blocks;
[0016] The first recognition result is determined according to the context feature vector.
[0017] As an optional implementation manner, determining the first recognition result according to the fused image includes:
[0018] The fused image is input into a pre-trained first large visual model to obtain the first recognition result output by the first large visual model.
[0019] As an optional implementation manner, determining the second recognition result based on the current frame image and the preset multiple rounds of question texts includes:
[0020] The current frame image and the preset multiple rounds of question text are input into a pre-trained second visual large model to obtain the second recognition result output by the second visual large model.
[0021] As an optional implementation manner, inputting the current frame image and the preset multiple rounds of question text into a pre-trained second visual large model to obtain the second recognition result output by the second visual large model includes:
[0022] The current frame image and the preset multiple-round question texts are input into a pre-trained second visual large model, and the second visual large model determines the reply text corresponding to the first question text based on the current frame image and the first question text in the multiple-round question texts. Then, the second visual large model determines the reply text corresponding to the next question text based on the current frame image, the next question text in the multiple-round question texts, and the answered question texts and corresponding reply texts, until the reply text corresponding to the last question text in the multiple-round question texts is determined, and the second recognition result is obtained according to the reply texts corresponding to the multiple-round question texts.
[0023] As an optional implementation manner, the first recognition result is used to indicate the presence of a first abnormal behavior, the presence of a second abnormal behavior, or the absence of an abnormal behavior; the penultimate question text in the multiple rounds of question texts is used to inquire whether the first abnormal behavior exists, and the last question text in the multiple rounds of question texts is used to inquire whether the second abnormal behavior exists;
[0024] Obtaining the second recognition result according to the answer texts corresponding to the multiple rounds of question texts includes:
[0025] The second recognition result is obtained based on the answer text corresponding to the second to last question text and the last question text in the multiple rounds of question texts, and the second recognition result is used to indicate the existence of the first abnormal behavior, the existence of the second abnormal behavior, or the absence of abnormal behavior.
[0026] In a second aspect, an embodiment of the present application provides an abnormal behavior recognition device, comprising:
[0027] An acquisition module, configured to acquire multiple frames of images of a target area, wherein the multiple frames of images include a current frame image and at least one frame image before the current frame;
[0028] a determination module, configured to determine a first recognition result based on the multiple frames of image;
[0029] The determination module is also used to determine a second recognition result based on the current frame image and preset multiple-round question texts when the first recognition result indicates the presence of abnormal behavior, wherein the multiple-round question texts include multiple questions for inquiring about the characteristics of the person in the current frame image.
[0030] As an optional implementation, the abnormal behavior identification device further includes: a processing module;
[0031] The processing module is used to splice the multiple frames of image into a fused image in chronological order;
[0032] The determination module is further configured to determine the first recognition result based on the fused image.
[0033] As an optional implementation, the processing module is further configured to divide the fused image into a plurality of non-overlapping image blocks;
[0034] The acquisition module is further configured to extract a feature vector of each image block, and based on the feature vector of each image block, extract a context feature vector of the multiple frames of image;
[0035] The determination module is further configured to determine the first recognition result according to the context feature vector.
[0036] As an optional implementation, the processing module is further configured to input the fused image into a pre-trained first large visual model to obtain the first recognition result output by the first large visual model.
[0037] As an optional implementation, the processing module is further used to input the current frame image and the preset multiple rounds of question text into a pre-trained second visual large model to obtain the second recognition result output by the second visual large model.
[0038] As an optional implementation, the determination module is also used to input the current frame image and the preset multiple-round question texts into a pre-trained second visual large model, and the second visual large model determines the reply text corresponding to the first question text based on the current frame image and the first question text in the multiple-round question texts, and then the second visual large model determines the reply text corresponding to the next question text based on the current frame image, the next question text in the multiple-round question texts, and the answered question text and the corresponding reply text, until the reply text corresponding to the last question text in the multiple-round question texts is determined, and the second recognition result is obtained according to the reply text corresponding to the multiple-round question texts.
[0039] As an optional implementation, the processing module is further configured to: the first recognition result is used to indicate the presence of a first abnormal behavior, the presence of a second abnormal behavior, or the absence of an abnormal behavior; the penultimate question text in the multiple rounds of question texts is used to inquire whether the first abnormal behavior exists, and the last question text in the multiple rounds of question texts is used to inquire whether the second abnormal behavior exists;
[0040] The determining module is further configured to obtain the second recognition result based on the answer texts corresponding to the multiple rounds of question texts, including:
[0041] The second recognition result is obtained based on the answer text corresponding to the second to last question text and the last question text in the multiple rounds of question texts, and the second recognition result is used to indicate the existence of the first abnormal behavior, the existence of the second abnormal behavior, or the absence of abnormal behavior.
[0042] In a third aspect, an embodiment of the present application provides an abnormal behavior recognition device, comprising: a receiver, a transmitter, a memory, and a processor;
[0043] A receiver, for receiving instructions and data;
[0044] Transmitter, used to send instructions and data;
[0045] The memory stores computer-executable instructions;
[0046] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.
[0047] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.
[0048] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0049] The abnormal behavior recognition method provided in the present application obtains multiple frames of images of a target area, wherein the multiple frames include a current frame image and at least one frame image before the current frame, determines a first recognition result based on the multiple frames, and when the first recognition result indicates the presence of abnormal behavior, determines a second recognition result based on the current frame image and preset multiple rounds of question text, wherein the multiple rounds of question text include multiple questions for inquiring about the characteristics of the person in the current frame image; the method effectively improves the accuracy of abnormal behavior recognition by analyzing multiple frames of images, performing two behavior recognitions, and conducting multiple rounds of question and answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0051] Figure 1 A flowchart of the abnormal behavior identification method provided in this application;
[0052] Figure 2 A schematic diagram of the structure of the abnormal behavior identification device provided in this application;
[0053] Figure 3 This is a schematic diagram of the structure of the abnormal behavior identification device provided in this application.
[0054] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0056] In today's complex and ever-changing social environment, with the acceleration of urbanization, the surge in public traffic, and the rapid development of network technology, incidents of unusual behavior, such as violent conflict, theft, terrorist attacks, and online fraud, are becoming increasingly frequent, posing a serious threat to public safety, social stability, and individual rights. Faced with the demands of massive amounts of data and real-time processing, traditional manual monitoring methods are no longer sufficient to efficiently and accurately identify these unusual behaviors.
[0057] Existing technologies use traditional deep learning models (such as the YOLO object detection model) to independently analyze each frame in a video stream or image sequence. Leveraging their powerful deep feature extraction and classification capabilities, they can detect, segment, or identify objects, scenes, or actions within each frame. The final output is a static recognition result based on a single frame (e.g., determining whether the frame contains abnormal behavior).
[0058] However, existing behavior recognition methods based on traditional visual deep learning models usually adopt a strategy of independent frame-by-frame analysis. Although this method can effectively extract static features in a single-frame image, it cannot accurately reflect human behavior, resulting in a low accuracy rate in abnormal behavior recognition.
[0059] To address the above issues, the abnormal behavior recognition method provided in this application obtains multiple frames of images within the target area, including the current frame and at least one previous frame, and first performs preliminary behavior recognition based on these images to determine a first recognition result. If possible abnormal behavior is detected, a more detailed analysis is performed by combining the current frame image with multiple rounds of preset question texts, which are specifically designed to inquire about the characteristics of the person in the current frame. This method not only utilizes the temporal information between multiple frames to capture dynamic changes, but also deeply explores detailed features through two rounds of behavior recognition and targeted multiple rounds of question and answer sessions, thereby improving the accuracy of abnormal behavior recognition.
[0060] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0061] Figure 1 This is a flow chart of the abnormal behavior identification method provided by this application. The execution subject of this embodiment is, for example, an abnormal behavior identification system. Figure 1 As shown, the method includes:
[0062] S101: Acquire multiple frames of images of a target area, where the multiple frames of images include a current frame image and at least one frame image before the current frame.
[0063] The target area refers to the spatial scope of surveillance designated when analyzing abnormal behavior. For example, in urban security management, a core area is defined as a high-risk crime zone (such as the area around subway stations or secluded alleyways), while a buffer zone is defined as a 1-kilometer radius of surrounding commercial and residential areas. The core area and buffer zone together constitute the target area.
[0064] Multiple frames of images form a time series. The current frame is the latest acquired image, representing the state of the target area at the current moment; and at least one frame of image before the current frame records the different states of the target area over a period of time in the past.
[0065] Understandably, a single-frame image can only provide static information about the target area at a specific moment, which is far from sufficient for many tasks that require analyzing the target's behavior and motion state. For example, in surveillance scenarios, it is difficult to determine whether a person is walking normally or suddenly running based on just a single frame. Multi-frame images, on the other hand, contain the state changes of the target area at different moments. By analyzing the differences between these images, dynamic features such as the target's motion trajectory, speed, and acceleration can be extracted, allowing for more accurate identification of the target's behavior. Secondly, multi-frame images help eliminate noise and interference. By integrating information from multiple frames, the correlation between images can be exploited to reduce the impact of noise, improving the robustness of target recognition.
[0066] For example, a multi-frame image is a three-frame image, and the three-frame image is ,in, is the current frame, and It's the first two frames.
[0067] S102: Determine a first recognition result based on the multiple frames of images.
[0068] Among them, the first recognition result is a preliminary judgment result output after comprehensively analyzing the target features, behavior patterns, dynamic changes and other information in multiple frames of images to determine whether the target object has abnormal behavior.
[0069] It's understandable that using temporal image sequences rather than single images can improve recognition accuracy and robustness by comparing inter-frame differences, tracking target trajectories, or capturing dynamic features. For example, in surveillance scenarios, an abnormal behavior recognition system can identify abnormal behavior (such as falls) by analyzing multiple frames of pedestrians walking continuously, rather than relying solely on static features in a single frame.
[0070] While a single frame image can only reflect a momentary state, multiple frames can capture the target's motion trajectory (e.g., behavioral changes and displacement). Inter-frame comparison can filter out background noise (e.g., sudden changes in illumination and occlusions) and enhance target features. Consecutive frames provide behavioral logic, improving the ability to understand complex scenes.
[0071] S103: When the first recognition result indicates the presence of abnormal behavior, determine a second recognition result based on the current frame image and a preset multi-round question text, wherein the multi-round question text includes a plurality of questions for inquiring about the characteristics of the person in the current frame image.
[0072] Among them, the second recognition result is the judgment result of integrating the visual information of the current frame image and the reasoning output of the question-answering logic.
[0073] After the first recognition result determines that abnormal behavior exists, the abnormal behavior recognition system starts the second recognition process: based on the current frame image and the preset multiple rounds of question text, the image content is further analyzed through the interactive question-and-answer mechanism to generate a more refined recognition conclusion.
[0074] When the first recognition result identifies abnormal behavior in the current frame, the Abnormal Behavior Recognition System automatically activates a pre-set multi-round question text, conducting an in-depth analysis of the image through structured question-answering. These multi-round questions focus on character features, covering dimensions such as "person presence," "interactional relationships," "posture characteristics," "objects held," and "nature of behavior." For example, questions like "Is there a person in the image?" "Is there someone on the ground?" "Is there a weapon?" are asked. The Abnormal Behavior Recognition System matches the current frame with the question text one by one, combining static visual features of the image (such as person outlines and spatial relationships) with semantic information from the question and answer (such as "non-upright posture" and "abnormal stillness") to construct a dynamic reasoning chain, gradually correcting any ambiguity or uncertainty in the first recognition result.
[0075] The abnormal behavior recognition method provided by the embodiment of the present application obtains multiple frames of images of the target area, where the multiple frames include the current frame image and at least one frame image before the current frame, determines a first recognition result based on the multiple frames, and when the first recognition result indicates the presence of abnormal behavior, determines a second recognition result based on the current frame image and preset multiple rounds of question text, wherein the multiple rounds of question text include multiple questions for inquiring about the characteristics of the person in the current frame image; the method effectively improves the accuracy of abnormal behavior recognition by analyzing multiple frames of images, performing two behavior recognitions, and conducting multiple rounds of question and answer.
[0076] As an optional implementation manner, determining the first recognition result according to the multiple frames of images includes:
[0077] Splice multiple frames of images into a fused image in chronological order from front to back;
[0078] A first recognition result is determined according to the fused image.
[0079] Multi-frame image fusion recognition involves splicing multiple consecutive frames in a time series into a single fused image in chronological order. This method integrates spatial and temporal information from these frames to generate a global visual feature representation. This method transcends the static limitations of single-frame images and leverages the dynamic relationships between adjacent frames (such as the movement of people and the deformation of objects) to construct a feature matrix encompassing both temporal and spatial dimensions.
[0080] Single-frame recognition methods are susceptible to interference from factors such as occlusion, lighting changes, and perspective deviation, leading to misjudgments or missed detections. For example, a target may be blurred in a single-frame image due to rapid movement. Multi-frame fusion, on the other hand, can eliminate instantaneous noise by integrating information redundancy in the temporal dimension. Accidental interference in a single frame (such as swaying leaves) will be smoothed out in the accumulation of multiple frames. Dynamic features are captured, and dynamic information such as changes in human posture and object motion trajectories are presented as continuous trajectories in the fused image, facilitating behavioral pattern recognition. Semantic associations are enhanced. After multiple frames are spliced together, the spatial-temporal relationships between people and the environment, and between people, are clarified. For example, the behavioral chain of "group gathering → conflict outbreak" can be intuitively reflected through position changes in the fused image. In the task of abnormal behavior detection, multi-frame fusion can improve recognition accuracy.
[0081] For example, a multi-frame image is three consecutive frames of images, and the time sequence of the three consecutive frames of images is ,in, is the current frame, and The first two frames are stitched together in order from front to back to form a fused image. , thereby determining whether the first recognition result is abnormal behavior or normal behavior.
[0082] As an optional implementation, determining the first recognition result according to the fused image includes:
[0083] Divide the fused image into multiple non-overlapping image blocks;
[0084] Extracting a feature vector of each image block, and extracting a context feature vector of multiple frames of images based on the feature vectors of each image block;
[0085] A first recognition result is determined according to the context feature vector.
[0086] First, a partitioning operation is performed on the fused image to divide it into multiple non-overlapping image blocks; then, feature extraction is performed on each image block to obtain the feature vector of each image block, and further based on these feature vectors, the context feature vectors of multiple frames of images are extracted; finally, the first recognition result is determined based on the extracted context feature vectors.
[0087] The fused image is divided into multiple non-overlapping image blocks. Because different areas of an image may contain different information, this division allows for more detailed capture of local features, preventing local information from being obscured during overall processing. Extracting the feature vector of each image block quantifies the key information in that block, making subsequent processing more convenient and efficient. Contextual feature vectors are extracted for multiple frames based on the feature vectors of each image block, as contextual information is crucial for image recognition. Multiple frames often exhibit a certain degree of correlation and continuity. Contextual feature vectors can comprehensively reflect the semantic relationships and overall information between these images, helping to improve the accuracy and reliability of recognition results.
[0088] For example, in an intelligent security surveillance scenario, an abnormal behavior recognition system receives multiple frames of continuously captured surveillance video. It first fuses these images to produce a single fused image, which integrates information from multiple frames and facilitates a more comprehensive analysis of the scene. The abnormal behavior recognition system then divides this fused image into multiple non-overlapping image blocks, such as 10x10 square blocks of equal size. For each block, a deep learning model is used to extract feature vectors. These feature vectors represent the visual information within the block, such as the object's shape and texture. The abnormal behavior recognition system then integrates the feature vectors of each block and analyzes their spatial and temporal relationships (if there is temporal information between the multiple frames) to extract a contextual feature vector for the multiple frames. This vector incorporates global information about the entire scene. Finally, the abnormal behavior recognition system inputs this contextual feature vector into a classifier, which, after calculation, determines the first recognition result, such as the presence of "abnormal behavior" in the surveillance scene.
[0089] As an optional implementation, determining the first recognition result according to the fused image includes:
[0090] The fused image is input into a pre-trained first visual model to obtain a first recognition result output by the first visual model.
[0091] For example, the first large visual model is the ViT (Vision Transformer) model.
[0092] The first visual big model is trained based on the ViT architecture to obtain a pre-trained first visual big model. The specific method is as follows: First, collect and annotate an image dataset containing abnormal behavior. Each image sample contains three consecutive frames of images as a time series input. Let each sample be ,in, is the current frame, and are the previous two frames. For each time series , mark whether the sequence contains abnormal behavior. Splice the three frames of images in each time series in order from front to back to form an image , the image size changes from the original Became .
[0093] Next, for each frame sequence , using the ViT model to extract visual features. The ViT model first divides each image into N non-overlapping image blocks and converts these blocks into N D-dimensional feature vectors using a linear mapping module. These feature vectors are then input into the Transformer Encoder layer, which uses a self-attention mechanism to process the spatial information in the image and extract contextual features from previous and next frames. This process integrates local and global image information and captures temporal correlations.
[0094] Finally, the features extracted by the Transformer Encoder layer are input into the Multi-Layer Perceptron Head (MLP Head) for classification. The MLP Head further processes the features and outputs a classification result. To optimize model performance, the training objective is set to minimize the cross-entropy loss function. By continuously adjusting the model parameters, the output classification result is made as close as possible to the true label, thus obtaining the pre-trained first visual large model.
[0095] It can be understood that the fused image formed by sequentially splicing a time sequence containing three consecutive frames of images is input into the pre-trained first large visual model. First, the input fused image is divided into N non-overlapping image blocks and converted into a D-dimensional feature vector sequence through linear mapping. Then, the Transformer Encoder layer is used to capture the local and global spatial features of the image and the temporal context correlation between the three frames through the self-attention mechanism. Finally, the encoded feature vector is input into the MLP Head for classification prediction, and finally, the first recognition result is output.
[0096] As an optional implementation, determining a second recognition result based on the current frame image and multiple rounds of preset question texts includes:
[0097] The current frame image and the preset multiple rounds of question text are input into the pre-trained second visual large model to obtain a second recognition result output by the second visual large model.
[0098] For example, the second visual large model is the MiniCPM-V model.
[0099] The specific method for pre-training the large Second Vision model is as follows: First, it involves data collection and labeling. Using devices such as surveillance cameras and web crawlers, we collect a large amount of diverse image and video data. These images and videos need to cover a wide range of scenarios and behavioral patterns so that the model can learn a wide range of behavioral characteristics. In particular, during this stage, we ensure that the dataset includes positive and negative examples of abnormal behavior, which helps the model more accurately identify different types of normal and abnormal behavior.
[0100] The MiniCPM-V model architecture consists of three main components: a vision module, a language module, and a cross-modal projection layer. The vision module specifically analyzes image content, such as extracting information about an object's shape, color, or motion. The language module focuses on understanding the meaning of textual descriptions, such as interpreting the specific context of the word "fall." The cross-modal projection layer matches visual and language information, enabling it to associate an image depicting a certain behavior (such as punching) with a corresponding textual description (such as fighting). During the fine-tuning phase, to optimize the MiniCPM-V model's performance for a specific task, the vision and language modules are frozen, and only the parameters of the cross-modal projection layer are adjusted. By leveraging previously collected and annotated image data, the model output is gradually refined to closely match the actual labeling results. This approach not only improves the model's ability to understand specific behaviors but also enhances its accuracy in real-world applications.
[0101] As you can understand, when processing the current frame, it is fed into the pre-trained Second Vision Large Model along with the pre-set multiple-round question text. This leverages the knowledge the MiniCPM-V model has previously acquired through training on a large amount of diverse image and video data, including its ability to understand various scenarios and behavioral patterns. The multiple-round question text is a sequence of questions designed to guide the MiniCPM-V model for more accurate behavior recognition. These questions can be descriptive questions about specific objects, actions, or scenes in the image.
[0102] Once the input is complete, the Second Vision Large Model analyzes and processes this information based on its internal structure—comprising a vision module, a language module, and a cross-modal projection layer. The vision module first analyzes the image content and extracts key features; simultaneously, the language module understands and processes the meaning of multiple rounds of question text. The cross-modal projection layer then combines the information from these two modules, matching the visual information in the image with the corresponding text description. Finally, after calculation and transformation, the Second Vision Large Model outputs a recognition result, known as the second recognition result.
[0103] As an optional implementation, the current frame image and the preset multiple rounds of question text are input into a pre-trained second visual large model to obtain a second recognition result output by the second visual large model, including:
[0104] The current frame image and the preset multiple-round question texts are input into a pre-trained second visual large model. The second visual large model determines the reply text corresponding to the first question text based on the current frame image and the first question text in the multiple-round question texts. Then, the second visual large model determines the reply text corresponding to the next question text based on the current frame image, the next question text in the multiple-round question texts, the answered question texts and the corresponding reply texts, until the reply text corresponding to the last question text in the multiple-round question texts is determined, and the second recognition result is obtained according to the reply texts corresponding to the multiple-round question texts.
[0105] The current frame image and the pre-set multiple-round question text are input into a pre-trained second vision large model. Rather than outputting the entire result all at once, the model uses a step-by-step approach: First, based on the current frame image and the first question text in the multiple-round question text, the model determines the corresponding answer text. Next, based on the current frame image, the next question text, and the already answered question text and corresponding answer text, the model determines the answer text for the next question. This cycle continues until the final question text in the multiple-round question text is answered. Finally, all the answers are combined to produce the second recognition result.
[0106] For example, when processing a frame of surveillance video, the current frame and a series of pre-set question text are input into a pre-trained large second vision model. First, the model analyzes the current frame and the first question in the multiple rounds of questions, "Is there a person in the picture?" and determines the corresponding answer text. For example, if there is indeed a person in the image, the output answer may be "yes."
[0107] Next, the model uses the current frame image, the next question text (such as "Are there multiple people interacting in the image?"), and the question and answer obtained in the previous step (i.e., "Are there people in the image?" with the answer "yes") to further analyze and determine the answer to the second question. For example, if the image shows two people talking, the corresponding answer text might be "yes."
[0108] This process continues one by one, working through a pre-set list of questions until all are answered. For example, if the image shows someone bending over to pick something up, the response might be a specific description of the action.
[0109] Finally, when all questions have been answered, the responses are combined to form a comprehensive second recognition result. For example, by combining responses to questions like "Is there a person in the image?", "Are there multiple people interacting in the image?", and "What postures are the people in the image?", the final recognition result might be: "There are two people in the image, they are interacting, and one of them is bending over to pick up something." This multi-round question-and-answer approach not only allows for detailed analysis of image content but also effectively identifies complex scenarios and behavioral patterns, thereby improving the accuracy of abnormal behavior detection.
[0110] As an optional implementation, the first recognition result is used to indicate the presence of the first abnormal behavior, the presence of the second abnormal behavior, or the absence of the abnormal behavior; the penultimate question text in the multiple rounds of question texts is used to inquire whether the first abnormal behavior exists, and the last question text in the multiple rounds of question texts is used to inquire whether the second abnormal behavior exists;
[0111] The second recognition result is obtained based on the answer texts corresponding to the multiple rounds of question texts, including:
[0112] A second recognition result is obtained based on the answer text corresponding to the second-to-last question text and the last question text in multiple rounds of question texts. The second recognition result is used to indicate the existence of the first abnormal behavior, the existence of the second abnormal behavior, or the absence of abnormal behavior.
[0113] For example, if the first abnormal behavior is a violent behavior and the second abnormal behavior is a falling behavior, the penultimate question text in the multiple-round question text is used to inquire whether there is a violent behavior, and the last question text in the multiple-round question text is used to inquire whether there is a falling behavior.
[0114] First, the current frame image and these preset questions are input into the pre-trained second vision model.
[0115] Suppose that during a multi-round question-and-answer process, the answer to the penultimate question, "Is there any violence in the image?", is affirmative, indicating the presence of the first abnormal behavior: violence. Then, if the answer to the final question, "Is there a person falling in the image?", is also "yes," this indicates the presence of the second abnormal behavior: falling. In this case, combining these two answers, the second recognition result will indicate the presence of both the violent behavior and the falling behavior.
[0116] On the other hand, if the answer to both key questions is negative, meaning neither violent behavior nor falls are detected in the image, the final second recognition result will indicate no abnormal behavior. This means that none of the predefined abnormal behavior types were detected in the scenario covered by the current frame. This precise, targeted, multi-round question-and-answer approach effectively and accurately identifies and distinguishes different types of abnormal behavior in complex scenarios, thereby improving the efficiency and accuracy of monitoring and early warning systems.
[0117] Figure 2 This is a schematic diagram of the structure of the abnormal behavior identification device provided by this application, such as Figure 2 As shown, the abnormal behavior identification device 200 provided in this embodiment includes:
[0118] An acquisition module 201 is configured to acquire multiple frames of images of a target area, where the multiple frames of images include a current frame and at least one frame of image before the current frame.
[0119] A determination module 202 is configured to determine a first recognition result based on multiple frames of images;
[0120] The determination module 202 is also used to determine a second recognition result based on the current frame image and preset multiple-round question texts when the first recognition result indicates the presence of abnormal behavior, wherein the multiple-round question texts include multiple questions for inquiring about the characteristics of the person in the current frame image.
[0121] As an optional implementation, the abnormal behavior identification device further includes: a processing module 203;
[0122] The processing module 203 is used to stitch multiple frames of images into a fused image in a time sequence from the front to the back;
[0123] The determination module 202 is further configured to determine a first recognition result according to the fused image.
[0124] As an optional implementation, the processing module 203 is further configured to divide the fused image into a plurality of non-overlapping image blocks;
[0125] The acquisition module 201 is further configured to extract a feature vector of each image block, and based on the feature vectors of each image block, extract a context feature vector of multiple frames of images;
[0126] The determination module 202 is further configured to determine a first recognition result according to the context feature vector.
[0127] As an optional implementation, the processing module 203 is further configured to input the fused image into a pre-trained first large visual model to obtain a first recognition result output by the first large visual model.
[0128] As an optional implementation, the processing module 203 is further configured to input the current frame image and the preset multiple rounds of question texts into a pre-trained second visual large model to obtain a second recognition result output by the second visual large model.
[0129] As an optional implementation, the determination module 202 is also used to input the current frame image and the preset multiple-round question texts into a pre-trained second visual large model, and the second visual large model determines the reply text corresponding to the first question text based on the current frame image and the first question text in the multiple-round question texts, and then the second visual large model determines the reply text corresponding to the next question text based on the current frame image, the next question text in the multiple-round question texts, and the answered question texts and the corresponding reply texts, until the reply text corresponding to the last question text in the multiple-round question texts is determined, and the second recognition result is obtained according to the reply texts corresponding to the multiple-round question texts.
[0130] As an optional implementation, the processing module 203 is further configured to: the first recognition result is used to indicate the presence of the first abnormal behavior, the presence of the second abnormal behavior, or the absence of the abnormal behavior; the second-to-last question text in the multiple rounds of question texts is used to inquire whether the first abnormal behavior exists, and the last question text in the multiple rounds of question texts is used to inquire whether the second abnormal behavior exists;
[0131] The determination module 202 is further configured to obtain a second recognition result based on the answer texts corresponding to the multiple rounds of question texts, including:
[0132] A second recognition result is obtained based on the answer text corresponding to the second-to-last question text and the last question text in multiple rounds of question texts. The second recognition result is used to indicate the existence of the first abnormal behavior, the existence of the second abnormal behavior, or the absence of abnormal behavior.
[0133] Figure 3 This is a schematic diagram of the structure of the abnormal behavior recognition device provided by this application. Figure 3 As shown, the present application provides an abnormal behavior identification device, which includes: a receiver 301, a transmitter 302, a processor 303 and a memory 304.
[0134] Receiver 301, for receiving instructions and data;
[0135] Transmitter 302, used to send instructions and data;
[0136] Memory 304, for storing computer-executable instructions;
[0137] The processor 303 is configured to execute the computer-executable instructions stored in the memory 304 to implement the various steps of the abnormal behavior identification method in the above embodiment. For details, please refer to the relevant description in the above abnormal behavior identification method embodiment.
[0138] Optionally, the memory 304 may be independent or integrated with the processor 303 .
[0139] When the memory 304 is independently provided, the electronic device further includes a bus for connecting the memory 304 and the processor 303 .
[0140] The present application also provides a computer-readable storage medium, which stores computer-executable instructions. When a processor executes the computer-executable instructions, the abnormal behavior recognition method performed by the abnormal behavior recognition device described above is implemented.
[0141] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units in the methods, systems, and devices disclosed above may be implemented as software, firmware, hardware, or any combination thereof. In hardware implementations, the division between functional modules / units described above does not necessarily correspond to the division between physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is well known to those skilled in the art, the term computer storage media encompasses both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0142] So far, the technical solution of the present application has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the scope of protection of the present application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solution of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solution to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying abnormal behavior, characterized in that: include: Acquire multiple frames of images of the target area, wherein the multiple frames of images include a current frame image and at least one frame image before the current frame; Determine a first recognition result according to the multiple frames of image; When the first recognition result indicates the presence of abnormal behavior, a second recognition result is determined based on the current frame image and preset multiple rounds of question texts, wherein the multiple rounds of question texts include multiple questions for inquiring about the characteristics of the person in the current frame image.
2. The method according to claim 1, characterized in that Determining a first recognition result according to the multiple frames of image includes: splicing the multiple frames of image into a fused image in chronological order; The first recognition result is determined according to the fused image.
3. The method according to claim 2, characterized in that Determining the first recognition result according to the fused image includes: Dividing the fused image into a plurality of non-overlapping image blocks; Extracting a feature vector of each of the image blocks, and extracting a context feature vector of the multiple frames of images based on the feature vectors of each of the image blocks; The first recognition result is determined according to the context feature vector.
4. The method according to claim 2, characterized in that Determining the first recognition result according to the fused image includes: The fused image is input into a pre-trained first large visual model to obtain the first recognition result output by the first large visual model.
5. The method according to any one of claims 1 to 4, characterized in that Determining a second recognition result based on the current frame image and the preset multiple rounds of question texts includes: The current frame image and the preset multiple rounds of question text are input into a pre-trained second visual large model to obtain the second recognition result output by the second visual large model.
6. The method according to claim 5, characterized in that Inputting the current frame image and the preset multiple rounds of question text into a pre-trained second visual large model to obtain the second recognition result output by the second visual large model includes: The current frame image and the preset multiple-round question texts are input into a pre-trained second visual large model, and the second visual large model determines the reply text corresponding to the first question text based on the current frame image and the first question text in the multiple-round question texts. Then, the second visual large model determines the reply text corresponding to the next question text based on the current frame image, the next question text in the multiple-round question texts, and the answered question texts and corresponding reply texts, until the reply text corresponding to the last question text in the multiple-round question texts is determined, and the second recognition result is obtained according to the reply texts corresponding to the multiple-round question texts.
7. The method according to claim 6, characterized in that The first recognition result is used to indicate the presence of a first abnormal behavior, the presence of a second abnormal behavior, or the absence of an abnormal behavior; the penultimate question text in the multiple rounds of question texts is used to inquire whether the first abnormal behavior exists, and the last question text in the multiple rounds of question texts is used to inquire whether the second abnormal behavior exists; Obtaining the second recognition result according to the answer texts corresponding to the multiple rounds of question texts includes: The second recognition result is obtained based on the answer text corresponding to the second to last question text and the last question text in the multiple rounds of question texts, and the second recognition result is used to indicate the existence of the first abnormal behavior, the existence of the second abnormal behavior, or the absence of abnormal behavior.
8. An abnormal behavior recognition device, characterized in that: include: An acquisition module, configured to acquire multiple frames of images of a target area, wherein the multiple frames of images include a current frame image and at least one frame image before the current frame; a determination module, configured to determine a first recognition result based on the multiple frames of image; The determination module is also used to determine a second recognition result based on the current frame image and preset multiple-round question texts when the first recognition result indicates the presence of abnormal behavior, wherein the multiple-round question texts include multiple questions for inquiring about the characteristics of the person in the current frame image.
9. An abnormal behavior recognition device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Audio and video intelligent processing method and device, storage medium and electronic equipment
CN117633604A
Calling behavior identification method, device and equipment, medium and product
CN117994855A
Behavior recognition method and device based on multi-modal large model and electronic equipment
CN118314624A
Out-of-season dressing identification method and device, electronic equipment and storage medium
CN118762303A
Abnormal behavior detection method, system, device and medium
CN119785416A
Cited By
Cleaning method and device, cleaning robot and readable storage medium
CN120783394A