Multimodal Behavior Analysis Method, Device, Equipment and Storage Medium

By cutting the image to be analyzed into multiple sub-maps to be analyzed in the monitoring and security scenario, and performing feature extraction and pooling operations, the analysis time and information loss problems caused by the small proportion of human images in the scene are solved, and efficient multimodal behavior analysis is achieved.

CN119723679BActive Publication Date: 2025-06-10ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510225882.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-10
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In monitoring and security scenarios, human body images account for too little proportion in the overall scene image, resulting in a long time to image analysis and when cutting independent human body images from scene images, it is easy to lose the topological and positional relationships between people, which leads to impaired information integrity, damage to space-time coherence, decreased analysis efficiency, increased risk of misjudgment and missed detection, and affected user experience.

Method used

By cutting the image to be analyzed, a plurality of sub-maps to be analyzed are obtained, and each sub-map includes a plurality of detection objects whose object distance is not greater than the set distance threshold. Then, feature extraction is performed on these subgraphs, local feature sets of each detection object are obtained, and the global behavioral characteristics of each detection object are obtained through pooling operations. Finally, based on the object behavior text and global behavior characteristics, the target object in the image that performs a specific object behavior is determined.

Benefits of technology

This method can perform behavioral analysis through small image scales, which can not only ensure the detection effect, but also retain the feasibility of analyzing complex problems of multi-person interactions, reduce time-consuming, achieve a balance between time-consuming and effect, and make up for the information loss during image segmentation through the fusion of local features of multiple perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723679B_ABST
    Figure CN119723679B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and provides a multi-modal behavior analysis method, apparatus, device and storage medium. The method includes: cutting the image to be analyzed to obtain a plurality of sub-images to be analyzed. A relatively small image scale not only ensures the detection effect, but also eliminates the interference of large areas where the detection object completely does not exist, balancing the time consumption and the effect; extracting features from the plurality of sub-images to be analyzed to obtain a plurality of local feature maps. From the plurality of local feature maps, local feature sets of each detection object are respectively obtained, and then by pooling each local feature set, global behavior features of each detection object are obtained, reconstructing global information based on multi-perspective local information to make up for the information loss caused by splitting the large-size image to be analyzed into small-size sub-images to be analyzed; finally, based on the object behavior text describing the behavior of a specific object and the global behavior features of each detection object, the target object performing the specific object behavior in the image to be analyzed is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Multi-modal behavior analysis is an important sub-topic at the intersection of computer vision and natural language processing. Multi-modal refers to the simultaneous use of multiple types of data or signals (such as images, videos, texts, audio, etc.) for information processing, analysis, and generation. In recent years, with the rapid development of multi-modal technologies represented by Large Language and Vision Assistant (LLAVA), researchers have seen great potential in open-set human behavior analysis based on multi-modal large models.

[0003] Although multi-modal technologies represented by LLAVA have developed rapidly and performed well in behavior analysis, applying them to surveillance and security scenarios still faces certain challenges. During the actual implementation of these technologies, two major challenges caused by the too-small proportion of human images in the overall scene image need to be overcome: one is the long time-consuming for image analysis; the other is that when simply cutting out independent human images from the overall scene image, it is easy to lose the topological and positional relationships between people, which in turn brings a series of problems, including damaged information integrity, disrupted spatio-temporal coherence, decreased analysis efficiency, increased risks of misjudgment and missed detection, and affected user experience. These problems not only limit the effectiveness of multi-modal technologies in practical applications but also increase the difficulty of system design and optimization. Summary of the Invention

[0004] Embodiments of this application provide a multi-modal behavior analysis method, apparatus, device, and storage medium to solve a series of problems caused by the too-small proportion of human images in the scene images of surveillance and security scenarios.

[0005] In a first aspect, embodiments of this application provide a multi-modal behavior analysis method, including:

[0006] Cut the image to be analyzed to obtain multiple sub-images to be analyzed, where each sub-image to be analyzed includes multiple detection objects with an object distance not greater than a set distance threshold;

[0007] Extract features from the multiple sub-images to be analyzed to obtain multiple local feature maps;

[0008] Respectively obtain the local feature sets of each detection object from the multiple local feature maps, and obtain the global behavior features of each detection object by pooling the local feature sets;

[0009] Based on the object behavior text describing the behavior of a specific object and the global behavior features of each detection object, obtain the target object performing the specific object behavior in the image to be analyzed.

[0010] Optionally, the step of cutting the image to be analyzed to obtain a plurality of sub-images to be analyzed includes:

[0011] Performing object detection on the image to be analyzed to obtain an object detection box for each detected object;

[0012] Grouping the detected objects based on the object detection boxes to obtain a plurality of object groups;

[0013] Cutting a plurality of detected objects in the image to be analyzed that are in the same object group into one sub-image to be analyzed, thereby obtaining a plurality of sub-images to be analyzed.

[0014] Optionally, the step of grouping the detected objects based on the object detection boxes to obtain a plurality of object groups includes traversing the detected objects in a loop, and performing the following operations for each traversed detected object:

[0015] When it is determined that the one detected object is not grouped, expanding the object detection box of the one detected object outward by a first set step length to obtain a first expanded detection box;

[0016] Dividing the one detected object and other detected objects that intersect with the first expanded detection box into one object group.

[0017] Optionally, the step of cutting a plurality of detected objects in the image to be analyzed that are in the same object group into one sub-image to be analyzed includes:

[0018] Expanding the first expanded detection box of the one detected object outward by a second set step length to obtain a second expanded detection box, where the second expanded detection box covers other detected objects that are in the same object group as the one detected object;

[0019] Cutting the image to be analyzed along the second expanded detection box to obtain one sub-image to be analyzed.

[0020] Optionally, the step of respectively obtaining a local feature set for each detected object from a plurality of local feature maps includes:

[0021] Respectively obtaining the local behavior features of each detected object in the plurality of local feature maps;

[0022] Dividing the local behavior features associated with the same object identifier into one local feature set to obtain a local feature set for each detected object.

[0023] Optionally, the step of obtaining a target object that performs a specific object behavior in the image to be analyzed based on the object behavior text describing the specific object behavior and the global behavior features of each detected object includes:

[0024] Extract features from the object behavior text that describes the behavior of a specific object to obtain text features;

[0025] Based on the text features and the global behavior features of each detection object, obtain the feature correlation degree of each detection object, and determine the detection objects whose feature correlation degree exceeds the set correlation degree threshold as the target objects that have performed the specific object behavior.

[0026] In a second aspect, an embodiment of the present application further provides a multimodal behavior analysis device, including:

[0027] An image cutting unit, configured to cut the image to be analyzed to obtain a plurality of sub-images to be analyzed, and a plurality of detection objects with an object distance not greater than a set distance threshold are included in one sub-image to be analyzed;

[0028] A feature fusion unit, configured to extract features from a plurality of sub-images to be analyzed to obtain a plurality of local feature maps;

[0029] Respectively obtain the local feature sets of each detection object from a plurality of local feature maps, and obtain the global behavior features of each detection object by pooling each local feature set;

[0030] A behavior analysis unit, configured to obtain the target objects that perform the specific object behavior in the image to be analyzed based on the object behavior text that describes the specific object behavior and the global behavior features of each detection object.

[0031] Optionally, the image cutting unit is configured to:

[0032] Perform object detection on the image to be analyzed to obtain the object detection frames of each detection object;

[0033] Based on each object detection frame, group the detection objects to obtain a plurality of object groups;

[0034] Cut the multiple detection objects located in the same object group in the image to be analyzed into one sub-image to be analyzed, and obtain a plurality of sub-images to be analyzed.

[0035] Optionally, the image cutting unit traverses each detection object in a loop, and performs the following operations for each traversed detection object:

[0036] When it is determined that the one detection object is not grouped, expand the object detection frame of the one detection object outward by a first set step length to obtain a first expanded detection frame;

[0037] Group the one detection object and other detection objects that intersect with the first expanded detection frame into one object group.

[0038] Optionally, the image cutting unit is configured to:

[0039] Expand the first externally extended detection frame of the one detection object outward by a second set step length to obtain a second externally extended detection frame, where the second externally extended detection frame covers other detection objects in the same object group as the one detection object;

[0040] Cut the image to be analyzed along the second externally extended detection frame to obtain a sub-image to be analyzed.

[0041] Optionally, the feature fusion unit is used for:

[0042] Obtain the local behavior features of each detection object in multiple local feature maps respectively;

[0043] Divide the local behavior features associated with the same object identifier into a local feature set to obtain the local feature set of each detection object.

[0044] Optionally, the behavior analysis unit is used for:

[0045] Extract features from the object behavior text describing the behavior of a specific object to obtain text features;

[0046] Based on the text features and the global behavior features of each detection object, obtain the feature correlation degree of each detection object, and determine the detection object whose feature correlation degree exceeds the set correlation degree threshold as the target object that has performed the specific object behavior.

[0047] In a third aspect, an embodiment of the present application further provides a computer device, including a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of any one of the above multi-modal behavior analysis methods.

[0048] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of any one of the above multi-modal behavior analysis methods.

[0049] In a fifth aspect, an embodiment of the present application further provides a computer program product, including computer instructions, and the computer instructions are executed by the processor to perform the steps of any one of the above multi-modal behavior analysis methods.

[0050] The beneficial effects of the present application are as follows:

[0051] An embodiment of the present application provides a multi-modal behavior analysis method, device, equipment, and storage medium. The method includes: cutting the image to be analyzed to obtain multiple sub-images to be analyzed, where each sub-image to be analyzed includes multiple detected objects with object distances not greater than a set distance threshold; extracting features from the multiple sub-images to be analyzed to obtain multiple local feature maps, respectively obtaining the local feature sets of each detected object from the multiple local feature maps, and then pooling each local feature set to obtain the global behavior features of each detected object; finally, based on the object behavior text describing the behavior of a specific object and the global behavior features of each detected object, obtaining the target object that performs the specific object behavior in the image to be analyzed.

[0052] In the embodiment of the present application, behavior analysis is performed in units of sub-images to be analyzed. A relatively small image scale can not only ensure the detection effect but also retain the feasibility of theoretically performing behavior analysis on complex problems of multi-person interaction. There is no need to worry about whether to expand the small-size image of a single detected object and how much to expand, and it also eliminates the interference of large areas without detected objects on the behavior analysis effect. Moreover, in the embodiment of the present application, a single CLIP forward is performed on the sub-images to be analyzed to extract the local behavior features of multiple people, which can greatly reduce the time consumption and achieve a balance between time consumption and effect.

[0053] The same detected object may exist in multiple sub-images to be analyzed. The feature fusion module realizes the reconstruction of global information based on multi-view local information through multi-view local feature fusion, compensating for the information loss caused by cutting the large-size image to be analyzed into small-size sub-images.

[0054] Other features and advantages of the present application will be described in the subsequent description, and some of them will become obvious from the description or be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0056] Figure 1 It is a schematic diagram of the architecture of the multi-modal behavior analysis model provided by the embodiment of the present application;

[0057] Figure 2 It is a schematic diagram of the interface of the data / label organization form and routing settings in the model training stage provided by the embodiment of the present application;

[0058] Figure 3AA schematic flowchart of multimodal behavior analysis for a monitoring and security scenario provided by an embodiment of the present application;

[0059] Figure 3B A schematic logic diagram of multimodal behavior analysis for a monitoring and security scenario provided by an embodiment of the present application;

[0060] Figure 3C A schematic flowchart of dividing each detection object into multiple object groups provided by an embodiment of the present application;

[0061] Figure 3D A schematic logic diagram of grouping a detection object provided by an embodiment of the present application;

[0062] Figure 3E A schematic logic diagram of cutting an image to be analyzed to obtain a sub-image to be analyzed provided by an embodiment of the present application;

[0063] Figure 3F A schematic logic diagram of the operation of a feature fusion module provided by an embodiment of the present application;

[0064] Figure 3G A schematic diagram of object behavior text provided by an embodiment of the present application;

[0065] Figure 4 A schematic structural diagram of a multimodal behavior analysis device provided by an embodiment of the present application;

[0066] Figure 5 A schematic diagram of a hardware composition structure of a computer device applying an embodiment of the present application;

[0067] Figure 6 A schematic diagram of a hardware composition structure of another computer device applying an embodiment of the present application. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments described in this application document without creative efforts fall within the scope of protection of the technical solutions of the present application.

[0069] The following briefly introduces the design concept of the embodiments of the present application:

[0070] Multi-modal behavior analysis is an important subtopic at the intersection of computer vision and natural language processing. Multi-modal refers to the simultaneous use of multiple types of data or signals (such as images, videos, texts, audio, etc.) for information processing, analysis, and generation. In recent years, with the rapid development of multi-modal technologies represented by LLAVA, researchers have seen great potential in open-set human behavior analysis based on multi-modal large models.

[0071] Although multi-modal technologies represented by LLAVA have developed rapidly and performed well in behavior analysis, applying them to surveillance and security scenarios still faces certain challenges. During the actual implementation of these technologies, two major challenges caused by the small proportion of human images in the overall scene image need to be overcome: one is the long time-consuming for image analysis; the other is that when simply cutting out independent human images from the overall scene image, it is easy to lose the topological and positional relationships between people, which in turn brings a series of problems, including damaged information integrity, disrupted spatio-temporal coherence, decreased analysis efficiency, increased risks of misjudgment and missed detection, and affected user experience. These problems not only limit the effectiveness of multi-modal technologies in practical applications but also increase the difficulty of system design and optimization.

[0072] Therefore, to solve the above problems, this application proposes a multi-modal behavior analysis method. The method includes: cutting the image to be analyzed to obtain multiple sub-images to be analyzed, where each sub-image to be analyzed includes multiple detection objects with an object distance not greater than a set distance threshold; extracting features from the multiple sub-images to be analyzed to obtain multiple local feature maps, and respectively obtaining the local feature sets of each detection object from the multiple local feature maps, and then pooling each local feature set to obtain the global behavior features of each detection object; finally, based on the object behavior text describing the behavior of a specific object and the global behavior features of each detection object, the target object performing the specific object behavior in the image to be analyzed is obtained.

[0073] The embodiments of this application perform behavior analysis in units of sub-images to be analyzed. The relatively small image scale can not only ensure the detection effect but also retain the feasibility of theoretically analyzing complex problems of multi-person interaction in behavior analysis. There is no need to worry about whether to expand the small-size image of a single detection object and how much to expand, and it also eliminates the interference of large areas without detection objects on the behavior analysis effect. Moreover, the embodiments of this application perform a single CLIP forward on the sub-images to be analyzed to extract the local behavior features of multiple people, which can significantly reduce the time consumption and achieve a balance between time consumption and effect.

[0074] The same detection object may exist in multiple sub-images to be analyzed. The feature fusion module realizes the reconstruction of global information based on multi-view local information through multi-view local feature fusion, compensating for the information loss caused by cutting the large-size image to be analyzed into small-size sub-images to be analyzed.

[0075] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0076] The embodiment of the present application designs a multi-modal behavior analysis model, as Figure 1 shown, mainly composed of an object detection sub-model, a greedy scheduling module, a Contrastive Language-Image Pre-training (CLIP) visual encoder, a projector, a feature fusion module, and a Large Language Model (LLM).

[0077] Among them, the open-source object detection sub-model is responsible for object detection on the large-size image to be analyzed in the monitoring and security scene, and determining the positions of each detected object in the image. The greedy scheduling module is responsible for dividing the detected objects with relatively close spatial positions in the large-size image to be analyzed in the monitoring and security scene into an object group, providing a cutting benchmark for cutting the image, and obtaining multiple small-size sub-images to be analyzed. The open-source CLIP visual encoder is responsible for extracting the visual representations of the sub-images to be analyzed. The projector is a simple linear layer or a Multilayer Perceptron (MLP), which is responsible for aligning the local behavior features to language features. The feature fusion module realizes the reconstruction of global information based on multi-view local information through the local feature fusion of multi-view regions, compensating for the information loss caused by cutting the large-size image to be analyzed into small-size sub-images to be analyzed. The segmented object behavior text is concatenated with the visual features aligned by the projector and sent into the LLM to obtain the corresponding behavior analysis results.

[0078] Different from using the images to be analyzed collected in the security monitoring scenario as the original form of data in the model application stage, the training samples in the model training stage are small-sized subgraphs to be analyzed. The behavior annotation results on the large-sized images to be analyzed are converted into the behavior annotation results on the small-sized subgraphs to be analyzed as the original form of the training samples. Meanwhile, in the training stage, a conversation may involve multiple subgraphs to be analyzed, which can be from the same image to be analyzed or from different images to be analyzed. Through the permutation and combination of the subgraphs to be analyzed, data augmentation of the training set is achieved. Behavior analysis is carried out on a per-person basis. When the same object appears in multiple subgraphs to be analyzed, by analyzing multiple subgraphs to be analyzed involved in the conversation and obtaining the local feature set of the object in the graph, then the local behavior features of the object that appears in multiple subgraphs to be analyzed also have multiple sources.

[0079] The open-source CLIP visual encoder only performs visual-text alignment at the image level and lacks region-level alignment. If feature fusion is performed on the local behavior features extracted by the native CLIP visual encoder, only suboptimal detection effects can be achieved. Therefore, based on the native LLAVA, the embodiments of this application refine the training process of the multi-modal LLM: (1) The first training stage is the same as the native LLAVA. Fix the CLIP visual encoder and the LLM, and only adjust the parameters of the projector to generally achieve the alignment from visual features to text features; (2) Add a second training stage. Fix the parameters of the LLM and adjust the parameters of the CLIP visual encoder and the projector to adapt the CLIP to the data in the security scenario field and, at the same time, achieve the purpose of optimizing its region-level feature extraction ability; (3) The third training stage is also the same as the native LLAVA. Fix the CLIP visual encoder and the projector and only adjust the parameters of the LLM.

[0080] For the data / label organization form and routing settings in the model training stage, please refer to Figure 2 Through the logical process set in the routing table, connect the output of the CLIP visual encoder with the input of the LLM. Any placeholder of the local behavior feature in the input text of the LLM can index to the corresponding local behavior feature in the local feature map associated with the relevant subgraph to be analyzed. Among them, Figure 2 the value of the first value field in represents the question, and the value of the second value field represents the answer. The content not wrapped by angle brackets <> in the question is the natural language form of the participating word segmentation, and the content wrapped by angle brackets is the placeholder of the local behavior feature. Before being sent into the LLM, it is replaced one by one with the local behavior features of the corresponding detected object in the subgraph to be analyzed. The id is the unique identifier of the human body area of the detected object.

[0081] To expand the detection range of the training set, data augmentation is performed on the training set. Private behavior analysis-related data is mixed with the public dataset, and special tags are used to distinguish public data from private data. In fact, there are other implementation methods for the logical process of connecting the CLIP visual encoder and the LLM, but they will not be elaborated here. After the above three-stage training, the multi-modal behavior analysis model for the monitoring and security scenario introduced in the embodiments of this application can be obtained.

[0082] Combined with Figures 3A - 3B , the process of performing multi-modal behavior analysis on the monitoring and security scenario based on the method provided in the embodiments of this application is introduced.

[0083] S301: Cut the image to be analyzed to obtain multiple sub-images to be analyzed. Each sub-image to be analyzed includes multiple detection objects whose object distances are not greater than the set distance threshold.

[0084] With the help of an open-source object detection sub-model, object detection is performed on the image to be analyzed to locate the positions of each detection object in the image, and the object detection box of each detection object is obtained. Then, with the help of the greedy scheduling module designed in the embodiments of this application, grouping is performed based on each object detection box, and each detection object is divided into multiple object groups. Multiple object groups with relatively close spatial positions are grouped into one object group. Therefore, the greedy scheduling module uses the group as the basic unit for cropping the image, and cuts multiple detection objects located in the same object group in the image to be analyzed into one sub-image to be analyzed, obtaining multiple sub-images to be analyzed.

[0085] Among them, combined with Figures 3C - 3D , the process of dividing each detection object into multiple object groups by using a cyclic iteration method is introduced.

[0086] S3011: Obtain the i-th detection object.

[0087] S3012: Determine whether the i-th detection object is in the grouped set. If so, jump to step 3011; otherwise, execute step 3013.

[0088] S3013: When it is determined that the i-th detection object is not grouped, expand the object detection box of the i-th detection object outward by the first set step length to obtain the first expanded detection box.

[0089] S3014: Divide the i-th detection object and other detection objects that intersect with the first expanded detection box into one object group, and add the i-th detection object and other detection objects that intersect with the first expanded detection box to the grouped set.

[0090] S3015: Determine whether all detection objects have been traversed. If so, obtain multiple object groups; otherwise, jump to step 3011.

[0091] In the embodiment of this application, the detection object and other detection objects that intersect with the first externally expanded detection frame are divided into an object group. As shown in Figure 3E , taking the group as the basic unit for cropping the image, the first externally expanded detection frame of the i-th detection object in the group is further expanded outward by a second set step length to obtain a second externally expanded detection frame. The second externally expanded detection frame covers other detection objects in the same object group as the i-th detection object. Then, the image to be analyzed is cut along the second externally expanded detection frame to obtain a sub-image to be analyzed.

[0092] Mark the relative positions of multiple detection objects in an object group in the corresponding sub-images to be analyzed. If the same detection object appears in multiple sub-images to be analyzed, the unique visual representation of this detection object in the future is the combination of the local behavior features extracted from the above-mentioned multiple sub-images to be analyzed.

[0093] S302: Extract features from multiple sub-images to be analyzed to obtain multiple local feature maps.

[0094] With the help of the open-source CLIP visual encoder, extract the visual representations of multiple sub-images to be analyzed. With the help of the projector, complete the alignment of local behavior features to language features to obtain the local feature maps corresponding to each sub-image to be analyzed.

[0095] A local feature map contains the local behavior features of each detection object in the image. This local behavior feature is also called the Region of Interest (ROI) feature, which refers to the area selected in the image for further analysis or processing. These areas are considered to contain more useful information than other parts of the image. Only need to focus on and analyze the interesting parts of the image, rather than the whole image. This is especially useful when processing large images. By using local behavior features, the processing time of the image can be significantly reduced, the computational complexity can be reduced, and the processing accuracy can be improved.

[0096] S303: Respectively obtain the local feature sets of each detection object from multiple local feature maps, and obtain the global behavior features of each detection object by pooling each local feature set.

[0097] Combined with Figure 3F , the feature fusion module respectively obtains the local behavior features of each detection object in multiple local feature maps, divides the local behavior features associated with the same object identifier into a local feature set, and obtains the local feature sets of each detection object. The feature fusion module performs global average pooling on each local feature set to obtain the global behavior features of each detection object. So far, each detection object in the large-size image to be analyzed is described by a unique global behavior feature, and the visual preparation work before feeding into the LLM has been completed.

[0098] S304: Based on the object behavior text describing the behavior of a specific object and the respective global behavior characteristics of each detected object, obtain the target object in the image to be analyzed that performs the specific object behavior.

[0099] As Figure 3G shown, the embodiment of the present application designs an object behavior text for describing the behavior of a specific object. The value of the first value field in this text represents the question, and the value of the second value field represents the answer. The content not wrapped by angle brackets <> in the question is the natural language form of the participating participles, and the content wrapped by angle brackets is the placeholder for local features. Before being fed into the LLM, they are replaced one by one with the local behavior characteristics of the corresponding detected objects in the sub-image to be analyzed. After the object behavior text and the local behavior characteristics are concatenated in order, they are fed into the LLM for behavior analysis.

[0100] The LLM extracts features from the object behavior text describing the behavior of a specific object to obtain text features. The LLM then, based on the text features and the respective global behavior characteristics of each detected object, obtains the feature correlation degree of each detected object, and determines the detected objects whose feature correlation degree exceeds the set correlation degree threshold as the target objects that have performed the specific object behavior.

[0101] As Figure 3G shown, the output result of the LLM is also the behavior name followed by the identification number of the detected object that generates the corresponding behavior. The behavior analysis alarm result of the embodiment of the present application is based on the identification number or the basic alarm unit of a single-person seat. Whether it is a single-person behavior or a multi-person interaction, this is the case. When multiple people in the scene trigger relevant behaviors, the identification numbers of all personnel will be output for warning.

[0102] LLAVA in the related art performs CLIP forward with the large-size original scene image as the basic unit, or performs CLIP forward with the small-size image of a single detected object as the basic unit, and cannot ensure both the processing consumption time and the detection effect at the same time. However, the embodiment of the present application performs behavior analysis with the sub-image to be analyzed as the unit. The relatively small image scale can not only ensure the detection effect but also retain the feasibility of theoretically performing behavior analysis on the complex problem of multi-person interaction. There is no need to worry about whether to expand the small-size image of a single detected object and how much to expand, and it also eliminates the interference of large areas where there are no detected objects on the behavior analysis effect. Moreover, the embodiment of the present application performs one CLIP forward on the sub-image to be analyzed to extract the local behavior characteristics of multiple people, which can greatly reduce the time consumption and achieve a balance between time consumption and effect.

[0103] The same detected object may exist in multiple sub-images to be analyzed. The feature fusion module realizes the reconstruction of global information based on multi-view local information through multi-view local feature fusion, compensating for the information loss caused by slicing the large-size image to be analyzed into small-size sub-images to be analyzed.

[0104] Based on the same inventive concept as the above method embodiments, an embodiment of the present application further provides a multi-modal behavior analysis device. As Figure 4 shown, the multi-modal behavior analysis device 400 may include:

[0105] An image cutting unit 401, configured to cut the image to be analyzed to obtain a plurality of sub-images to be analyzed, and a plurality of detected objects with an object distance not greater than a set distance threshold are included in one sub-image to be analyzed;

[0106] A feature fusion unit 402, configured to extract features from the plurality of sub-images to be analyzed to obtain a plurality of local feature maps;

[0107] Respectively obtain the local feature sets of each detected object from the plurality of local feature maps, and obtain the global behavior features of each detected object by pooling each local feature set;

[0108] A behavior analysis unit 403, configured to obtain a target object that performs a specific object behavior in the image to be analyzed based on the object behavior text describing the behavior of a specific object and the global behavior features of each detected object.

[0109] Optionally, the image cutting unit 401 is configured to:

[0110] Perform object detection on the image to be analyzed to obtain the object detection frames of each detected object;

[0111] Based on each object detection frame, group each detected object to obtain a plurality of object groups;

[0112] Cut the multiple detected objects in the same object group in the image to be analyzed into one sub-image to be analyzed, and obtain a plurality of sub-images to be analyzed.

[0113] Optionally, the image cutting unit 401 traverses each detected object in a loop, and performs the following operations for each traversed detected object:

[0114] When it is determined that the detected object is not grouped, expand the object detection frame of the detected object outward by a first set step length to obtain a first expanded detection frame;

[0115] Group the detected object and other detected objects that intersect with the first expanded detection frame into one object group.

[0116] Optionally, the image cutting unit 401 is configured to:

[0117] Expand the first externally expanded detection frame of the one detection object outward by a second set step length to obtain a second externally expanded detection frame, where the second externally expanded detection frame covers other detection objects in the same object group as the one detection object;

[0118] Cut the image to be analyzed along the second externally expanded detection frame to obtain a sub-image to be analyzed.

[0119] Optionally, the feature fusion unit 402 is configured to:

[0120] Obtain the local behavior features of each detection object in multiple local feature maps respectively;

[0121] Classify the local behavior features associated with the same object identifier into one local feature set to obtain the local feature set of each detection object.

[0122] Optionally, the behavior analysis unit 403 is configured to:

[0123] Extract features from the object behavior text describing the behavior of a specific object to obtain text features;

[0124] Based on the text features and the global behavior features of each detection object, obtain the feature correlation degree of each detection object, and determine the detection object whose feature correlation degree exceeds the set correlation degree threshold as the target object that has performed the specific object behavior.

[0125] For the convenience of description, the above parts are divided into each module (or unit) according to functions and described separately. Of course, when implementing the present application, the functions of each module (or unit) can be implemented in one or more software or hardware.

[0126] After introducing the multi-modal behavior analysis method and device of the exemplary embodiment of the present application, next, a computer device according to another exemplary embodiment of the present application will be introduced.

[0127] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, that is: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0128] Based on the same inventive concept as the above method embodiment, an embodiment of the present application also provides a computer device. In one embodiment, the computer device can be a server. In this embodiment, the structure of the computer device 500 is as Figure 5 shown, and may at least include a memory 501, a communication module 503, and at least one processor 502.

[0129] A memory 501 for storing a computer program executed by a processor 502. The memory 501 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.

[0130] The memory 501 can be a volatile memory, such as a random-access memory (RAM); the memory 501 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 501 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 501 can be a combination of the above memories.

[0131] The processor 502 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 502 is used to implement the above multi-modal behavior analysis method when calling the computer program stored in the memory 501.

[0132] A communication module 503 is used to communicate with a terminal device and other servers.

[0133] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 501, communication module 503 and processor 502 is not limited. In the embodiments of the present application Figure 5 it is described that the memory 501 and the processor 502 are connected through a bus 504, and the bus 504 is described in thick lines in Figure 5 For the connection methods between other components, only a schematic description is made and is not taken as a limitation. The bus 504 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of description, Figure 5 only a thick line is used to describe it in

[0134] The memory 501 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the parameter calibration method of the pan-tilt camera in the embodiments of the present application. The processor 502 is used to execute the above multi-modal behavior analysis method, as Figure 3A shown.

[0135] In another embodiment, the computer device may also be other computer device terminal devices. In this embodiment, the structure of the computer device may be as follows Figure 6 shown, including: communication component 610, memory 620, display unit 630, camera 640, sensor 650, audio circuit 660, Bluetooth module 670, processor 680 and other components.

[0136] The communication component 610 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-distance wireless transmission technology, and the electronic device can help the object send and receive information through the WiFi module.

[0137] The memory 620 can be used to store software programs and data. The processor 680 executes various functions and data processing of the terminal device by running the software programs or data stored in the memory 620. The memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. The memory 620 stores an operating system that enables the terminal device to run. In this application, the memory 620 can store the operating system and various application programs, and can also store the computer program for executing the multi-modal behavior analysis method of the embodiments of this application.

[0138] The display unit 630 can also be used to display the information input by the object or the information provided to the object, as well as the graphical user interface (GUI) of various menus of the terminal device. Specifically, the display unit 630 may include a display screen 632 disposed on the front of the terminal device. Among them, the display screen 632 can be configured in the form of a liquid crystal display, light-emitting diode, etc. The display unit 630 can be used to display the bitstream access selection interface, panoramic video display interface, etc. in the embodiments of this application.

[0139] The display unit 630 can also be used to receive input digital or character information, and generate signal inputs related to the object settings and function controls of the terminal device. Specifically, the display unit 630 may include a touch screen 631 disposed on the front of the terminal device, which can collect touch operations of the object on or near it, such as clicking buttons, dragging scroll boxes, etc.

[0140] Among them, the touch screen 631 can cover the display screen 632, or the touch screen 631 and the display screen 632 can be integrated to implement the input and output functions of the terminal device. After integration, it can be simply referred to as a touch display screen. In this application, the display unit 630 can display application programs and corresponding operation steps.

[0141] The camera 60 can be used to capture static images, and the object can publish the images captured by the camera 640 through an application. There can be one or more cameras 660. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 680 to convert it into a digital image signal.

[0142] The terminal device may further include at least one sensor 650, such as an acceleration sensor 651, a distance sensor 652, a fingerprint sensor 653, and a temperature sensor 654. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0143] The audio circuit 660, the speaker 661, and the microphone 662 can provide an audio interface between the object and the terminal device. The audio circuit 660 can transmit the electrical signal converted from the received audio data to the speaker 661, and the speaker 661 converts it into a sound signal for output. The terminal device may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 662 converts the collected sound signal into an electrical signal, which is received by the audio circuit 660, converted into audio data, and then the audio data is output to the communication component 610 to be sent to, for example, another terminal device, or the audio data is output to the memory 620 for further processing.

[0144] The Bluetooth module 670 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 670 to perform data interaction.

[0145] The processor 680 is the control center of the terminal device, connecting various parts of the entire terminal through various interfaces and lines. By running or executing software programs stored in the memory 620 and invoking data stored in the memory 620, it performs various functions of the terminal device and processes data. In some embodiments, the processor 680 may include one or more processing units; the processor 680 may also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 680. In this application, the processor 680 can run the operating system, application programs, user interface display, and touch response, as well as the multi-modal behavior analysis method of the embodiments of this application. In addition, the processor 680 is coupled to the display unit 630.

[0146] In some possible implementation manners, each aspect of the multi-modal behavior analysis method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on a computer device, the computer program is used to cause the computer device to execute the steps in the multi-modal behavior analysis method according to various exemplary embodiments of this application described above in this specification. For example, the computer device can execute the steps as shown in Figure 3A shown.

[0147] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0148] The program product of the embodiments of this application can adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can run on an electronic device. However, the program product of this application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with a command execution system, device, or component.

[0149] A readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.

[0150] The computer program contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0151] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's computer device, partially on the user's computer device, executed as an independent software package, partially on the user's computer device and partially on a remote computer device, or entirely on the remote computer device. In the case of a remote computer device, the remote computer device can be connected to the user's computer device through any type of network including a local area network (LAN) or a wide area network (WAN), or, can be connected to an external computer device (e.g., connected through the Internet using an Internet service provider).

[0152] It should be noted that although several units or subunits of the apparatus are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0153] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0154] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.

[0155] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0158] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0159] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A multimodal behavior analysis method, characterized in that: include: Cutting the image to be analyzed to obtain multiple sub-images to be analyzed, wherein one sub-image to be analyzed includes multiple detection objects whose object distance is not greater than a set distance threshold; Extracting features from the multiple sub-graphs to be analyzed respectively to obtain multiple local feature graphs, wherein each local feature graph includes: local behavior features of the multiple detection objects in the corresponding sub-graphs to be analyzed; The local behavior features of each detection object in the multiple local feature maps are respectively obtained, and the local behavior features associated with the same object identifier are divided into a local feature set to obtain the local feature set of each detection object, and the global behavior features of each detection object are obtained by pooling the local feature sets; Based on the object behavior text describing the specific object behavior and the global behavior characteristics of each detected object, the target object performing the specific object behavior in the image to be analyzed is obtained.

2. The method according to claim 1, characterized in that The cutting of the image to be analyzed to obtain a plurality of sub-images to be analyzed includes: Performing object detection on the image to be analyzed to obtain an object detection frame for each detected object; Grouping the detected objects into a plurality of object groups based on the object detection frames; A plurality of detection objects in the same object group in the image to be analyzed are cut into a sub-image to be analyzed, thereby obtaining a plurality of sub-images to be analyzed.

3. The method according to claim 2, characterized in that The detection objects are grouped based on the object detection frames, and the detection objects are divided into multiple object groups. The detection objects are traversed in a loop, and the following operations are performed each time a detection object is traversed: When it is determined that the one detection object is not grouped, expanding the object detection frame of the one detection object outward by a first set step length to obtain a first outwardly expanded detection frame; The one detection object and other detection objects intersecting with the first outward expanded detection frame are divided into an object group, and the one detection object and other detection objects intersecting with the first outward expanded detection frame are added to the grouped set.

4. The method according to claim 2, characterized in that The step of cutting a plurality of detection objects in the same object group in the image to be analyzed into a sub-image to be analyzed includes: Expanding the first outward-expanded detection frame of the one detection object outwardly by a second set step length to obtain a second outward-expanded detection frame, wherein the second outward-expanded detection frame covers other detection objects in the same object group as the one detection object; The image to be analyzed is cut along the second outward-expanded detection frame to obtain a sub-image to be analyzed.

5. The method according to claim 1, characterized in that The step of obtaining the target object that performs the specific object behavior in the image to be analyzed based on the object behavior text that describes the specific object behavior and the global behavior features of each detected object includes: Extracting features from the object behavior text describing the behavior of a specific object to obtain text features; Based on the text features and the global behavior features of each detection object, the feature correlation of each detection object is obtained, and the detection object whose feature correlation exceeds the set correlation threshold is determined as the target object that has performed the specific object behavior.

6. A multimodal behavior analysis device, characterized in that: include: An image cutting unit is used to cut the image to be analyzed to obtain multiple sub-images to be analyzed, wherein one sub-image to be analyzed includes multiple detection objects whose object distance is not greater than a set distance threshold; A feature fusion unit is used to extract features from the multiple sub-graphs to be analyzed respectively to obtain multiple local feature graphs, where each local feature graph includes: local behavior features of the multiple detection objects in the corresponding sub-graphs to be analyzed; Respectively obtain local behavior features of each detection object in multiple local feature maps, divide the local behavior features associated with the same object identifier into a local feature set, obtain a local feature set of each detection object, and obtain a global behavior feature of each detection object by pooling the local feature sets; The behavior analysis unit is used to obtain the target object performing the specific object behavior in the image to be analyzed based on the object behavior text describing the specific object behavior and the global behavior characteristics of each detected object.

7. A computer device, characterized in that: It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: It includes program code, and when the program code is run on a computer device, the program code is used to enable the computer device to execute the steps of the method described in any one of claims 1 to 5.

9. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Behavior recognition method and device and computer readable storage medium

    CN110738101A

  • KR1018681030000B1