An event reasoning method and system based on multi-modal features
By integrating feature and identifier image data from multiple image acquisition devices, combined with multimodal large models and prompt words, the problem of insufficient cross-view information fusion in multimodal monitoring is solved, achieving more accurate event reasoning and object recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU EZVIZ SOFTWARE CO LTD
- Filing Date
- 2025-12-17
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies lack effective cross-perspective information fusion mechanisms in multimodal monitoring, making it difficult for large multimodal models to reason about events from a global perspective. They also fail to fully utilize the correlation information between image data sources from different channels, resulting in low accuracy of event reasoning results. Furthermore, traditional REID technology cannot effectively collaborate with large multimodal models, making it difficult to accurately identify objects and determine behavioral trajectories.
By acquiring image data from multiple image acquisition devices at different poses, features are extracted and fused to identify objects in the images. The fused features and identified images are then input into a multimodal large model, which is combined with prompt words to guide the model in event reasoning and to achieve comprehensive understanding using semantic and visual features.
It enables accurate reasoning about events from a global perspective, improves the accuracy of event reasoning results, and allows for a more comprehensive and accurate understanding of image information, object identification, and judgment of behavioral trajectories.
Smart Images

Figure CN121366385B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent security technology, and in particular to an event reasoning method and system based on multimodal features. Background Technology
[0002] In the field of intelligent security technology, multimodal monitoring technology is playing an increasingly important role. For example, in scenarios such as home entry, multi-channel, multi-angle image acquisition devices are typically deployed to collect multiple image data sources in order to obtain more comprehensive and accurate information. However, in the process of event reasoning based on the acquired multi-channel image data using existing multimodal large-scale models, each image data source is usually analyzed independently or simply stitched together, lacking an effective cross-view information fusion mechanism. This approach makes it difficult for multimodal large-scale models to reason about events from a global perspective, failing to fully utilize the correlation information between different channel image data sources. This results in a superficial understanding of the scene, making it difficult to uncover the deeper logical relationships behind the events. Furthermore, current technologies use REID (Person Re-Identification) technology in the reasoning process, but traditional REID technology can only output a numerical ID (Identity number), which cannot effectively collaborate with multimodal large-scale models. This results in the system being unable to distinguish "who is who," leading to low accuracy in event reasoning results obtained from multi-channel image data sources. For example, in a home entry scenario, the images captured by cameras at different angles may have overlapping areas. However, existing technologies cannot effectively integrate the information from these overlapping areas, and when there are multiple objects in the image, it is impossible to accurately identify the objects, making it difficult to accurately determine the complete behavioral trajectory and event development of a specific object in the scene. Summary of the Invention
[0003] The purpose of this application is to provide an event reasoning method and system based on multimodal features, so as to improve the accuracy of event reasoning. The specific technical solution is as follows:
[0004] In a first aspect, embodiments of this application provide an event reasoning method based on multimodal features, the method comprising:
[0005] Multiple image data points are acquired from multiple image acquisition devices during a target time period; wherein the multiple image acquisition devices are used to capture images of the target area from different poses.
[0006] Image features are extracted from each of the image data separately, and the extracted image features are fused to obtain fused features;
[0007] Each object present in the image data is identified to obtain an identified image; wherein, different objects in the identified image are identified in different ways;
[0008] The fused features, each of the labeled images, and the prompt words are input into a multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target region within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
[0009] In one possible implementation, the fused image features obtained from the fusion extraction are combined to form fused features, including:
[0010] Each of the aforementioned image features is input into an event detector to obtain events occurring in the target region within the target time period as the second event; wherein, the algorithm complexity of the event detector is lower than that of the multimodal large model;
[0011] According to the first dynamic weight of each image acquisition device, the image features are weighted and fused to obtain fused features; wherein, the first dynamic weight of each image acquisition device is positively correlated with the importance of the region captured by the image acquisition device in the second event.
[0012] In one possible implementation, the method further includes:
[0013] If the first event and the second event meet a preset difference condition, then the event detector and / or the first dynamic weight are adjusted in a direction that makes the second event the same as the first event.
[0014] In one possible implementation, the method further includes:
[0015] Based on the type of objects present in the target area, obtain a second dynamic weight that has been pre-set for the type;
[0016] The fused features, each of the labeled images, and the prompt words are input into a multimodal large model, so that the multimodal large model, guided by the prompt words, considers and outputs events occurring in the target region within the target time period as the first event, including:
[0017] The fused features, each of the labeled images, and the prompt words are input into the multimodal large model, so that the multimodal large model, guided by the prompt words, uses the second dynamic weights to identify objects existing in the target area, and, based on the identification results, considers and outputs the events that occur in the target area within the target time period as the first event.
[0018] Secondly, embodiments of this application provide an event reasoning system based on multimodal features, the system including multiple image acquisition devices and a processor; each of the image acquisition devices has a different pose;
[0019] Each of the aforementioned image acquisition devices is used to capture image data of the target area from different poses and upload the captured image data to the processor;
[0020] The processor is configured to execute any of the methods described in the first aspect.
[0021] Thirdly, embodiments of this application provide an event reasoning device based on multimodal features, the device comprising:
[0022] The acquisition module is used to acquire multiple image data captured by multiple image acquisition devices during a target time period; wherein, the multiple image acquisition devices are used to capture the target area from different poses;
[0023] The fusion module is used to extract image features from each of the image data separately and fuse the extracted image features to obtain fused features;
[0024] The identification module is used to identify the objects present in each of the image data to obtain an identified image; wherein, different objects in the identified image are identified in different ways;
[0025] The output module is used to input the fused features, each of the labeled images, and the prompt words into the multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target area within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
[0026] Fourthly, embodiments of this application provide an electronic device, including:
[0027] Memory, used to store computer programs;
[0028] When a processor executes a program stored in memory, it implements any of the event reasoning methods described above.
[0029] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the event reasoning methods described above.
[0030] Sixthly, embodiments of this application also provide a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the event reasoning methods described above.
[0031] Beneficial effects of the embodiments in this application:
[0032] This application provides an event reasoning method and system based on multimodal features. By fusing features extracted from image data captured from different poses by multiple image acquisition devices, a fused feature is obtained. Compared to existing technologies that independently analyze or simply stitch together image data sources, event reasoning based on fused features can utilize the correlation information between image data sources from different angles to achieve event reasoning from a global perspective. By identifying objects in each image data and inputting the identified images into a multimodal large-scale model, and using prompts to guide the model's thinking, the prompts characterize the identification methods of each object in each identified image. Different objects have different identification methods, providing the multimodal large-scale model with clear object differentiation information, enabling it to more accurately identify objects. This method, through the introduction of fused features and identified images (which can be considered semantic features and identified images as visual features), combined with prompts to guide the multimodal large-scale model's thinking, integrates semantic and visual features for a more comprehensive and accurate understanding of image information. This results in a more accurate output of the event occurring in the target area within the target time period, i.e., the first event, thus improving the accuracy of the event reasoning results.
[0033] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0035] Figure 1 This is a schematic diagram of a first type of event reasoning method provided in an embodiment of this application;
[0036] Figure 2 This is a second flowchart illustrating the event reasoning method provided in the embodiments of this application;
[0037] Figure 3 This is a schematic diagram of a third type of event reasoning method provided in the embodiments of this application;
[0038] Figure 4This is a schematic diagram of the fourth process of the event reasoning method provided in the embodiments of this application;
[0039] Figure 5 A fifth flowchart illustrating the event reasoning method provided in this application embodiment;
[0040] Figure 6 A sixth flowchart illustrating the event reasoning method provided in this application embodiment;
[0041] Figure 7 This is a schematic diagram of the structure of the event reasoning system provided in the embodiments of this application;
[0042] Figure 8 This is a schematic diagram of the structure of the event reasoning device provided in the embodiments of this application;
[0043] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0045] In the field of intelligent security technology, multimodal monitoring technology is playing an increasingly important role. For example, in scenarios such as home entry, multi-channel, multi-angle image acquisition devices are typically deployed to collect multiple image data sources in order to obtain more comprehensive and accurate information. However, existing technologies using multimodal large models to perform event reasoning based on the acquired multi-channel image data have the following problems:
[0046] Problem 1: Existing technologies often analyze each image data source independently or simply stitch them together when processing multi-channel image data, lacking an effective cross-view information fusion mechanism. This approach makes it difficult for multimodal large models to reason about events from a global perspective, failing to fully utilize the correlation information between different channel image data sources. This results in a superficial understanding of the scene, making it difficult to uncover the deeper logical relationships behind the events. Furthermore, while current technologies use REID technology in the reasoning process, traditional REID technology only outputs numerical IDs and cannot effectively collaborate with multimodal large models. This makes it difficult for the system to distinguish "who is who," leading to low accuracy in event reasoning results obtained from multi-channel image data sources. For example, in a home entrance scene, cameras at different angles capture images from different perspectives, but existing technologies cannot effectively integrate this information from different perspectives. Moreover, when multiple objects are present in the image, objects cannot be accurately identified, making it difficult to accurately determine the complete behavioral trajectory and event development of a specific object in the scene.
[0047] Problem 2: A single multimodal large model is difficult to handle tasks with huge differences in different scenarios (such as human, vehicle and pet behavior recognition) simultaneously and well. It usually requires full parameter fine-tuning of the model, which is costly and difficult to scale flexibly.
[0048] To address the aforementioned issues, embodiments of this application provide an event reasoning method based on multimodal features, see [link to relevant documentation]. Figure 1 , Figure 1 This is a first flowchart illustrating an event reasoning method provided in an embodiment of this application. The method includes:
[0049] S101: Acquire multiple image data captured by multiple image acquisition devices during the target time period.
[0050] Among them, multiple image acquisition devices are used to capture images of the target area from different poses.
[0051] S102, extract the image features of each image data separately, and fuse the extracted image features to obtain the fused features.
[0052] S103, identify the objects present in each image data to obtain the identified image.
[0053] In this process, different objects in the image are identified in different ways.
[0054] S104, input the fused features, each labeled image and the prompt words into the multimodal large model, so that the multimodal large model can think about and output the events that occur in the target area within the target time period under the guidance of the prompt words, as the first event.
[0055] The prompt words are used to characterize the way each object is identified in each image.
[0056] By applying the above embodiments, fused features are obtained by fusing features extracted from image data captured from different poses by multiple image acquisition devices. Compared with existing technologies that analyze each image data source independently or simply stitch them together, event reasoning based on fused features can utilize the correlation information between image data sources from different angles to achieve event reasoning from a global perspective. By identifying objects in each image data and inputting the identified images into a multimodal large-scale model, combined with prompts to guide the model's thinking, and using prompts to characterize the identification methods of each object in each identified image (different objects have different identification methods), the multimodal large-scale model receives clear object differentiation information, enabling it to more accurately identify objects. This method, through the introduction of fused features and identified images (which can be considered semantic features and identified images as visual features), combined with prompts to guide the multimodal large-scale model's thinking, integrates semantic and visual features, allowing for a more comprehensive and accurate understanding of image information. This results in a more accurate output of the event occurring in the target area within the target time period, i.e., the first event, improving the accuracy of event reasoning results and thus solving the aforementioned problem one.
[0057] The following will explain steps S101-S104:
[0058] In step S101, the image acquisition device can be a PTZ camera, network camera, infrared camera, or other device with image acquisition capabilities. The image acquisition device's pose when capturing the target area varies, meaning at least one of its position or angle differs. For example, in one example, two image acquisition devices are included, denoted as image acquisition device 1 and image acquisition device 2. Image acquisition device 1 is installed at a high position for overhead shooting, while image acquisition device 2 is installed at a low position for upward shooting; or image acquisition device 1 is installed on the right side of the target area for shooting, and image acquisition device 2 is installed on the left side of the target area for shooting. By capturing images from different poses, information about the target area can be obtained from multiple angles, providing rich image data for event inference. The target time period is any time period during which events of interest to the user may occur in the target area, and these events should be events that can be inferred through a multimodal large model.
[0059] In step S102, feature extraction is performed on the image data acquired in step S101. To comprehensively extract key information from the image, specifically, taking image acquisition devices including image acquisition device 1, image acquisition device 2, ..., image acquisition device n as an example, a shared weight backbone network can be used. The image data acquired by image acquisition devices 1, 2, ..., n are first passed through a shared weight deep convolutional neural network, which extracts visual feature maps {F1, F2, ..., Fn}. Each visual feature map is used to characterize the image features of each channel image (i.e., image data acquired by different image acquisition devices). Specifically, F1 is the visual feature map obtained by feature extraction from the image data acquired by image acquisition device 1, F2 is the visual feature map obtained by feature extraction from the image data acquired by image acquisition device 2, and so on, with Fn being the visual feature map obtained by feature extraction from the image data acquired by image acquisition device n. Deep convolutional neural networks include, but are not limited to, network architectures with feature extraction capabilities such as ResNet and Vision Transformer; specific limitations are not specified here.
[0060] Image features can include one or more of the following: color features (such as the average color and color distribution of different regions in an image), texture features (describing the texture pattern of the surface of an object in an image, such as roughness or smoothness), shape features (the outline shape and geometric shape of an object), and spatial relationship features (the position of an object in an image and its spatial arrangement with each other).
[0061] Since images captured by different image acquisition devices reflect information about the target area from different angles, fusing the features extracted from each image allows for the comprehensive utilization of information from multiple perspectives, resulting in representative and comprehensive features. This fused feature can better describe the overall situation of the target area within the target time period.
[0062] Fusion methods can take many forms, such as weighted fusion (which will be illustrated below and will not be repeated here), embedding a learnable channel semantic layer, and feature fusion methods based on machine learning algorithms.
[0063] In step S103, image recognition technology is used to detect and identify various objects present in the image, such as people, animals, vehicles, and buildings. Image recognition technology can include deep learning-based object detection algorithms, biometric recognition technology, template matching-based recognition technology, etc. Deep learning-based object detection algorithms can include the YOLO (You Only Look Once) algorithm.
[0064] After identifying objects in an image, the image is labeled, visually marking the identified objects to generate a labeled image. Different objects can be distinguished using any combination of elements such as different colors, shapes, text, and numbers. For example, in traffic monitoring images, vehicles driving normally are marked with green rectangles; vehicles running red lights are highlighted with red rectangles; and vehicles that have broken down and are stopped on the roadside are marked with yellow rectangles.
[0065] In addition, the type of each object and the attributes of each image acquisition device can be labeled.
[0066] The type to which each object belongs is the class to which each object belongs. When labeling the type to which each object belongs, the labeling methods include, but are not limited to, text labels (such as adding the type name directly next to the label box, such as "pedestrian" or "non-motorized vehicle"), numerical codes (assigning a unique numerical ID to different types to facilitate quick retrieval and classification by the system), icon symbols, and color coding systems (establishing a mapping relationship between color and type, such as blue representing "passable area"), so that multimodal large models can better understand the labeled images.
[0067] Each image acquisition device possesses an embedding vector that is learned jointly with the network during training. This embedding vector characterizes the attributes of each image acquisition device and is initialized during system initialization through short-term self-supervised learning of the scene. The attributes of each image acquisition device refer to its geometric and physical characteristics and semantic functional characteristics.
[0068] The geometrical and physical characteristics of each image acquisition device refer to the spatial observation characteristics determined by the hardware configuration and installation method of the image acquisition device. Examples include: field of view, such as the wide-angle (>90°) of an upper-view camera covering the entire area, and the narrow-angle (<45°) of an lower-view camera focusing on a local area; and installation angle, such as looking down (e.g., ceiling camera), looking straight ahead (e.g., wall camera), or looking up (e.g., ground-mounted camera). The semantic attributes of each image acquisition device refer to its functional positioning and expected observation content within the scene. Examples include: monitoring area type, such as the "global view" of an upper-view camera corresponding to open spaces (e.g., courtyards, parking lots), and the "focus on the ground" of an lower-view camera corresponding to densely populated areas such as entrances and corridors. The annotation method is similar to the annotation method for the types of objects mentioned above, and will not be repeated here.
[0069] In one possible embodiment, the objects present in the image, the types of the objects, and the attributes of each image acquisition device together constitute a semantic embedding vector. Specifically, taking the image acquisition devices including image acquisition device 1, image acquisition device 2, ..., image acquisition device n as an example, the semantic embedding vector is {E1, E2, ..., En}.
[0070] In step S104, the multimodal large model is a type of model that combines the natural language processing capabilities of a large language model with the ability to understand and generate data from other modalities (such as visual and semantic data).
[0071] The fused features obtained in the preceding steps, the various labeled images, and the prompt words are input into the multimodal large model. The fused features integrate feature representations from multi-angle images, the labeled images contain object recognition and annotation information, and the prompt words guide the model in understanding the object identification methods within the labeled images. Using the labeled images as input to the multimodal large model is equivalent to injecting channel semantic features into the input data.
[0072] The purpose of cue words is to provide the multimodal large model with information about how objects are identified in the labeled images, helping the model better understand the meaning of different objects in the image. After receiving this input data, the multimodal large model uses its internal knowledge and algorithms to reason and analyze, combining fused features and information from the labeled images to infer the events that occurred in the target area during the target time period, and outputs this inference as the first event. For example, if the labeled image shows vehicles A, B, and C all located in specific positions, and the fused features include features reflecting traffic congestion, the multimodal large model will output "A traffic congestion event occurred in the target area during the target time period" as the first event.
[0073] In step S102 above, when fusing the extracted image features to obtain the fused features, as explained above, a weighted average can be used to fuse the image features. For details, see [link to relevant documentation]. Figure 2 , Figure 2 This is a second flowchart illustrating the event reasoning method provided in an embodiment of this application, the method comprising:
[0074] S101: Acquire multiple image data captured by multiple image acquisition devices during the target time period.
[0075] Among them, multiple image acquisition devices are used to capture images of the target area from different poses.
[0076] S1021, extract the image features of each image data respectively, and input each image feature into the event detector to obtain the events that occur in the target area within the target time period output by the event detector, as the second event.
[0077] Among them, the algorithm complexity of the event detector is lower than that of the multimodal large model.
[0078] S1022, according to the first dynamic weight of each image acquisition device, the image features are weighted and fused to obtain the fused features.
[0079] Among them, the first dynamic weight of each image acquisition device is positively correlated with the importance of the area captured by the image acquisition device in the second event.
[0080] S103, identify the objects present in each image data to obtain the identified image.
[0081] In this process, different objects in the image are identified in different ways.
[0082] S104, input the fused features, each labeled image and the prompt words into the multimodal large model, so that the multimodal large model can think about and output the events that occur in the target area within the target time period under the guidance of the prompt words, as the first event.
[0083] The prompt words are used to characterize the way each object is identified in each image.
[0084] Steps S1021 and S1022 are detailed steps of the aforementioned step S102. Steps S101, S103 and S104 have been explained in the preceding text and will not be repeated here.
[0085] In step S1021, the event detector is an algorithm or model used to identify specific events in image data based on input image features. The extracted image features are input into the event detector, which analyzes the features using a preset algorithm or lightweight model and filters events based on a preset confidence level to determine whether a specific event, such as traffic violations, crowd gatherings, or abnormal behavior, has occurred in the target area within a target time period. The detected events are then output as second events for subsequent event inference.
[0086] Specifically, taking an event detector as an example of a target detection or action recognition model, the target detection or action recognition model analyzes the features of each input image to detect predefined atomic events and compound events.
[0087] Atomic events consist of single events, such as "person appears," "package appears," and "movement begins." Composite events are composed of multiple atomic events. For example, the composite event "person returns home" includes the atomic events "person appears from the external camera," "door opens," and "person appears from the internal camera." The external camera and the internal camera refer to the image acquisition devices in different poses mentioned above.
[0088] After a valid event (atomic event or composite event) is detected, a structured event description vector V_event is automatically generated to represent the second event. V_event = [event type code, confidence level, trigger channel (i.e., the image acquisition device that captured the corresponding event), timestamp, spatial location]. For example, the structured event description vector for the composite event "People go home" is [HOME_ARRIVAL_01, 0.92, [Cam_Outdoor, Door_Sensor, Cam_Indoor], 2025-09-25_14:30:45, [LivingRoom_Zone]].
[0089] The algorithm complexity of the event detector is lower than that of the aforementioned multimodal large model, which means that the event detector is more efficient at outputting the second event than the multimodal large model is at outputting the first event.
[0090] In step S1022, a first dynamic weight can be assigned to each image acquisition device by introducing algorithms or models such as gating networks (i.e., channel gating) and Transformer models. The weight value is dynamically adjusted according to the importance of the area captured by the image acquisition device in the second event. The gating network is usually a simple MLP (Multilayer Perceptron).
[0091] The weights are proportional to the importance of the region. The image features of each device are weighted and summed according to their importance to generate a fused feature. For example, if the second event is a car accident, if image acquisition device A captures the core area of the event (such as the accident scene), its weight will be higher; while image acquisition device B, which captures the edge area, will have a lower weight. As another example, if the second event is an abnormal entry into a warehouse, the weight of image acquisition devices near the entrance will be significantly higher than that of image acquisition devices in other areas. The image features of image acquisition devices A and B are weighted and summed according to their importance to generate a fused feature.
[0092] For example, the global importance weights {α1, α2, ..., αn} of the regions captured by each image acquisition device are calculated. For instance, when the second event is the aforementioned composite event "people return home," the value of α for the upper-view image is the highest. Channel gating generates a spatial weight map for the feature map of each channel, identifying the importance of different regions in the feature map. For example, when the second event is the aforementioned atomic event "package appears," the region below the feature map of the lower-view image has the highest weight. Finally, the weighted multi-channel features (αi) are... Fi) is further combined with the spatial weight map to generate fused features through a fusion network (such as a convolutional neural network after splicing).
[0093] It is understandable that the positive correlation between the first dynamic weight of each image acquisition device and the importance of the area captured by the image acquisition device in the second event means that, assuming that other factors affecting the first dynamic weight remain unchanged, the first dynamic weight increases monotonically with the increase of importance. The monotonous increase in this article can refer to a strict monotonous increase or a non-strict monotonous increase.
[0094] Applying the above embodiments, in real-world scenarios, the importance of different regions in different events dynamically changes. Image features are weighted and fused according to the first dynamic weight of each image acquisition device, and the weight is positively correlated with the importance of the region captured by the image acquisition device in the second event. This method, by employing dynamic weights related to event importance, allows the fused features to adaptively adjust the fusion ratio of each image feature according to the actual event, thus highlighting image information related to key events, enhancing the ability of the fused features to capture important event features, and improving the effectiveness and relevance of the fused features. This results in multimodal large models outputting more accurate event inference results based on the fused features, thereby improving the accuracy of event inference.
[0095] To further improve the accuracy of event reasoning, a feedback correction mechanism can be introduced. For details, see [link to relevant documentation]. Figure 3 , Figure 3 This is a third flowchart illustrating the event reasoning method provided in the embodiments of this application, the method comprising:
[0096] S101: Acquire multiple image data captured by multiple image acquisition devices during the target time period.
[0097] Among them, multiple image acquisition devices are used to capture images of the target area from different poses.
[0098] S1021, extract the image features of each image data respectively, and input each image feature into the event detector to obtain the events that occur in the target area within the target time period output by the event detector, as the second event.
[0099] Among them, the algorithm complexity of the event detector is lower than that of the multimodal large model.
[0100] S1022, according to the first dynamic weight of each image acquisition device, the image features are weighted and fused to obtain the fused features.
[0101] Among them, the first dynamic weight of each image acquisition device is positively correlated with the importance of the area captured by the image acquisition device in the second event.
[0102] S103, identify the objects present in each image data to obtain the identified image.
[0103] In this process, different objects in the image are identified in different ways.
[0104] S104, input the fused features, each labeled image and the prompt words into the multimodal large model, so that the multimodal large model can think about and output the events that occur in the target area within the target time period under the guidance of the prompt words, as the first event.
[0105] The prompt words are used to characterize the way each object is identified in each image.
[0106] S105, if the first event and the second event meet the preset difference condition, then the event detector and / or the first dynamic weight are adjusted in the direction that makes the second event the same as the first event.
[0107] Steps S101-S104 have been explained in the preceding text and will not be repeated here.
[0108] In step S105, the judgment result of whether the first event and the second event meet the preset difference condition can be obtained by the system directly judging based on the first event and the second event, or it can be obtained by the user confirming the system's judgment result through application software, such as "false alarm" or "missed alarm".
[0109] The preset difference condition is a pre-set standard used to determine whether the first event and the second event are inconsistent. The difference is reflected in one or more dimensions such as event type, occurrence location, and occurrence time. The difference is caused by one or more factors such as anomalies detected by the event detector and errors in the first dynamic weight allocation.
[0110] Based on this, the preset difference conditions include, but are not limited to, abnormal event type detection, abnormal event location detection, and abnormal event time detection. For example, if the first event shows a package appearing in a specific area, but the second event fails to detect the package, or misidentifies other objects (such as shadows) as packages, this indicates a difference between the two events, satisfying the preset difference conditions. When the first and second events are detected to meet the preset difference conditions, the event detector and the first dynamic weight need to be adjusted. During adjustment, the event detector can be adjusted alone, the first dynamic weight can be adjusted alone, or both can be adjusted simultaneously. The specific adjustment method depends on the main reason for the difference between the first and second events in the actual system. If the main issue is detection error by the event detector, then the event detector is adjusted, such as adjusting the confidence threshold of the event detector; if the main issue is unreasonable weight allocation, then the first dynamic weight is adjusted, such as injecting a correction vector containing semantic interpretation into the gating network to achieve automatic correction of the feature space of the gating network and optimize the weight allocation logic; if both have problems, then both are adjusted simultaneously.
[0111] By injecting correction vectors containing semantic interpretations into the gating network, the feature space of the gating network can be automatically corrected, and the weight allocation logic can be optimized. This can be achieved by collecting current hard examples for online incremental learning, adjusting the weight distribution of each branch of the gating network, and correcting the weight allocation logic. Furthermore, the gating network and embedding vectors can be fine-tuned periodically (e.g., when the gating network is idle) using accumulated running data, thereby optimizing its weight allocation strategy.
[0112] To more clearly illustrate the aforementioned feedback correction mechanism, the following explanation will take the example of image acquisition devices including image acquisition device 1, image acquisition device 2, and image acquisition device 3, where each image acquisition device is assigned a first dynamic weight through a gating network. (See [link to relevant documentation]). Figure 4 , Figure 4 This is a schematic diagram of the fourth process of the event reasoning method provided in the embodiments of this application. Since the image data acquired by the image acquisition device needs to be transmitted to the processor through its corresponding channel to realize event reasoning, therefore... Figure 4 Channel 1, Channel 2, and Channel 3 are used to characterize image acquisition device 1, image acquisition device 2, and image acquisition device 3, respectively.
[0113] The visual feature extraction module performs feature extraction based on the image data transmitted through channels 1, 2, and 3, obtaining feature 1, feature 2, and feature 3. This step is equivalent to the aforementioned steps S101 and S102.
[0114] Input features 1, 2 and 3 into the event detector to obtain the event vector output by the event detector (i.e. the aforementioned second event), which is equivalent to the aforementioned step S1021.
[0115] After obtaining the event vector, channel semantic features (i.e., ...) are embedded in feature 1, feature 2, and feature 3. Figure 4 The semantic feature embedding of channel 1, channel 2 and channel 3 in the above-mentioned steps S103 are equivalent to the aforementioned steps S103.
[0116] Gated networks (i.e.) Figure 4 The gated network layer takes as input the three channels of globally average pooled features (feature 1, feature 2, and feature 3) and the event vector features input from the event detector, concatenated through channels, and outputs a ternary weight vector (i.e., Figure 4 The dynamic channel 1 weight coefficient, dynamic channel 2 weight coefficient, and dynamic channel 3 weight coefficient (also known as the aforementioned first dynamic weight) are used to obtain the fused feature based on the ternary weight vector. This is equivalent to the aforementioned step S1022.
[0117] Finally, the obtained fused features are input into a multimodal large model, which then interprets and outputs the event reasoning results (i.e., Figure 4 The multimodal understanding output (i.e., the first event mentioned above) is used to diagnose false detection events in real time and generate feedback signal 1 and feedback signal 2. Feedback signal 1 is used to dynamically adjust the confidence threshold of the event detection module (e.g., the package recognition threshold is increased from 0.7 to 0.85), and feedback signal 2 is used to inject a correction vector into the gating network to correct the weight allocation logic.
[0118] A confidence threshold is a quantitative metric used to measure the reliability of event detection results. Multimodal large models analyze and judge based on the fused features of the input. When the calculated confidence score exceeds a pre-set threshold, the corresponding event is considered detected. For example, in a package recognition scenario, if a package is determined to exist in a certain area, and the calculated confidence score is 0.8, while the preset package recognition confidence threshold is 0.7, then it will be determined that a package does indeed exist in that area.
[0119] Applying the above embodiments, when there is a difference between the first event (which can be considered the actual situation or a more reliable reference standard) and the second event (the output result of the event detector), it indicates that the event detector may have made a misjudgment or missed judgment, or that the first dynamic weight allocation is incorrect. By adjusting the event detector and / or the first dynamic weight in the same direction as the first event, the judgment standard of the event detector or the weight allocation of feature fusion can be corrected, thereby reducing the occurrence of misjudgments and missed judgments, and thus improving the accuracy of event reasoning.
[0120] As described in step S103 above, the image can also be identified with the type of the object. Based on this, the type of the object in the target region can be obtained during the event reasoning process. After obtaining the object type, a second dynamic weight pre-set for that type is found.
[0121] The second dynamic weight is the LoRa (Low-Rank Adaptation) weight, which dynamically selects and loads the corresponding contextualized LoRa weight based on the object type. For example, if the current scene object type is a person, then the LoRa weight module for "person" is loaded; if the current scene object type is both a person and a car, then the LoRa weights for both "person" and "car" are loaded.
[0122] Specifically, in different scenarios (i.e., different target areas), different types of objects have varying importance in event recognition. To more accurately achieve event reasoning in different scenarios, corresponding weights are pre-set for each object type. For example, in traffic monitoring scenarios, the weight for vehicles may be set higher; while in pedestrian safety monitoring scenarios, the weight for pedestrians will be relatively higher. This dynamic weight setting method allows the multimodal large model to adapt more flexibly to different task requirements. For details on how to determine the type, please refer to the relevant explanation in step S103 above.
[0123] Based on this, see Figure 5 , Figure 5 A fifth flowchart illustrating the event reasoning method provided in this application embodiment, the method comprising:
[0124] S101: Acquire multiple image data captured by multiple image acquisition devices during the target time period.
[0125] Among them, multiple image acquisition devices are used to capture images of the target area from different poses.
[0126] S102, extract the image features of each image data separately, and fuse the extracted image features to obtain the fused features.
[0127] S103, identify the objects present in each image data to obtain the identified image.
[0128] In this process, different objects in the image are identified in different ways.
[0129] S1041, input the fused features, each labeled image and the prompt words into the multimodal large model, so that the multimodal large model can identify the objects existing in the target area under the guidance of the prompt words and use the second dynamic weights, and think about and output the events that occurred in the target area within the target time period as the first event based on the identification results.
[0130] The prompt words are used to characterize the way each object is identified in each image.
[0131] Step S1041 is a detailed step of the aforementioned step S104. Steps S101-S103 have been explained in the preceding text and will not be repeated here.
[0132] In step S1041, after receiving the input fused features, labeled image, and prompt words, the multimodal large model uses the acquired second dynamic weights to identify objects within the target area. For example, if pedestrian A is marked with a red box, pedestrian N with a blue box, vehicle 1 with the text "vehicle 1", and vehicle 2 with the text "vehicle 2" in the prompt word labeled image, and the second dynamic weights pre-set for vehicles and pedestrians are acquired, with a higher weight for vehicles in the second dynamic weights, then the model will pay more attention to the dynamics and location information of vehicles to determine whether there are behaviors such as illegal parking or speeding.
[0133] After identifying objects within the target area, the model performs further analysis and reasoning based on the identification results. Combining information from fused features and prompts, it generates an accurate description of events occurring in the target area within the target time period. For example, the multimodal large model outputs "At 3:15 PM, vehicle 2 illegally parked in parking space B of parking area A, occupying two parking spaces."
[0134] Applying the above embodiments, by obtaining a pre-set second dynamic weight based on the object type in the target region, the model uses this weight to identify objects and output events under the guidance of prompt words. This method improves the multi-modal large model's adaptability to multiple scenarios. When facing event reasoning in different scenarios, only the second dynamic weight needs to be adjusted according to the object type, avoiding full parameter fine-tuning of the multimodal large model, reducing the cost of event reasoning, quickly adapting to new scenarios, improving the flexibility and accuracy of event reasoning, and solving the aforementioned problem two.
[0135] To more clearly illustrate the event reasoning method provided in the embodiments of this application, the following explanation will be based on a flowchart. (See attached flowchart.) Figure 6 , Figure 6 A sixth flowchart illustrating the event reasoning method provided in this application embodiment includes:
[0136] The input video image data is acquired. This video image data can be single-channel uploaded video image data or multi-channel uploaded video image data. This is equivalent to the aforementioned step S101.
[0137] Target detection and REID recognition are performed based on the input video image data, and the detection and recognition results are input into the identity semantic injection module to obtain the data output by the identity semantic injection module (that is, the identification image and prompt words in the aforementioned steps S103 and S104).
[0138] The dynamic feature fusion module extracts the image features of each video image data to obtain the data output by the dynamic feature fusion module (that is, the fusion features in the aforementioned step S102).
[0139] The fused features, labeled images, and prompt words are input into the multimodal large model. The multimodal large model outputs the inference result (i.e. the first event in step S104 above) under the action of Lora weight 1, Lora weight 2, and Lora weight N (N is the number of object types in the video image data).
[0140] Corresponding to the aforementioned event reasoning method based on multimodal features, this application also provides an event reasoning system based on multimodal features, see [link to relevant documentation]. Figure 7 , Figure 7 The diagram below shows the structure of an event reasoning system provided in an embodiment of this application. The event reasoning system 700 includes multiple image acquisition devices (only image acquisition devices 701, 702, and 703 are shown in the diagram for ease of description) and a processor 704.
[0141] Each image acquisition device is used to capture image data of the target area from different poses and upload the captured image data to the processor 704;
[0142] Processor 704 is used to execute any of the event reasoning methods described above.
[0143] Each image acquisition device has a communication connection with the processor.
[0144] Using the above embodiments, the system includes multiple image acquisition devices and a processor. The image acquisition devices capture image data from different poses and upload it to the processor. The processor fuses the features extracted from the image data captured from different poses to obtain fused features. Compared with existing technologies that analyze each image data source independently or simply stitch them together, event reasoning based on fused features can utilize the correlation information between image data sources from different angles to achieve event reasoning from a global perspective. By identifying existing objects in each image data and inputting the identified images into a multimodal large model, combined with prompts to guide the multimodal large model's thinking, the prompts characterize the identification method of each object in each identified image. Different objects have different identification methods, providing the multimodal large model with clear object differentiation information, enabling it to more accurately identify objects. This method introduces fusion features and labeled images. Fusion features can be regarded as semantic features, and labeled images can be regarded as visual features. Combined with prompt words, it guides the multimodal large model to think. Because it integrates semantic and visual features, the multimodal large model can understand both semantics and vision together, and can understand image information more comprehensively and accurately. This results in outputting more accurate events that occur in the target area within the target time period, i.e., the first event, which improves the accuracy of event reasoning results and thus solves the aforementioned problem one.
[0145] Corresponding to the aforementioned event reasoning method based on multimodal features, this application also provides an event reasoning apparatus based on multimodal features, see [link to relevant documentation]. Figure 8 , Figure 8 A schematic diagram of the structure of the event reasoning device provided in the embodiments of this application includes:
[0146] The acquisition module 801 is used to acquire multiple image data captured by multiple image acquisition devices during a target time period; wherein, the multiple image acquisition devices are used to capture the target area from different poses;
[0147] The fusion module 802 is used to extract image features from each of the image data separately and fuse the extracted image features to obtain fused features;
[0148] The identification module 803 is used to identify the objects present in each of the image data to obtain an identification image; wherein, different objects in the identification image are identified in different ways;
[0149] The output module 804 is used to input the fused features, each of the labeled images, and the prompt words into the multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target area within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
[0150] By applying the above embodiments, fused features are obtained by fusing features extracted from image data captured from different poses by multiple image acquisition devices. Compared with existing technologies that analyze each image data source independently or simply stitch them together, event reasoning based on fused features can utilize the correlation information between image data sources from different angles to achieve event reasoning from a global perspective. By identifying objects in each image data and inputting the identified images into a multimodal large-scale model, combined with prompts to guide the model's thinking, and using prompts to characterize the identification methods of each object in each identified image (different objects have different identification methods), the multimodal large-scale model receives clear object differentiation information, enabling it to more accurately identify objects. This method, through the introduction of fused features and identified images (which can be considered semantic features and identified images as visual features), combined with prompts to guide the multimodal large-scale model's thinking, integrates semantic and visual features, allowing for a more comprehensive and accurate understanding of image information. This results in a more accurate output of the event occurring in the target area within the target time period, i.e., the first event, improving the accuracy of event reasoning results and thus solving the aforementioned problem one.
[0151] In one possible implementation, the fusion module includes:
[0152] The first submodule is used to input the image features into the event detector and obtain the events that occur in the target region within the target time period as the second event; wherein, the algorithm complexity of the event detector is lower than that of the multimodal large model;
[0153] The second sub-module is used to perform weighted fusion of the image features according to the first dynamic weight of each image acquisition device to obtain fused features; wherein, the first dynamic weight of each image acquisition device is positively correlated with the importance of the region captured by the image acquisition device in the second event.
[0154] In one possible implementation, the device further includes:
[0155] An adjustment module is configured to adjust the event detector and / or the first dynamic weight in a direction that makes the second event the same as the first event if a preset difference condition is met between the first event and the second event.
[0156] In one possible implementation, the device further includes:
[0157] The acquisition module is used to obtain a second dynamic weight pre-set for the type of object present in the target area;
[0158] The output module includes:
[0159] The output first submodule is used to input the fused features, each of the labeled images, and the prompt words into the multimodal large model, so that the multimodal large model, guided by the prompt words, uses the second dynamic weight to identify objects existing in the target area, and, based on the identification results, considers and outputs the events that occur in the target area during the target time period as the first event.
[0160] This application also provides an electronic device, such as... Figure 9 As shown, it includes:
[0161] Memory 901 is used to store computer programs;
[0162] When processor 902 executes a program stored in memory 901, it performs the following steps:
[0163] Multiple image data points are acquired from multiple image acquisition devices during a target time period; wherein the multiple image acquisition devices are used to capture images of the target area from different poses.
[0164] Image features are extracted from each of the image data separately, and the extracted image features are fused to obtain fused features;
[0165] Each object present in the image data is identified to obtain an identified image; wherein, different objects in the identified image are identified in different ways;
[0166] The fused features, each of the labeled images, and the prompt words are input into a multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target region within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
[0167] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 902, communication interface, and memory 901 communicating with each other via the communication bus.
[0168] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0169] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0170] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0171] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0172] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described event reasoning methods.
[0173] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the event reasoning methods described in the above embodiments.
[0174] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.
[0175] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0176] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0177] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. An event reasoning method based on multimodal features, characterized in that, The method includes: Multiple image data points are acquired from multiple image acquisition devices during a target time period; wherein the multiple image acquisition devices are used to capture images of the target area from different poses. Image features are extracted from each of the image data separately, and the extracted image features are fused to obtain fused features; The objects present in each of the image data are identified to obtain an identified image; wherein, different objects in the identified image are identified in different ways; the different ways of identifying different objects include: using any one or more combinations of different colors, shapes, text, and numbers to identify different objects; The fused features, each of the labeled images, and the prompt words are input into a multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target region within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
2. The method according to claim 1, characterized in that, The extracted image features are combined to obtain fused features, including: Each of the aforementioned image features is input into an event detector to obtain events occurring in the target region within the target time period as the second event; wherein, the algorithm complexity of the event detector is lower than that of the multimodal large model; According to the first dynamic weight of each image acquisition device, the image features are weighted and fused to obtain fused features; wherein, the first dynamic weight of each image acquisition device is positively correlated with the importance of the region captured by the image acquisition device in the second event.
3. The method according to claim 2, characterized in that, The method further includes: If the first event and the second event meet a preset difference condition, then the event detector and / or the first dynamic weight are adjusted in a direction that makes the second event the same as the first event.
4. The method according to claim 1, characterized in that, The method further includes: Based on the type of objects present in the target area, obtain a second dynamic weight that has been pre-set for the type; The fused features, each of the labeled images, and the prompt words are input into a multimodal large model, so that the multimodal large model, guided by the prompt words, considers and outputs events occurring in the target region within the target time period as the first event, including: The fused features, each of the labeled images, and the prompt words are input into the multimodal large model, so that the multimodal large model, guided by the prompt words, uses the second dynamic weights to identify objects existing in the target area, and, based on the identification results, considers and outputs the events that occur in the target area within the target time period as the first event.
5. An event reasoning system based on multimodal features, characterized in that, The system includes multiple image acquisition devices and a processor; each of the image acquisition devices has a different pose. Each of the aforementioned image acquisition devices is used to capture image data of the target area from different poses and upload the captured image data to the processor; The processor is configured to execute the method according to any one of claims 1-4.
6. An event reasoning device based on multimodal features, characterized in that, The device includes: The acquisition module is used to acquire multiple image data captured by multiple image acquisition devices during a target time period; wherein, the multiple image acquisition devices are used to capture the target area from different poses; The fusion module is used to extract image features from each of the image data separately and fuse the extracted image features to obtain fused features; The identification module is used to identify the objects present in each of the image data to obtain an identification image; wherein, different objects in the identification image are identified in different ways; the different ways of identifying different objects include: using any one or more combinations of different colors, shapes, text, and numbers to identify different objects; The output module is used to input the fused features, each of the labeled images, and the prompt words into the multimodal large model, so that the multimodal large model, guided by the prompt words, thinks about and outputs the events that occur in the target area within the target time period as the first event; wherein, the prompt words are used to characterize the identification method of each object in each of the labeled images.
7. The apparatus according to claim 6, characterized in that, The fusion module includes: The first submodule is used to input the image features into the event detector and obtain the events that occur in the target region within the target time period as the second event; wherein, the algorithm complexity of the event detector is lower than that of the multimodal large model; The second sub-module is used to perform weighted fusion of the image features according to the first dynamic weight of each image acquisition device to obtain fused features; wherein, the first dynamic weight of each image acquisition device is positively correlated with the importance of the region captured by the image acquisition device in the second event.
8. The apparatus according to claim 7, characterized in that, The device further includes: An adjustment module is used to adjust the event detector and / or the first dynamic weight in a direction that makes the second event the same as the first event if a preset difference condition is met between the first event and the second event. The device further includes: The acquisition module is used to obtain a second dynamic weight pre-set for the type of object present in the target area; The output module includes: The output first submodule is used to input the fused features, each of the labeled images, and the prompt words into the multimodal large model, so that the multimodal large model, guided by the prompt words, uses the second dynamic weight to identify objects existing in the target area, and, based on the identification results, considers and outputs the events that occur in the target area during the target time period as the first event.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-4.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-4.