A Full-Scene Intelligent Analysis System and Method Based on Multimodal Large Model
By integrating cross-modal data and mining causal and temporal correlations through a multimodal big data model full-scene intelligent analysis system, the system solves the problem of insufficient detection capabilities of existing security systems in unknown risk scenarios, and achieves accurate identification and dynamic supervision of violations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-04-03
AI Technical Summary
Existing security systems struggle to detect new types of violations when an individual's identity features are not found in the initial permission database, and lack the ability to generalize detection capabilities for complex or unknown risk scenarios.
A full-scenario intelligent analysis system based on a multimodal large model is adopted. The feature extraction module integrates cross-modal data to generate a set of temporal features and spatial structures. Combined with the semantic fusion module and the causal generation module, the system explores the causal temporal correlations between cross-modal data to achieve accurate identification of violations and abnormal behaviors.
It significantly improves the intelligent supervision capabilities of key locations, can cover complex or unknown risk scenarios, supports real-time visualization and dynamic prediction, and enhances the generalized detection capabilities for new types of violations.
Smart Images

Figure CN120953913B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of security monitoring systems, and in particular to a full-scene intelligent analysis system and method based on a multimodal large model. Background Technology
[0002] With the acceleration of urbanization and the increasing severity of social security problems, the construction of security monitoring systems has gradually become an important part of modern social security management. Security monitoring systems can monitor specific areas in real time through video surveillance technology, promptly detect and respond to various emergencies, and provide a guarantee for maintaining public safety.
[0003] Existing security systems rely on multimodal monitoring data, such as surveillance image data, audio data, and environmental information data, to predict the behavior of individuals when they are in security blind spots within a security area. At the same time, they combine image feature classification algorithms to classify individual behavioral characteristics, thereby enabling the classification of the behavior level of individuals when they are in security blind spots and to make corresponding alarm responses based on the behavior level.
[0004] However, existing security systems, when determining that an individual's identity characteristics are not in the initial permission database, mark the permissions associated with those characteristics as the monitoring level and perform corresponding monitoring. Therefore, they lack the ability to generalize detection capabilities for new types of violations and struggle to cover complex or unknown risk scenarios. Summary of the Invention
[0005] To improve the ability to detect new types of violations and thus cover complex or unknown risk scenarios, this application provides a full-scenario intelligent analysis system and method based on a multimodal large model.
[0006] Firstly, this application provides a full-scene intelligent analysis system based on a multimodal large model, employing the following technical solution:
[0007] A full-scene intelligent analysis system based on a multimodal large model includes:
[0008] The feature extraction module is used to acquire cross-modal data, which includes cross-modal tensors and corresponding cross-modal images. The feature extraction module is also used to generate a temporal feature set based on the cross-modal tensors and a spatial structure set based on the cross-modal image data.
[0009] A semantic fusion module is network-connected to the feature extraction module to receive the temporal feature set and the spatial structure set, and performs semantic interaction on the temporal feature set and the spatial structure set based on cross-modal structure to obtain a cross-modal data set.
[0010] A causal generation module is network-connected to the feature extraction module to receive the temporal feature set and the spatial structure set, and to generate causal associations based on the temporal feature set and the spatial structure set; the causal generation module is also network-connected to the semantic fusion module to receive the cross-modal data set, and to generate causal temporal associations based on the cross-modal data set and the causal associations.
[0011] A dynamic prediction module, together with the causal generation module, obtains the causal time series association and dynamically learns the association mechanism based on the causal time series association to generate a dynamic prediction model.
[0012] By adopting the above technical solutions, the feature extraction module extracts features from cross-modal data, integrates traditional time-series sensor data with video surveillance data across modalities, breaks down information silos, and achieves semantic-level linkage and unified processing between data. It performs high-level feature extraction and deep semantic fusion on time-series sensor signals and video images to generate cross-modal data sets. Through the causal generation module, it mines potential causal temporal correlations between cross-modal data, thereby accurately identifying illegal and abnormal behaviors in the controlled area, significantly improving the intelligent supervision capabilities of key locations, and further enhancing the generalized detection capabilities for new types of illegal behaviors, thus covering complex or unknown risk scenarios.
[0013] In some embodiments, a preliminary data processing module is also included. This module performs timestamp alignment on the heterogeneous sensing data corresponding to the region under control based on timestamp alignment technology to generate mapped timeline data. The preliminary data processing module is also used to filter and standardize the mapped timeline data to obtain standardized data, and to construct a cross-modal model based on the standardized data to obtain cross-modal tensors.
[0014] By adopting the above technical solution, the heterogeneous sensing data in the controlled area is preliminarily processed based on the data preliminary processing module, thereby ensuring that the input data for subsequent cross-modal fusion is accurate and consistent.
[0015] In some embodiments, the data preprocessing module is further configured to preprocess the image data of heterogeneous sensor data to obtain key target regions, perform human image detection on the key target regions based on the detection model to obtain corresponding human image annotation information, and obtain cross-modal images based on the human image annotation information.
[0016] In some embodiments, an automatic tracing module is also included, which is further used to receive video image data of the area to be controlled in real time, and to perform anomaly detection on the video image data based on the dynamic prediction module to obtain detection results, generate tracing data based on the detection results, and send the detection results and the tracing data to the management terminal.
[0017] By adopting the above technical solution, the automatic tracing module receives video image data of the area to be controlled in real time, performs anomaly detection on the video image data to obtain detection results, generates tracing data based on the detection results, and sends the detection results and the tracing data to the management terminal. Through a dynamic causal inference method, combined with the changing trend of environmental parameters and the historical trajectory of individual behavior, the module automatically learns the dynamic correlation mechanism between the two, thereby realizing the causal explanation and tracing analysis of complex behavioral patterns.
[0018] In some embodiments, an interactive decision-making module is also included, which is used to receive video image data of the area to be controlled in real time, and automatically identify key information based on the video image data by the dynamic prediction module; and generate key target behavior trajectories and long-term trend statistical charts based on the key information.
[0019] By adopting the above technical solutions, administrators can quickly understand the overall dynamics at the macro level and accurately locate key abnormal events at the micro level, thereby improving the comprehensiveness and immediacy of scene perception.
[0020] Secondly, this application provides a full-scene intelligent analysis method based on a multimodal large model, employing the following technical solution:
[0021] A full-scene intelligent analysis method based on a multimodal large model, executed based on the full-scene intelligent analysis system based on a multimodal large model as described in the first aspect, includes the following steps:
[0022] Acquire cross-modal data, and generate a temporal feature set and a spatial structure set based on the cross-modal data;
[0023] Based on the cross-modal structure, semantic interaction is performed on the temporal feature set and the spatial structure set to obtain the cross-modal data set;
[0024] Causal relationships are generated based on the temporal feature set and the spatial structure set, and causal temporal relationships are generated based on the cross-modal data set and the causal relationships.
[0025] Based on the aforementioned causal temporal correlation dynamic learning association mechanism, a dynamic prediction model is generated.
[0026] In some embodiments, semantic interaction is performed on the temporal feature set and the spatial structure set based on cross-modal structure to obtain a cross-modal data set, including the following steps:
[0027] The temporal feature set and the spatial structure set are mapped to the same dimension to obtain combined dimensional features. The combined maintenance features include temporal features and structural features, and the temporal features and structural features are in one-to-one correspondence.
[0028] The temporal and structural features are interactively processed based on the temporal guidance space to obtain first interactive data, and the temporal and structural features are interactively processed based on the spatial constraint temporal to obtain second interactive data.
[0029] Based on the temporal and structural features, corresponding data weights are generated, and the first and second interactive data are adjusted based on the data weights to generate a cross-modal data set.
[0030] By adopting the above technical solution, the temporal feature set and the spatial structure set are processed based on the cross-modal structure. The first interactive data is generated based on the temporal guidance space, and the second interactive data is obtained based on the spatial constraint temporal. The first interactive data and the second interactive data are filtered based on the data weight to obtain the cross-modal data set. Thus, the temporal features and structural features can be processed from two aspects to obtain an accurate cross-modal data set, thereby generating an accurate dynamic prediction model and realizing the safety monitoring of the area to be controlled.
[0031] In some embodiments, generating corresponding data weights based on the temporal and structural features includes the following steps:
[0032] The first indicator feature is obtained based on the time series features, and the first indicator feature is normalized to obtain the time series information.
[0033] Based on the structural features, a second indicator feature is obtained, and the second indicator feature is normalized to obtain spatial information.
[0034] Based on the temporal and structural features, corresponding cross-modal data are obtained, and cross-modal correlation is generated based on the cross-modal data;
[0035] The data weight is calculated based on the temporal information, spatial information, and cross-modal correlation quantification attribute index, and a preset fusion strategy.
[0036] By adopting the above technical solution, the data proportion is calculated based on the first indicator feature, the second indicator feature, and the cross-modal correlation degree, thereby obtaining an accurate cross-modal data set, generating an accurate dynamic prediction model, improving the intelligent supervision capability of key locations, and enhancing the generalized detection capability of new types of violations, thus covering complex or unknown risk scenarios.
[0037] In some embodiments, generating causal relationships based on the temporal features and the spatial structure information includes the following steps:
[0038] Based on the temporal features and spatial structure, causal elements are obtained, and the temporal features are detected based on the causal elements to obtain key events;
[0039] Based on the causal factors, the spatial structure is analyzed to obtain the changes corresponding to the key events.
[0040] A causal relationship set is generated based on the key events and the changes, and the causal relationship set is verified based on the verification mode to obtain the causal relationships.
[0041] In some embodiments, detecting the temporal features based on the causal elements to obtain key events includes the following steps:
[0042] Based on the temporal features and the spatial structure, feature spatial data on the time chain is obtained. The feature spatial data includes time points, temporal-series sensor data, and corresponding spatial data.
[0043] The abrupt change time is obtained based on the time-series sensor data, and key events are obtained based on the abrupt change time and the spatial data.
[0044] In summary, this application includes at least one of the following beneficial technical effects:
[0045] 1. The feature extraction module extracts features from cross-modal data, integrating traditional time-series data with video surveillance data across modalities, breaking down information silos, and achieving semantic linkage and unified processing between data. It performs high-level feature extraction and deep semantic fusion on time-series sensor signals and video images to generate cross-modal data sets. Through the causal generation module, it mines potential causal temporal correlations between cross-modal data, thereby accurately identifying illegal and abnormal behaviors in the controlled area, significantly improving the intelligent supervision capabilities of key locations, and further enhancing the generalized detection capabilities for new types of illegal behaviors, thus covering complex or unknown risk scenarios.
[0046] 2. The preliminary data processing module performs preliminary processing on the heterogeneous sensor data within the controlled area to ensure that the input data for subsequent cross-modal fusion is accurate and consistent;
[0047] 3. Supports real-time visualization, including multi-channel monitoring video footage, automatically identified event summary information, key target behavior trajectories, and long-term trend statistical charts, helping administrators quickly understand the overall dynamics at a macro level and accurately locate key abnormal events at a micro level, improving the comprehensiveness and immediacy of scene perception. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the structure of a full-scene intelligent analysis system based on a multimodal large model provided in an embodiment of this application;
[0049] Figure 2 This is a block diagram of the full-scene intelligent analysis method based on a multimodal large model provided in the embodiments of this application;
[0050] Figure 3 This is a block diagram of a method for acquiring cross-modal data sets provided in an embodiment of this application;
[0051] Figure 4 This is a block diagram of the data weighting acquisition method provided in the embodiments of this application;
[0052] Figure 5 This is a block diagram of the causal association method provided in the embodiments of this application.
[0053] Figure labeling: 10, Feature extraction module; 20, Semantic fusion module; 30, Causal generation module; 40, Dynamic prediction module; 50, Preliminary data processing module; 60, Automatic source tracing module; 70, Interactive decision-making module. Detailed Implementation
[0054] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.
[0055] This application discloses a full-scene intelligent analysis system based on a multimodal large model.
[0056] like Figure 1As shown, the full-scene intelligent analysis system based on a multimodal large model includes a feature extraction module 10, a semantic fusion module 20, a causal generation module 30, and a dynamic prediction module 40. The feature extraction module 10 acquires cross-modal data and generates a temporal feature set and a spatial structure set based on the cross-modal data. The semantic fusion module 20 is network-connected to the feature extraction module 10 to receive the temporal feature set and the spatial structure set, and performs semantic interaction on the temporal feature set and the spatial structure set based on the cross-modal structure to acquire the cross-modal data set. The causal generation module 30 is network-connected to the feature extraction module 10 to receive the temporal feature set and the spatial structure set, and generates causal relationships based on the temporal feature set and the spatial structure set. The causal generation module 30 is network-connected to the semantic fusion module 20 to receive the cross-modal data set and generates causal temporal relationships based on the cross-modal data set and causal relationships. The dynamic prediction module 40 is connected to the causal generation module 30 to acquire causal temporal relationships and dynamically learns the relationship mechanism based on the causal temporal relationships to generate a dynamic prediction model.
[0057] Cross-modal data includes cross-modal tensors and corresponding cross-modal images. The cross-modal tensor representation is an input tensor of a unified dimension formed by standardizing temporal sensor data and concatenating it along its dimensions. The cross-modal image representation is an image obtained by preprocessing video image data. The preprocessing includes enhancements such as brightness and contrast, which improves the visual information quality of the obtained cross-modal images, facilitating subsequent feature extraction.
[0058] The time-series sensor data mentioned here refers to the data collected by the data acquisition subsystem. This subsystem includes temperature and humidity sensors, access control / entry controllers, RFID identification devices, and video surveillance equipment. The temperature and humidity sensors collect environmental parameters in real time, the access control / entry controllers record personnel entry and exit information, and the RFID identification devices determine whether personnel are exhibiting any inappropriate behavior by linking their identity to their behavioral trajectory. The video image data is collected by the video surveillance equipment, which monitors the controlled area in real time and acquires video image data.
[0059] Furthermore, regarding how to perform cross-modal processing on time-series sensor and video image data to obtain corresponding cross-modal data, the specific method is to integrate traditional IoT data, such as data collected from temperature and humidity, access control, and RFID, with video surveillance data to achieve semantic-level linkage and unified processing between data.
[0060] It should be noted that cross-modal data is mainly obtained by fusing data collected by the data acquisition subsystem. The data collected by the data acquisition subsystem comes from different sources, including time-series sensor data and video image data. Data from different sources may have semantic conflicts, temporal or spatial alignment issues. Therefore, the data preliminary processing module 50 needs to perform preliminary processing on the data collected by the data acquisition subsystem. The resulting cross-modal data is then readily extracted by the feature extraction module 10 to obtain a set of temporal features and a set of spatial structures.
[0061] The temporal feature set mainly includes the state values at each time point, the frequency and amplitude of periodic fluctuations. These features are automatically extracted using temporal models such as LSTM, ultimately forming a fixed-dimensional feature vector. The spatial structure set, on the other hand, is a set of features extracted from cross-modal images related to spatial morphology and topology. It primarily reflects the spatial distribution patterns and structural characteristics of data during the cross-membrane process. It typically includes key time point markers, state durations, and state interval statistics. Key time point markers include moments of state abrupt changes and extreme points; state durations include the duration of a stable state; and state interval statistics include the time difference between two consecutive state changes. Here, "state" refers to the behavioral state of pedestrians in the controlled area.
[0062] In one embodiment, a preliminary data processing module 50 is also included. This module aligns the heterogeneous sensing data corresponding to the controlled region using timestamp alignment technology to generate mapped timeline data. The preliminary data processing module 50 further filters and standardizes the mapped timeline data to obtain standardized data, and constructs a cross-modal model based on the standardized data to obtain cross-modal tensors.
[0063] The heterogeneous sensor data includes time-series sensor data and video image data. Timestamp alignment techniques include interpolation, resampling, time offset correction, and event mapping. In this embodiment, interpolation is used to align the timestamps of the heterogeneous sensor data. The specific timetamp alignment operation based on interpolation is carried out using existing technologies, which will not be elaborated on here.
[0064] Mapped timeline data represents data generated by aligning heterogeneous sensors using timestamp technology, and this mapped timeline data is displayed using a unified target timeline method. For ease of access, mapped timeline data is obtained by timestamp alignment of multi-source heterogeneous data, and it possesses spatiotemporal consistency.
[0065] For example, suppose there are three sensors sampling at frequencies of 10Hz, 5Hz, and 20Hz respectively, with different start times. After timestamp alignment, the mapped timeline data will generate a timeline with a fixed frequency, which can be set to 10Hz. The data from the three sensors are converted into sequences on this timeline, with corresponding values from the three sensors at specific time points, such as t=0.1s, t=0.2s, etc.
[0066] Timestamp alignment eliminates the temporal heterogeneity of multi-source data, enabling direct comparison and correlation of data from different sources over time. This provides a data foundation for subsequent synchronous analysis. Standardized time intervals facilitate the calculation of time-series features. Synchronous analysis includes multi-parameter correlation calculations at the same time point, and time-series feature calculations such as sliding window statistics and time difference analysis.
[0067] Standardized data refers to data obtained by denoising and standardizing the mapped time axis data. The denoising process for the mapped time axis data mainly involves using filters. This application uses Kalman filtering to remove environmental and equipment noise. The details of how the filtering is performed will not be elaborated here.
[0068] First, determine the time range for data collection from different sensors. Since different sensors have different collection frequencies, data collected within the same time range may be incomplete. Therefore, the collection frequency of different sensors can be adjusted according to the time node for data processing of the area under control. Of course, this means adjusting before the data processing time node, so that all sensors collect data at the same time when the data is processed, thus avoiding data loss.
[0069] It should be noted that, regarding how to adjust all sensors, a data processing signal can be generated based on the data processing time point, and this signal can be sent to a designated sensor to control the designated sensor to adjust its data acquisition frequency. The designated sensor mentioned here is one that does not perform any data acquisition at the specified time point, selected from all sensors through a screening process. The screening process utilizes existing technologies for comparison, which will not be elaborated upon here.
[0070] In addition, for heterogeneous sensor data, in order to reduce data loss, when processing heterogeneous sensor data, it is necessary to ensure that all sensors have data at the start and end points.
[0071] Next, a unified target timeline is generated from the heterogeneous sensor data, and data types are generated based on this heterogeneous data. Different alignment methods are automatically adopted based on the data type, including linear interpolation and resampling. Linear interpolation is suitable for continuously changing data, while resampling is suitable for discretely sampled data. The time-series sensor data collected by the temperature and humidity sensor is continuously changing, while the access control / entry controller collects discretely sampled data recording personnel entry and exit information. Here, different alignment methods can be used based on the data type of the heterogeneous sensor data, improving the automatic data processing capability of the preliminary data processing module 50.
[0072] Furthermore, the preliminary data processing module 50 employs the 3σ principle for outlier detection and filtering to ensure data quality. It uses standardization methods to eliminate dimensional differences between data from different sensors and retains timestamp information to ensure the time traceability of subsequent processing. Finally, a three-dimensional tensor is constructed based on a sliding time window. The tensor dimensions include the number of samples, time step, and number of sensors. Each tensor sample contains all sensor data within a specified time window, facilitating subsequent temporal feature extraction. The standardization method will not be elaborated upon here.
[0073] It should be noted that the construction of the 3D tensor uses existing technology, which will not be elaborated upon here. The sliding time window can be selected from preset trigger nodes in historical data. Trigger nodes mainly refer to the presence of suspicious environmental parameters. The sliding time window is set based on the intervals at which these suspicious environmental parameters appear. This sliding time window can be adjusted according to the actual situation, allowing for more precise data processing and accurate results. When no suspicious environmental parameters appear, the sliding time window can be set to a fixed processing frequency. The fixed processing frequency is a pre-set data processing frequency, and the value of the fixed processing frequency will not be elaborated upon here.
[0074] In one embodiment, the data preprocessing module 50 is further configured to preprocess the image data of the heterogeneous sensor data to obtain key target regions, perform human image detection on the key target regions based on the detection model to obtain corresponding human image annotation information, and obtain cross-modal images based on the human image annotation information.
[0075] The key target region refers to the area within the controllable region where environmental parameters change. These parameters can include temperature, humidity, behavioral states, etc. The key target region includes the main target region, the interaction region, and the salient feature region. The main target region refers to the object in the image that carries the core semantics, such as the outline of a pedestrian in a surveillance video. The interaction region is the area where actions occur between targets, such as the movements of pedestrians or the contact areas between pedestrians. The salient feature region contains highly identifiable local features, such as the facial features of a pedestrian or the area of a pedestrian's body movements.
[0076] Key target regions can be extracted using region of interest (ROI) techniques, which will not be elaborated upon here. Instead, a detection model is used to detect people within the key target regions to obtain corresponding person annotation information. Detection models such as Faster R-CNN, YOLO, and SSD can be employed. The person annotation information is obtained by detecting human figures within the key target regions. The annotated dataset is divided into training, validation, and test sets, which are then input into the model for training. During training, the model parameters are adjusted to minimize the loss function, enabling the model to learn the features of people in the images and accurately detect them.
[0077] The semantic fusion module 20 performs semantic interaction on the temporal feature set and spatial structure set based on cross-modal structure to obtain a cross-modal data set. Cross-modal images typically refer to pedestrian images in different modalities, such as visible light-infrared cross-modal pedestrian images. First, the specific modal type of the required cross-modal image must be clearly defined, such as obtaining the corresponding infrared image from a visible light image, or other modal transformations. If an image needs to be generated from a text description, a text-to-image generation model can be selected. Through deep learning algorithms, the text description is encoded into a vector representation, and then the image is generated through a decoder. If a person image in one modality already exists, and another modality image needs to be generated, cross-modal pedestrian re-identification methods based on data augmentation and graph matching can be referenced. For example, a recurrent generative adversarial network with a u-net structure can be used to input a visible light domain pedestrian image into the network to generate a pseudo-infrared image, or an infrared domain pedestrian image can be input to generate a pseudo-visible light image.
[0078] Cross-modal datasets represent temporal features within the same time chain and the corresponding spatial structural features. In this embodiment, the cross-modal structure can adopt a Transformer architecture, using a multi-head attention mechanism to achieve semantic alignment and deep interaction between different temporal features and visual features, strengthening the complementarity and consistent expression between modalities. This mechanism can automatically capture global dependencies and fine-grained differences in behavioral patterns, effectively improving cross-modal behavior understanding and prediction capabilities.
[0079] Building upon semantic fusion, we further construct causal time-series correlations between environmental data, such as temperature and humidity, and RFID signals, and human behavior. Through dynamic causal inference methods, combining the changing trends of environmental parameters with the historical trajectories of individual behaviors, we automatically learn the dynamic correlation mechanism between the two, thereby achieving causal explanation and source analysis of complex behavioral patterns.
[0080] The causal generation module 30 generates causal associations based on the temporal feature set and the spatial structure set; the causal generation module 30 is network-connected to the semantic fusion module 20 to receive the cross-modal data set, and generates causal temporal associations based on the cross-modal data set and the causal associations. The dynamic prediction module 40 obtains the causal temporal associations from the causal generation module 30, and dynamically learns the association mechanism based on the causal temporal associations to generate a dynamic prediction model.
[0081] Based on the rich prior knowledge of the pre-trained large model, the dynamic prediction model can accurately identify typical personnel behaviors in multi-modal input data, including but not limited to various behavioral states such as standing, walking, gathering, abnormal staying, running fast, and standing still for a long time. On the basis of the identification, the dynamic prediction model further combines the historical behavior data accumulated in the long term to construct a behavior pattern library, and compares the currently detected behavior features to automatically determine the degree of deviation from the normal pattern. When an abnormal deviation is detected, the system will perform multi-factor linkage reasoning by combining real-time environmental data, and comprehensively generate a risk assessment result, significantly improving the accurate recognition rate and scene adaptability of abnormal events. Among them, the environmental data includes abnormal temperature and humidity, abnormal RFID signals, abnormal access control, etc.
[0082] In addition, the dynamic prediction model also introduces generative model techniques, such as generative adversarial networks, diffusion models, etc., which will not be elaborated here. The generative model technique is used to automatically construct diverse simulated abnormal behavior events. The system is supplemented and trained with the generated abnormal data, effectively improving the generalization detection ability of the model for unseen behavior patterns, and enhancing the robustness and agile response ability of the system in complex, dynamic, and changeable environments.
[0083] Refer to Figure 1 , in one of the embodiments, it further includes an automatic traceability module 60. The automatic traceability module 60 is also used to receive the video image data of the controlled area in real time, perform abnormal detection on the video image data to obtain the detection result, generate traceability data based on the detection result, and send the detection result and the traceability data to the management end.
[0084] Among them, the video image data is the data monitored by the video monitoring equipment in the controlled area, and abnormal detection is performed on the video image data. The detection results include no risk, risk exists, and there is a risk, etc. No risk indicates that the video image data is monitored and analyzed, and there is no risk situation. Risk exists indicates that there is a possibility of risk in the controlled area, and there is a risk indicates that there is a risk in the controlled area, and a response plan needs to be provided in time, for example, informing the staff in advance or issuing a warning notice.
[0085] The automatic traceability module 60 incorporates security and traceability design for the entire controlled area, enabling rapid response, accurate analysis, and comprehensive tracing of detected abnormal behavior. When the full-scene intelligent analysis system based on a multimodal large model detects suspected abnormal behavior, the module automatically generates a complete structured evidence package, including multidimensional information: key video frames, IoT sensor data logs, personnel behavior tags, environmental parameter records, and time-series graphs based on multimodal features. Through a multi-layered, structured chain of evidence, the event process is fully reconstructed, ensuring the objectivity and completeness of subsequent evidence collection and tracing.
[0086] To trace the source of suspected abnormal behavior, multi-dimensional information can be used to mark the abnormal behavior, and the marked data can be displayed on a target timeline, making it easier to view the source data of suspected abnormal behavior later. It should be noted that suspected abnormal behavior at different times can be displayed on the target timeline.
[0087] Furthermore, the full-scenario intelligent analysis system based on a multimodal large model supports one-click behavior trajectory retrospection and risk cause analysis. Administrators can quickly replay the historical behavior trajectories of target individuals or groups through a visual interface, and automatically analyze the triggering factors, evolution process, and potential risk causes of abnormal behavior by combining environmental and IoT data. This function significantly improves the efficiency and accuracy of analyzing complex events, helping to quickly locate security vulnerabilities or management weaknesses.
[0088] Meanwhile, the full-scenario intelligent analysis system based on a multimodal large model can push early warning information, abnormal behavior summaries, risk levels, and on-site handling suggestions to the management terminal in real time. Combined with the generated decision support information, managers can grasp the on-site dynamics immediately and take timely and targeted measures, such as isolating areas, guiding evacuation, and on-site intervention, effectively improving the intelligence and automation level of the overall safety response.
[0089] Reference Figure 1 In one embodiment, the system further includes an interactive decision module 70, which is used to receive video image data of the area to be controlled in real time, and automatically identify key information based on the video image data; and generate key target behavior trajectories and long-term trend statistical charts based on the key information.
[0090] The interactive decision-making module 70 acts as a bridge between the system and administrators, providing a highly integrated, easy-to-use, and intelligent visual interactive experience based on multimodal intelligent analysis results. It supports real-time visualization, including multi-channel monitoring video feeds, automatically identified event summary information, key target behavior trajectories, and long-term trend statistical charts, helping administrators quickly understand the overall dynamics at a macro level and accurately locate key abnormal events at a micro level, improving the comprehensiveness and immediacy of scene perception.
[0091] The multimodal large-scale intelligent analysis system for all scenarios has built-in risk event management and reporting functions, supporting multi-condition queries, precise filtering, and structured downloads of historical risk event records. The generated risk event reports include detailed event descriptions, behavioral identification tags, environmental data snapshots, timeline graphs, and system analysis recommendations, facilitating subsequent archiving, external reporting, and accountability.
[0092] Furthermore, the interactive decision module 70 also provides an open model correction and feedback mechanism, allowing administrators to manually correct, label, or confirm the system's recognition results based on behavior labels, abnormal event logs, and trajectory playback results. Administrator feedback will be automatically recorded by the multimodal large-scale model-based full-scene intelligent analysis system and used for subsequent supervised learning and model adaptive optimization, continuously improving the system's recognition accuracy and scene generalization ability, achieving a closed-loop iteration of human-machine collaboration.
[0093] This application also discloses a full-scene intelligent analysis method based on a multimodal large model.
[0094] like Figure 2 As shown, the full-scene intelligent analysis method based on a multimodal large model includes the following steps:
[0095] S100: Acquire cross-modal data and generate a time-series feature set and a spatial structure set based on the cross-modal data.
[0096] S200 uses cross-modal structures to perform semantic interaction on temporal feature sets and spatial structure sets to obtain cross-modal data sets.
[0097] S300 generates causal relationships based on temporal feature sets and spatial structure sets, and generates causal temporal relationships based on cross-modal data sets and causal relationships.
[0098] S400 is based on a dynamic learning association mechanism for causal temporal correlation to generate dynamic prediction models.
[0099] The features mentioned in steps S100-S400 above are the same as or similar to those in the full-scene intelligent analysis system based on a multimodal large model, and will not be elaborated further here. Based on the dynamic prediction model generated in steps S100-S400 above, after the full-scene intelligent analysis system based on a multimodal large model is subsequently started, the real-time monitored heterogeneous sensor data is analyzed to obtain the corresponding monitoring results.
[0100] Reference Figure 3 In one embodiment, semantic interaction is performed on the temporal feature set and the spatial structure set based on cross-modal structure to obtain a cross-modal data set, including the following steps:
[0101] S210 maps the temporal feature set and the spatial structure set to the same dimension to obtain combined dimensional features.
[0102] S220, based on the temporal guidance space, interactive processing of temporal features and structural features is performed to obtain the first interactive data, and at the same time, based on the spatial constraint temporal, interactive processing of temporal features and structural features is performed to obtain the second interactive data.
[0103] S230, generate corresponding data weights based on temporal and structural features, and adjust the first and second interactive data based on the data weights to generate a cross-modal data set.
[0104] The combined maintenance features include temporal features and structural features, and these features are in a one-to-one correspondence. Temporal features are based on the temporal characteristics of heterogeneous sensor data, including temporal statistical features, frequency features, and temporal dependency features. Structural features are based on the spatial characteristics of heterogeneous sensor data over time, and these features include pedestrian states, which include, but are not limited to, various behavioral states such as standing, walking, gathering, abnormal lingering, running rapidly, and prolonged stillness.
[0105] The temporal guidance space is based primarily on temporal features and secondarily on spatial features. It learns which spatial regions are related to the current temporal state and uses the current temporal state and the spatial regions related to the current temporal state as the first interaction data. The first interaction data also shows the correlation between the temporal state and the spatial regions.
[0106] Spatial-constrained temporal series refers to learning which time nodes are related to the current spatial structure, with spatial features as the primary focus and temporal features as a secondary focus, and then using the obtained relevant temporal states and spatial structures as secondary interactive data.
[0107] Data weighting refers to the trade-off between the two types of interactive data. It mainly considers the relative impact of temporal and structural features on the monitoring system, and then adjusts the first and second interactive data accordingly. This enables the system to react quickly to different behavioral states in the controlled area, effectively improving the model's ability to generalize detection of unseen behavioral patterns and enhancing the system's robustness and agile response in complex, dynamic, and changeable environments.
[0108] In the initial application of a multimodal large-scale model-based full-scene intelligent analysis system, the data weight can be set to 0.5. By filtering the first and second interaction data according to their weights, the actual cross-modal data set can be obtained. In subsequent monitoring, the data weights are adaptively adjusted by continuously learning from the heterogeneous sensor data obtained during monitoring. The ratio of the number of first interaction data to the number of second interaction data in the cross-modal data set is the data ratio.
[0109] The adjustment of data weight can be achieved by real-time analysis of the currently monitored heterogeneous sensor data to obtain a predicted result, which is then compared with the actual result. If the difference between the predicted and actual results is within a preset range, the data weight does not need to be changed. If the difference between the predicted and actual results is significant, the data weight needs to be automatically adjusted. The direction of data weight adjustment depends on whether the predicted result shows a predictive tendency. If a predictive tendency exists, the data weight can continue to increase in the same direction as its previous growth pattern. If no predictive tendency exists, the adjustment should be made in the opposite direction to the previous growth pattern.
[0110] Regarding how to adjust the first and second interactive data based on data weight to generate a cross-modal data set, the first and second interactive data are quantitatively filtered to combine them into a new cross-modal data set.
[0111] Reference Figure 4 In one embodiment, generating corresponding data weights based on temporal and structural features includes the following steps:
[0112] S231, obtain the first indicator feature based on the time series features, and normalize the first indicator feature to obtain the time series information.
[0113] S232, obtain the second indicator feature based on the structural features, and normalize the second indicator feature to obtain the spatial information.
[0114] S233: Obtain corresponding cross-modal data based on temporal and structural features, and generate cross-modal correlation degree based on cross-modal data.
[0115] S234 calculates the data weight based on temporal information content, spatial information content, and cross-modal correlation quantification attribute indicators, and based on a preset fusion strategy.
[0116] The first indicator features include feature variance, information entropy, and key event density. The time series information is mainly used to normalize the above first indicator features based on the data fluctuation pattern, key event density, and statistical significance, and thus obtains the information.
[0117] The second indicator features include spatial entropy, target confidence, and gradient energy. By normalizing the second indicator features, spatial information is obtained. Cross-modal correlation strength is used to measure the semantic consistency between temporal features and spatial features. The higher the value, the more complementary or consistent the information conveyed by the two modalities in the current task.
[0118] The cross-modal correlation degree is generated by calculating the similarity between temporal features and structural features, then statistically analyzing the dependence relationship of features in the two modalities, and generating the cross-modal correlation degree based on the similarity and dependence relationship. Combining the temporal information volume, spatial information volume and cross-modal correlation degree, the temporal feature weight and spatial feature weight are obtained. The preset fusion strategy refers to the fusion ratio of temporal features and spatial features set in advance, which will not be elaborated here.
[0119] Referring to Figure 5 , in one of the embodiments, generating a causal association based on temporal features and spatial structure information includes the following steps:
[0120] S310, obtaining causal factors based on temporal features and spatial structure, and detecting the temporal features based on the causal factors to obtain key events.
[0121] S320, obtaining the change result corresponding to the key event for the spatial structure based on the causal factors.
[0122] S330, generating a causal association set based on the key event and the change result, and verifying the causal association set based on the verification mode to obtain the causal association.
[0123] Among them, the causal factors refer to the time sequence, spatial correlation and non-coincidence between temporal features and spatial features, etc. The key event refers to an event that occurs due to the change of temporal features. Specifically, a mutation point detection algorithm can be used to identify the significant change points in the time series. The significant change points include peaks / steps at time t0. Among them, the mutation point detection algorithm includes the ADF test and the PELT algorithm. The change result refers to the final presentation state corresponding to the causal factors in the spatial structure. The specific analysis process is to perform differential analysis on the spatial features to detect the spatial changes that occur after time t1, such as the artificial movement positions and human behavior states in the area to be controlled.
[0124] The causal association set represents a set formed by the corresponding relationship between the key event and the change result. The verification mode can adopt time sequence verification and spatial correlation verification. The time sequence verification ensures that t1 < t2, and the human behavior state changes only after the time sequence event ends, and the time interval Δt = t2 - t1. The spatial correlation verification is to calculate the distance between the source of the time sequence event, such as the position of the sensor, and the area to be controlled. If the distance is less than the causal action range, the causal possibility is enhanced.
[0125] In addition, it is also necessary to analyze through a causal graph (CausalGraph) to exclude third-party factors that affect both time sequence and space, such as equipment influence, other behavior influence, etc.
[0126] In one embodiment, detecting temporal features based on causal factors to obtain key events includes the following steps:
[0127] S311, based on temporal features and spatial structure, obtains feature spatial data on the time chain.
[0128] S312 obtains the time of abrupt change based on time-series sensor data, and obtains key events based on the time of abrupt change and spatial data.
[0129] The feature spatial data includes time points, time-series sensor data, and corresponding spatial data. The mutation time specifically refers to the point in time when the time-series sensor data experiences a sudden change, such as a sudden drop or rise in temperature. Critical events refer to events that occur after the mutation time, such as a person's agitated behavior.
[0130] The implementation principle is as follows:
[0131] Based on the dynamic prediction model generated in steps S100-S400, after the subsequent launch of the multimodal large-scale model-based full-scene intelligent analysis system, various sensing devices begin to synchronously collect environmental and image data. The data preprocessing module performs unified transformation and timestamp alignment of multimodal data to form standard input. The encoding module extracts local and global features from video and IoT data respectively, and then achieves deep fusion through multi-layer Transformers. The behavior recognition module makes real-time judgments on personnel activities and identifies potential anomalies in combination with environmental background, such as illegal gatherings, lingering, and the appearance of unidentified personnel. The model assesses the risk level and generates warning signals based on the current context and rule engine. If an anomaly is determined, the system automatically extracts key video clips, exports device logs, generates a structured evidence package, and pushes it to the management terminal. Administrators can quickly locate personnel trajectories and behavioral paths based on event tags, supporting compliance review and disciplinary analysis.
[0132] It should be understood that although the steps in the flowcharts in the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise expressly stated herein, there is no strict order in which these steps are performed, and they may be performed in other orders.
[0133] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A full-scene intelligent analysis system based on a multimodal large model, characterized in that, include: The feature extraction module (10) is used to acquire cross-modal data, which includes cross-modal tensors and corresponding cross-modal images. The feature extraction module (10) is also used to generate a temporal feature set based on the cross-modal tensors and a spatial structure set based on the cross-modal images. A semantic fusion module (20) is network-connected to the feature extraction module (10) to receive the temporal feature set and the spatial structure set, and performs semantic interaction on the temporal feature set and the spatial structure set based on the cross-modal structure to obtain a cross-modal data set. The causal generation module (30) is network-connected to the feature extraction module (10) to receive the temporal feature set and the spatial structure set, and to generate causal associations based on the temporal feature set and the spatial structure set; the causal generation module (30) is network-connected to the semantic fusion module (20) to receive the cross-modal data set, and to generate causal temporal associations based on the cross-modal data set and the causal associations. The dynamic prediction module (40) is connected to the causal generation module (30) to obtain the causal time series association and dynamically learns the association mechanism based on the causal time series association to generate a dynamic prediction model. The process of semantically interacting with the temporal feature set and the spatial structure set based on cross-modal structure to obtain a cross-modal data set includes the following steps: The temporal feature set and the spatial structure set are mapped to the same dimension to obtain combined dimensional features, which include temporal features and structural features, and the temporal features and structural features are in one-to-one correspondence. The temporal and structural features are interactively processed based on the temporal guidance space to obtain first interactive data, and the temporal and structural features are interactively processed based on the spatial constraint temporal to obtain second interactive data. Based on the temporal and structural features, corresponding data weights are generated, and the first and second interactive data are adjusted based on the data weights to generate a cross-modal data set.
2. The full-scene intelligent analysis system based on a multimodal large model according to claim 1, characterized in that, It also includes a data preliminary processing module (50), which performs timestamp alignment on the heterogeneous sensing data corresponding to the region under control based on timestamp alignment technology to generate mapped time axis data; the data preliminary processing module (50) is also used to filter and standardize the mapped time axis data to obtain standardized data, and construct a cross-modal model based on the standardized data to obtain cross-modal tensors.
3. The full-scene intelligent analysis system based on a multimodal large model according to claim 2, characterized in that, The data preprocessing module (50) is also used to preprocess the image data of heterogeneous sensor data to obtain key target areas, perform human image detection on the key target areas based on the detection model to obtain corresponding human image annotation information, and obtain cross-modal images based on the human image annotation information.
4. The full-scene intelligent analysis system based on a multimodal large model according to claim 1 or 3, characterized in that, It also includes an automatic tracing module (60), which is used to receive video image data of the area to be controlled in real time, and perform anomaly detection on the video image data based on the dynamic prediction module (40) to obtain detection results, generate tracing data based on the detection results, and send the detection results and the tracing data to the management terminal.
5. The full-scene intelligent analysis system based on a multimodal large model according to claim 4, characterized in that, It also includes an interactive decision module (70), which is used to receive video image data of the area to be controlled in real time, and automatically identify key information based on the video image data by the dynamic prediction module (40); and generate key target behavior trajectories and long-term trend statistical charts based on the key information.
6. A full-scene intelligent analysis method based on a multimodal large model, characterized in that, The execution of the full-scene intelligent analysis system based on a multimodal large model as described in any one of claims 1-5 includes the following steps: Acquire cross-modal data, and generate a temporal feature set and a spatial structure set based on the cross-modal data; Based on the cross-modal structure, semantic interaction is performed on the temporal feature set and the spatial structure set to obtain the cross-modal data set; Causal relationships are generated based on the temporal feature set and the spatial structure set, and causal temporal relationships are generated based on the cross-modal data set and the causal relationships. The temporal feature set includes several sets of temporal features, and the spatial structure set includes several sets of spatial features. Based on the aforementioned causal temporal correlation dynamic learning association mechanism, a dynamic prediction model is generated; The process of semantically interacting with the temporal feature set and the spatial structure set based on cross-modal structure to obtain a cross-modal data set includes the following steps: The temporal feature set and the spatial structure set are mapped to the same dimension to obtain combined dimensional features, which include temporal features and structural features, and the temporal features and structural features are in one-to-one correspondence. The temporal and structural features are interactively processed based on the temporal guidance space to obtain first interactive data, and the temporal and structural features are interactively processed based on the spatial constraint temporal to obtain second interactive data. Based on the temporal and structural features, corresponding data weights are generated, and the first and second interactive data are adjusted based on the data weights to generate a cross-modal data set.
7. The full-scene intelligent analysis method based on a multimodal large model according to claim 6, characterized in that, The process of generating corresponding data weights based on the temporal and structural features includes the following steps: The first indicator feature is obtained based on the time series features, and the first indicator feature is normalized to obtain the time series information. Based on the structural features, a second indicator feature is obtained, and the second indicator feature is normalized to obtain spatial information. Based on the temporal and structural features, corresponding cross-modal data are obtained, and cross-modal correlation is generated based on the cross-modal data; The data weight is calculated based on the temporal information, spatial information, and cross-modal correlation quantification attribute index, and a preset fusion strategy.
8. The full-scene intelligent analysis method based on a multimodal large model according to claim 6, characterized in that, Generating causal relationships based on the temporal feature set and the spatial structure set includes the following steps: Based on the set of temporal features and the set of spatial structures, causal elements are obtained, and the temporal features are detected based on the causal elements to obtain key events; Based on the causal factors, the spatial structure is analyzed to obtain the changes corresponding to the key events. A causal relationship set is generated based on the key events and the changes, and the causal relationship set is verified based on the verification mode to obtain the causal relationships.
9. The full-scene intelligent analysis method based on a multimodal large model according to claim 8, characterized in that, Detecting temporal features based on the aforementioned causal elements to obtain key events includes the following steps: Based on the temporal features and the spatial structure, feature spatial data on the time chain is obtained. The feature spatial data includes time points, temporal-series sensor data, and corresponding spatial data. The abrupt change time is obtained based on the time-series sensor data, and key events are obtained based on the abrupt change time and the spatial data.
Citation Information
Patent Citations
Security alarm information data interaction system and method
CN118486152A