A Simulation Method for Behavioral Analysis and Intelligent Risk Early Warning Based on Multimodal Information Fusion

By generating high-fidelity multimodal data in a virtual environment and conducting closed-loop simulation tests, the problem of insufficient adaptability of intelligent risk warning systems in complex environments is solved, achieving efficient system optimization and performance improvement, and is suitable for the research and development and testing of intelligent risk warning systems.

CN120562206BActive Publication Date: 2026-05-26SHENZHEN HAILINKE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HAILINKE INFORMATION TECH CO LTD
Filing Date
2025-07-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing intelligent risk warning systems lack the ability to adapt to environmental changes in the real world, resulting in unstable performance, high false alarm and false negative rates in complex and ever-changing environments. They are unable to meet the practical needs of high-risk law enforcement scenarios, mainly due to problems such as data scarcity, high acquisition costs, uncontrollable testing conditions, and difficulty in labeling.

Method used

By constructing a controllable virtual testing environment, high-fidelity multimodal data is generated using 3D modeling and physical simulation engines. Real-world conditions such as lighting and noise are simulated in the virtual environment. Closed-loop simulation and automated testing are performed in conjunction with behavioral analysis and intelligent risk warning systems, and algorithm parameters are optimized to improve the system's adaptability.

Benefits of technology

It enables the efficient generation of multimodal data consistent with the real world in a virtual environment, significantly shortens the R&D cycle, reduces verification costs, improves the system's recognition accuracy in harsh environments, reduces false alarms and false negatives, and provides reliable support for front-line law enforcement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562206B_ABST
    Figure CN120562206B_ABST
Patent Text Reader

Abstract

This invention discloses a simulation method for behavior analysis and intelligent risk early warning based on multimodal information fusion. First, a virtual 3D scene with precise physical properties and programmably controllable environmental parameters is constructed. Second, multimodal raw data streams are synchronously generated within the scene and processed using a physical degradation model based on set environmental parameters to generate data highly consistent with the real world. Third, the degraded data stream is fed into the intelligent system under test (SUT) in real time and automatically compared with the precise ground truth labels generated by the simulation, constructing a closed-loop simulation and performance evaluation system. This method can deeply verify the effectiveness of modules such as environmental perception, dynamic weight adjustment, and intelligent fusion decision-making within the SUT, and supports large-scale automated testing and parameter optimization, thereby systematically improving the system's environmental adaptability, recognition accuracy, and decision robustness under harsh conditions such as low light and high noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer simulation technology, and more specifically, to a multimodal information fusion-based behavioral analysis and intelligent risk warning simulation method. Background Technology

[0002] Portable security devices, exemplified by intelligent law enforcement recorders, are no longer simply data collection terminals, but have evolved into proactive decision-making support systems with edge AI. Their core task is to accurately identify and provide immediate warnings of potential risky behaviors by analyzing multimodal information streams collected on-site in real time. This provides crucial information support to frontline personnel in critical moments, effectively preventing escalation and ensuring personnel safety. However, the development and effectiveness verification of such advanced intelligent systems face a series of severe challenges stemming from their fundamental methodologies. These challenges constitute the core technological bottlenecks restricting their practical and widespread application.

[0003] Currently, the development and testing of systems in this field heavily rely on processing datasets collected from real-world scenarios. While this method theoretically offers the highest fidelity, it has revealed several insurmountable obstacles in practice.

[0004] The primary obstacle lies in the extreme scarcity and enormous cost of acquiring critical, high-value data. The performance of an early warning system ultimately hinges on its ability to respond to low-probability, high-risk events, such as sudden violent attacks, armed threats, or sudden collapses. These "long-tail events" occur very infrequently in daily operations, resulting in a sparse pool of truly effective samples from massive amounts of routine data, suitable for training and testing models' ability to handle extreme situations. Even when some data is acquired through specific drills, the process is extremely resource-intensive and often accompanied by strict privacy and data confidentiality requirements, making it difficult to widely and freely use this valuable data in the research and development phase. This severe imbalance in data distribution directly leads to AI models "overfitting" to routine behaviors while exhibiting a serious "lack of generalization ability" for critical, life-or-death anomalies.

[0005] Secondly, the inherent randomness of real-world testing environments leads to the uncontrollability and unreproducibility of testing conditions. The robustness of an early warning system hinges on its ability to maintain stable performance under various non-ideal, even extreme, physical environments. These environmental factors include, but are not limited to: dramatic changes in lighting from midday sun to midnight alleyways, differences in noise levels from quiet neighborhoods to noisy factories, and severe weather conditions such as rain, snow, fog, and haze. In the real world, artificially and precisely reproducing a specific combination of environments, such as simultaneously experiencing "high-decibel vehicle noise" and "multi-person physical conflict" in a "tunnel with intense glare," is virtually impossible in terms of logistical organization and security. This uncontrollability prevents us from conducting systematic stress tests and boundary condition analyses on the system, and makes it impossible to conduct fair performance comparisons under identical conditions after iterative updates to the algorithm model. This makes the algorithm optimization process highly accidental, lacking scientific rigor and direction.

[0006] Furthermore, there is the challenge of acquiring accurate ground-level labels. To effectively supervise AI model training and accurately evaluate its performance, precise and complete labels are needed for the collected multimodal data. This work is labor-intensive and highly susceptible to subjective bias. For example, accurately labeling the start and end frames of a fight in a video, the 3D pose of each participant, and the precise timing and source category of each impact sound in the audio is incredibly labor-intensive. Especially in the context of multimodal fusion, simultaneously labeling the complex causal and temporal relationships between visual and auditory events increases exponentially in difficulty. Missing, inaccurate, and asynchronous labels directly contaminate the model training process and significantly reduce the credibility of the final performance evaluation results.

[0007] Finally, all of the above problems collectively lead to lengthy R&D cycles and high iteration costs. The traditional R&D loop of "data collection - manual annotation - model training - field testing" is cumbersome, with each step interconnected and time-consuming. Any minor adjustment made by algorithm engineers to the model architecture or parameters may require repeating the entire lengthy verification process to confirm its effect, which greatly inhibits the speed of innovation, making the efficiency of technology iteration far from meeting the rapidly changing demands of real-world applications.

[0008] To overcome these bottlenecks, the industry has begun exploring the use of computer graphics and game engines to generate synthetic data. However, most synthetic data generation methods are still in their infancy. They often focus on generating single-modal data, such as generating only visual images, while neglecting the high-fidelity audio synchronized with them. They fail to deeply simulate the physical laws of the real world, especially the complex degradation process of sensor data quality caused by environmental factors. A typical example is that existing simulators may generate smooth fighting animations, but the rendered images are too "clean" and cannot simulate the specific patterns of shot noise and readout noise that CMOS sensors inevitably produce under low-light conditions; they may also fail to generate audio streams that accurately match the distance of characters, lip movements, and environmental spatial structures (such as reverberation) in the image, let alone physically mix these audio streams with simulated background noise. Using synthetic data with such a huge "domain difference" from the real physical world to train AI models often results in a precipitous drop in system performance once deployed on physical devices because the system cannot adapt to the noise characteristics of real sensors and environmental interference.

[0009] In summary, the core technological bottleneck of existing behavior analysis and early warning systems lies in their lack of dynamic perception and adaptability to the environment, whether they are single-modal or multimodal systems employing static fusion strategies. They cannot dynamically and intelligently adjust the contribution of different information modalities in fusion decision-making based on changes in actual environmental conditions such as lighting and noise. This results in unstable performance, poor reliability, and high false alarm and false negative rates in complex and ever-changing real-world environments, making it difficult to meet the practical needs of high-risk, highly dynamic law enforcement scenarios. The root cause of this bottleneck lies in the limitations of the research and development methodology, namely, the lack of a new paradigm that can systematically solve the aforementioned data, testing, and efficiency problems. Therefore, the industry urgently needs a new type of system that can adapt to environmental changes and intelligently optimize fusion strategies. To achieve this goal, a simulation method is first required that can generate, on a large scale and at low cost, multimodal data with precise labels that are highly consistent with real-world physical laws in a controllable and quantifiable virtual environment. This data will then be used for efficient training, rigorous testing, and intelligent optimization of the core algorithms of such new systems. Summary of the Invention

[0010] To address the aforementioned technical problems, this invention provides a simulation method for behavioral analysis and intelligent risk early warning based on multimodal information fusion. This method can not only reproduce the geometry and behavior of a scene but also simulate the impact of environmental factors on the data stream of multimodal sensors. Through this method, a controllable virtual test field can be created for the intelligent system under test (SUT). The simulation method of this invention is logically decomposed into the following tightly coupled, sequentially progressive technical modules.

[0011] The Virtual Environment and Dynamic Scene Construction Module provides an interactive, physically accurate 3D digital world as a stage for simulation. Its implementation begins with using professional 3D modeling software (such as Blender and Autodesk 3ds Max) or reality capture technologies like LiDAR scanning and photogrammetry to construct detailed geometric models of the simulation scene, encompassing everything from macroscopic urban blocks and building interiors to microscopic ground textures and vegetation morphology. These models are then imported into an engine with advanced physical simulation capabilities, such as the widely used Unreal Engine 5 or Unity. Within the engine, a crucial step is assigning all surfaces in the scene their corresponding material properties in the physical world. This is not merely visual mapping, but deep physical parameterization; for example, assigning a specific bidirectional reflectance distribution function (BRDF) and transmittance to glass windows, and corresponding acoustic impedance and sound absorption coefficient to concrete walls. The core of this module lies in its procedural and fine-grained control over key environmental parameters. Operators can arbitrarily set and dynamically change global and local environmental parameters affecting the quality of multimodal data through scripts or a graphical interface. The scope of specific controllable parameters includes, but is not limited to, lighting parameters such as global illumination intensity (in lux), light source type (point light, spotlight, parallel light to simulate the sun), light source position, color temperature, shadow softness, and the density and attenuation characteristics of atmospheric scattering media such as smoke, fog, and haze simulated through volumetric fog and particle systems. It also includes acoustic parameters, such as selectable and configurable global background noise (e.g., urban traffic noise, factory machinery noise, crowd noise, etc., in decibels, dB), and the definition of sound reverberation and frequency-related absorption characteristics for each individual material surface in the scene. Finally, this module uses a complex AI behavior system, such as a behavior tree or hierarchical state machine, to drive virtual agents in the scene to execute pre-defined, highly complex behavioral scripts. These scripts can simulate a range of behaviors highly relevant to security and early warning, from routine patrols and conversations to sudden running chases, heated arguments, various physical conflicts (such as shoving and punching), sudden falls in emergencies, and even wielding specific dangerous objects (such as knives and sticks). Each behavior pattern is associated with a precise motion capture animation library and a synchronized sound event library. These precisely script-controlled sequences of behaviors constitute the absolutely objective and complete "Ground Truth" benchmark for the entire simulation experiment.

[0012] The multimodal data synchronization generation and degradation module places a "virtual sensor suite" simulating real portable law enforcement equipment in a virtual scene and drives it to generate a multimodal data stream with physical characteristics highly consistent with the output of real sensors. This module ensures perfect, sample-point-level synchronization of all generated data modalities in time by sharing a unified clock with the simulation engine. The data generation process begins with the creation of high-fidelity visual data. In the virtual scene, one or more "virtual cameras" configured with precise physical camera parameters (such as focal length, aperture, and sensor size) are attached to the agent acting as the observer (e.g., simulating a recorder worn on the chest of a law enforcement officer). Advanced techniques in modern graphics rendering pipelines, such as real-time ray tracing or path tracing, are used to generate raw, lossless video frame sequences with physically correct lighting, shadows, and reflections. Simultaneously, high-fidelity audio data is generated. The system uses the 3D positions, sound intensities, and radiation patterns of all sound sources in the scene (e.g., the Agent's mouth, footsteps, object collision points) and the scene's acoustic materials and geometry to precisely calculate the propagation paths, reflections, diffractions, and absorption of sound in complex environments through complex acoustic rendering algorithms (such as ray tracing based on geometry acoustics or the finite element method based on the wave equation). This generates a high-fidelity, multi-channel raw audio stream with an analog microphone array. Simultaneously, the generation of high-fidelity inertial measurement unit (IMU) data is relatively straightforward. It generates an IMU data stream that is completely synchronized with the motion posture of the image by querying the linear acceleration and angular velocity in the world coordinate system in real time at high frequencies (e.g., 200Hz) from the physical rigid body components of the Agent carrying the virtual camera. The most critical step in this module is the environment-driven data degradation process. This step aims to perform a procedural and physically plausible "contamination" and "degradation" process on the previously generated, overly "pure" ideal data stream, based on the environmental parameters set in the first module. This simulates the inherent limitations and environmental interference faced by real sensors in the face of the imperfections of the physical world. For visual data, the degradation model is represented as follows: In this formula, It is a moment The original rendered image, Simulates the barrel or pincushion distortion, chromatic aberration, and vignetting effects unavoidable in real-world lenses. The core... The function is based on the current ambient light intensity. and simulated sensor photosensitivity The values ​​are applied to the image to produce shot noise and read noise that conform to the physical properties of semiconductors. The function depends on the camera's exposure time. Internal velocity and angular velocity Apply the appropriate motion blur effect. For audio data, the degradation model is as follows: .in, It is the original multi-channel audio. Represents convolution operation. It is based on the location of the sound source Receiver location and scene geometry at time The Room Impulse Response is calculated or queried in real time; this convolution operation simulates the reverberation and echo effects of sound. It generates a background noise signal based on the set noise type and decibel value, and then precisely mixes it with the foreground sound in terms of energy.

[0013] Closed-loop simulation and performance evaluation module. This module interfaces the simulation method with the behavioral analysis and intelligent risk warning system (SUT) under test to construct an automated, end-to-end evaluation closed loop. Its workflow begins with data injection, where the degraded multimodal data stream (…) (etc.) are fed as input signals to a running SUT instance in real time via low-latency network protocols (such as gRPC) or shared memory. Following this is result capture; the system captures all output information of the SUT in real time and non-intrusively. This includes not only its final decision results (such as the identified behavior category, assessed risk level, and confidence level), but also intermediate state data from its internal key modules, such as the illumination and noise levels assessed by its environmental perception module, and the real-time fusion weights calculated for each modality by its dynamic weight adjustment module. Finally, and most importantly, is the automatic performance comparison and measurement. A separate evaluator program rigorously compares the captured SUT output with the absolutely accurate ground truth labels preset for the simulation scene in the first module, frame by frame and event by event. Based on the comparison results, the system automatically calculates and records a series of detailed key performance indicators (KPIs). These KPIs extend beyond macro-level behavior recognition confusion matrices, precision, recall, and F1 scores for each behavior class, as well as overall false alarm rate (FAR) and missed detection rate (MDR). They also include in-depth verification of the SUT's internal logic, such as calculating the dynamic weights of the SUT output. The Pearson correlation coefficient between the data quality degradation parameter set in the simulation environment and the data quality degradation parameter is used to quantitatively evaluate the effectiveness of its environmental adaptability and response sensitivity.

[0014] The automated testing and parameter optimization module fully leverages the programmability and determinism of the simulation environment, elevating SUT testing from traditional, fragmented, individual tests to a systematic, large-scale automated testing approach. A typical application is parameter scanning testing, which uses scripts to automatically scan changes in one or more environmental parameters throughout their effective range. For example, a script depicting a "two-person armed standoff" can be fixed and run in a loop. During each run, the ambient light level can be gradually reduced from 10,000 lux (daytime) to 0.1 lux (nighttime) on a logarithmic scale, while the background noise level can be gradually increased from 30 dB (quiet) to 100 dB (extremely noisy). After each simulation, the complete set of performance metrics for the SUT under the current parameter combination is recorded. After testing, all data points can be visualized as a "performance map" of the SUT in a multi-dimensional environmental parameter space, clearly identifying the most vulnerable areas of the system. Furthermore, this module can automatically optimize model parameters. Key adjustable intrinsic parameters in the SUT, such as the prior reliability coefficient vector in its dynamic weight calculation formula, can be used. Or, the hyperparameters of the reward function in its reinforcement learning module used for adaptive learning. The parameters are set as variables to be optimized. Then, a composite index that comprehensively reflects the overall system performance (e.g., a weighted function combining F1 score, FAR, and MDR) is defined as the optimization objective function. Finally, an optimization algorithm (such as Bayesian optimization or evolutionary algorithm) is used to drive the entire simulation optimization loop. The optimizer proposes a set of candidate parameters, the system automatically runs a standardized benchmark simulation and calculates the objective function value, feeds the results back to the optimizer, the optimizer updates its surrogate model of the parameter space based on all historical data, and intelligently selects the next parameter point most likely to bring performance improvement for testing. By running thousands of such automated simulation iterations in parallel on a computing cluster, this method can automatically and data-drivenly find the parameter configuration that makes the overall performance of the SUT reach the global optimum or near-optimal.

[0015] The simulation method proposed in this invention effectively solves the technical bottleneck in the security AI field caused by the scarcity of key data and the high cost of manual annotation by constructing a controllable and repeatable virtual testing environment. This method can generate high-risk scene multimodal data with precise, synchronized real-world ground labels on demand and programmatically, and achieves precise control and scientific reproduction of environmental conditions such as lighting and noise. This provides a reliable data benchmark for systematic stress testing, boundary condition analysis, and algorithm iteration, thus constructing a rapid iterative closed loop of "virtual design - digital twin simulation - data-driven optimization," significantly shortening the R&D cycle and reducing verification costs. This method not only evaluates the final input-output performance of the system but also deeply verifies and optimizes a class of intelligent systems with advanced internal mechanisms. These systems integrate an environmental perception module to quantify environmental parameters in real time, and their dynamic weight adjustment module calculates real-time weights for each modality based on environmental conditions. The intelligent fusion decision module then uses these weights for weighted deep fusion, while also supporting data-driven optimization of their adaptive learning capabilities. Ultimately, the system, after thorough testing and optimization using this method, can truly adapt to changes in the environment, thereby significantly improving the recognition accuracy under harsh real-world conditions such as low light and high noise, greatly reducing false alarms and false negatives, and providing reliable technical support and security for frontline law enforcement personnel. Attached Figure Description

[0016] Figure 1 A flowchart of the method provided by the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, a detailed operational description of specific embodiments of the invention will be provided below. The specific embodiments described herein are intended to illustrate the methodology of this invention through a complete example, and are not intended to constitute any limitation on the scope of protection of this invention.

[0018] from Figure 1 As can be seen, the general flow of the method provided by this invention is as follows:

[0019] First, the virtual environment construction module performs 3D scene modeling, physical material configuration, and AI behavior script settings to build a controllable simulation scene. Once the scene is ready, the multimodal data generation module generates raw multimodal data streams through visual rendering, acoustic simulation, and IMU data acquisition.

[0020] Next, the environmental degradation processing module applies physically plausible degradation processing to the raw data based on the set environmental parameters, including noise simulation and illumination attenuation of the visual data. And the reverb effect of audio data superimposed. .

[0021] The degraded data is injected into the system under test (SUT) in real time via the gRPC protocol for processing. The performance evaluation module compares the output of the SUT with the preset Ground Truth labels and calculates key performance indicators such as F1 score, false alarm rate (FAR), and false negative rate (MDR).

[0022] Finally, based on the performance evaluation results, the parameter optimization module uses environmental parameter scanning and Bayesian optimization algorithms to find the optimal parameter configuration and feeds the optimized parameters back to the virtual environment construction module, forming a complete closed-loop optimization system. The entire process enables automated testing, evaluation, and optimization of the multimodal information fusion system.

[0023] In a preferred embodiment, the simulation method is implemented using a highly integrated and modular software platform. This platform can be flexibly deployed on local high-performance graphics workstations or on cloud computing clusters for large-scale parallel computing. Its core technology stack and architecture design are as follows: For the core of the simulation engine, we chose to use a leading commercial game engine, specifically Unreal Engine 5. This engine was chosen because it provides a powerful set of features for building realistic virtual worlds, including but not limited to its revolutionary Lumen dynamic global illumination system and Nanite virtualized micropolygon geometry technology, both of which ensure high realism and performance in the rendered images. Simultaneously, its built-in Chaos physics simulation system can accurately simulate rigid body dynamics and collisions, while the comprehensive Blueprint visualization scripting system and C++ programming interface provide a solid foundation for implementing complex simulation logic and deep control over the engine's underlying layers. To connect the simulated world with the AI ​​system under test, we designed a deep learning framework interface. This interface tightly couples the simulation engine to an independent Python process through efficient inter-process communication (IPC) mechanisms, such as gRPC or ZeroMQ. This Python process runs a System Under Test (SUT) based on a PyTorch or TensorFlow deep learning framework. This decoupled architecture allows various types of data generated by the simulation engine to be serialized and seamlessly, with low latency, flow to the input of the SUT. Meanwhile, the inference results from the SUT can be asynchronously fed back to the simulation engine or evaluation module. To support the needs of large-scale automated testing, the platform's backend employs a modern distributed computing and task scheduling system. For example, Kubernetes is used to containerize and schedule resources for both the simulation engine and SUT instances, and a distributed task queue framework like Celery is used to manage tens of thousands of parallel simulation tasks. Users can submit a test plan containing thousands of parameter combinations through a web front-end interface, and the backend system automatically distributes these tasks to idle nodes in the cluster for execution. Finally, to facilitate user interaction with the simulation platform, we provide a data management and visualization front-end based on web technologies (such as React or Vue.js frameworks). Users can intuitively configure simulation scenarios, design and submit test cases, monitor the execution status of simulation tasks in real time through this interface, and perform interactive visualization analysis on the massive amount of returned results data after the test is completed. For example, they can generate multi-dimensional performance heatmaps, plot ROC curves, and generate detailed comparative analysis reports with one click.

[0024] Taking a specific test scenario as an example, the specific implementation of each functional module is explained in detail according to the internal logical order of the method of this invention. The test scenario is set as follows: "In a dimly lit underground passage with obvious echoes, a law enforcement officer (whose perspective is the input source of SUT) encounters a suspicious person, and the two sides quickly escalate from verbal dispute to physical conflict."

[0025] The first step is the implementation of the virtual environment and dynamic scene construction module. We first created a 3D model of the underground passage in Blender, including its arched structure, mottled walls, slippery water stains on the ground, and scattered trash. After the model was completed, it was imported into Unreal Engine. Next, we assigned precise PBR (Physically Based Rendering) and acoustic materials to each surface in the scene. For example, we specified a high sound reflection coefficient and a low absorption coefficient for the concrete walls to produce noticeable echoes in subsequent acoustic simulations. Then, we created a blueprint class called "BP_EnvironmentController," which encapsulates the interface for controlling environmental parameters. This blueprint exposes several functions, such as UpdateLighting(float TargetLux, float TransitionTime) and SetAmbientNoise(SoundClass NoiseType, float TargetDB). In the scene, we placed two virtual agents driven by AI behavior trees, one named "Officer" and the other "Suspect." The behavior tree of "Suspect" is designed with the following logic: the initial state is "Wandering". When "Officer" enters its perception range, the triggered state changes to "Confrontation_Verbal". In this state, if "Officer" continues to approach, the triggered state changes to "Attack_Push", and finally enters the "Struggle" state. Each state is bound to a corresponding motion capture animation and synchronized sound events. For example, when the "Attack_Push" animation plays to frame 50, we record the ground real-world label: {"timestamp": 1.666, "event_type": "behavior_start", "behavior_class": "physical_assault", "participants": [0, 1]}. Simultaneously, when a pushing action occurs, an impact sound is triggered and a label is recorded: {"timestamp": 1.666, "event_type": "audio_event", "audio_class": "impact_sound", "source_location": [x, y, z]}. All these labels are written in real time to a Ground Truth log file associated with this simulation.

[0026] Secondly, the implementation of the multimodal data synchronization generation and degradation module is addressed. We attach an "AC_SimulatedRecorder" component to the virtual chest position of the "Officer" agent. This component internally contains a virtual camera, microphone array, and IMU. At each tick (time step) of the simulation engine, this component performs the following operations: First, it obtains the current precise simulation timestamp t. Then, it calls its internal CineCamera component to render the raw HDR image I_raw(t) from the current viewpoint using the Lumen system. Simultaneously, it calls an acoustic simulation plugin (such as Unreal Engine's built-in AudioGameplay Volumes or a third-party plugin) to calculate the 4-channel raw audio waveform A_raw(t) when the sound reaches the microphone array position on the "Officer's" chest, based on the position and content of the "Suspect's" mouth and other sounds generated by the action. At this point, the audio already includes early reflections and reverberation caused by the geometry of the underground passage. Next, it queries the "Officer's" current acceleration and angular velocity data IMU_raw(t) from the physics simulation component. The next step is the crucial data degradation. This component queries "BP_EnvironmentController" to obtain the current ambient light value L(t) (e.g., 50 lux at the start of the scene, which can be reduced to 5 lux by a script). Based on a pre-calibrated function noise_level = f(L(t), ISO), it calculates the noise intensity to be applied and adds the corresponding Poisson-Gaussian mixture noise to I_raw(t). Simultaneously, based on the displacement and rotation of the "Officer" agent within this frame, it applies motion blur to the image, ultimately obtaining I_final(t). For audio, it obtains background noise from the environment (e.g., the sound of distant water droplets and the low-frequency hum of ventilation ducts) and mixes it with A_raw(t) to obtain the final A_final(t). Finally, it streams I_final(t) (encoded as JPEG), A_final(t) (encoded as FLAC), and IMU_raw(t) along with a timestamp t via gRPC to the SUT waiting in another process.

[0027] Next is the implementation of the closed-loop simulation and performance evaluation module. The SUT under test is encapsulated as a Python service that listens on a gRPC port. Upon receiving a data packet from the simulator, it initiates its complete internal processing flow. The SUT's output, a JSON object containing behavior judgments, risk levels, and internal weights, is sent back to a separate "evaluator" service. The evaluator receives the SUT's output, for example, {"timestamp": 1.700, "behavior_class": "fighting", "risk_level": 2, "weights": {"video": 0.35, "audio": 0.55, "imu": 0.10}}. It immediately searches the Ground Truth log for the true label near the timestamp 1.700 (allowing a small tolerance window). It finds the true behavior to be "physical_assault," semantically matching the "fighting" identified by the SUT, and thus records a true positive (TP). It also records the weights calculated by the SUT for subsequent analysis. If the SUT does not output any alarms during the time period of the real event, the evaluator will record a false negative (FN). After the entire simulation scenario is completed, the evaluator summarizes all TP, FP, TN, and FN counts and calculates the final performance metrics.

[0028] Finally, we implement a large-scale automated testing and parameter optimization module. Based on the closed loop described above, we can design an automated testing script. This script uses a nested loop: the outer loop iterates through the illumination intensity `lux_levels = [100, 50, 20, 10, 5, 1]`, and the inner loop iterates through the reverberation intensity (by changing the acoustic absorption coefficient of the walls) `reverb_coeffs = [0.1, 0.3, 0.5, 0.7, 0.9]`. For each (lux, reverb) combination, the script automatically runs a complete "underground passage conflict" simulation and records the SUT's F1 score. After the test, a heatmap can be generated to visually demonstrate the SUT's performance sensitivity to illumination and reverberation. A more advanced application is parameter optimization. Suppose we want to optimize the three prior reliability coefficients `r_video`, `r_audio`, and `r_imu` in the dynamic weight formula of the SUT. We use the optuna library to define an objective function `objective(trial)`. Inside this function, the `trial` object will suggest a set of r coefficient values ​​to be tested. We configure these values ​​for the SUT and then run a standard benchmark test set containing various scenarios (underground passages, noisy streets, open squares, etc.). Finally, the overall performance score on the test set (e.g., F1_score -0.5FAR -2.0MDR) is used as the return value. The optuna library intelligently guides subsequent experiments based on historical test results. After hundreds or even thousands of fully automated "simulation-evaluation" cycles, it ultimately finds a set of r coefficients that enable the SUT to perform optimally under various complex environments.

[0029] Table 1 simulates a robustness test on a SUT, with the fixed scenario being "underground passage conflict," but the two environmental variables of lighting and acoustic reverberation are systematically changed, and the F1 score for recognizing the key behavior of "physical fighting" is recorded.

[0030] Table 1:

[0031] ;

[0032] The simulation data clearly shows that the SUT can maintain acceptable performance under a single environmental challenge, but its performance drops precipitously under extreme conditions where both lighting and acoustic conditions are severe. This finding directly exposes the vulnerability of its multimodal fusion strategy to multiple layers of information contamination, providing a clear target for subsequent algorithm optimization.

[0033] Table 2 shows how to use this method to verify and optimize the dynamic weight adjustment mechanism within the SUT. Through Bayesian optimization, we found two different sets of prior reliability coefficients r and compared their performance in harsh scenarios of "low illumination and high reverberation".

[0034] Table 2:

[0035] ;

[0036] The simulation results show that the prior weight configuration, which is more biased towards trusted audio modalities and was found through large-scale simulation optimization, enables the SUT to allocate its attention more intelligently in harsh environments, thereby significantly improving its final recognition performance.

[0037] Finally, it should be noted that the above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multimodal information fusion-based behavioral analysis and intelligent risk early warning simulation method, characterized in that, This method is specifically designed for the systematic testing, verification, and optimization of an intelligent system with environmental adaptability. The simulation method includes the following steps: Construct virtual environments and dynamic scenes with programmable controllable environmental parameters; In this virtual environment, multimodal raw data streams are generated synchronously; Based on the environmental parameters, the multimodal raw data stream is subjected to programmed degradation processing to simulate the inherent limitations and environmental interference of real sensors when facing the imperfections of the physical world. The degraded multimodal data stream is fed to the intelligent system under test in real time, and its output, including its internal state, is captured. The output of the intelligent system is automatically compared with the simulated ground real labels to complete the evaluation and optimization of the intelligent system's performance.

2. The method according to claim 1, characterized in that, The intelligent system has the following internal structure: The environmental perception module is used to quantitatively assess the current ambient light intensity and background noise level in real time based on the input multimodal data stream. The dynamic weight adjustment module calculates dynamic fusion weights in real time for different information modalities, including video, audio, and IMU, based on the evaluation results of the environmental perception module. The intelligent fusion decision-making module utilizes the dynamic weights to perform weighted deep fusion of features extracted from each modality, thereby achieving accurate analysis and risk level assessment of complex behaviors.

3. The method according to claim 2, characterized in that, The performance evaluation steps of the simulation method thoroughly verify the internal mechanisms of the intelligent system in the following ways: The system captures and records the illumination and noise parameters evaluated by the environmental perception module of the intelligent system in real time, and compares them with the real physical parameters set in the simulation environment. Simultaneously, the real-time weights calculated by the dynamic weight adjustment module for each mode are captured and recorded to quantitatively analyze the response relationship and rationality between the weight allocation strategy and environmental changes.

4. The method according to claim 2, characterized in that, The simulation method further includes a step of optimizing the adaptive learning capability of the intelligent system. This step: The prior reliability coefficient in the dynamic weight adjustment module, or the reward function hyperparameter in the reinforcement learning module used for adaptive learning, is taken as the variable to be optimized. Through parallel automated simulation optimization loops, a set of parameter configurations that enable the intelligent system to achieve optimal overall performance under harsh conditions such as low light and high noise are found in a data-driven manner, thereby continuously optimizing its model.

5. The method according to claim 1, characterized in that, The steps for constructing the virtual environment and dynamic scene include: assigning all surfaces in the scene their corresponding material properties in the physical world, including a bidirectional reflection distribution function for lighting simulation and acoustic impedance and sound absorption coefficient for acoustic simulation, to ensure the realism of the physical simulation.

6. The method according to claim 1, characterized in that, The method also includes a large-scale automated testing step, which automatically scans the changes of one or more environmental parameters throughout their effective range using scripts to generate a "performance map" of the intelligent system in a multi-dimensional environmental parameter space, and locates and analyzes the vulnerable areas of its performance under harsh conditions such as low light and high noise.

7. The method according to claim 1, characterized in that, The ground reality tag is generated by the AI ​​behavior system driving the virtual character to execute a preset behavior script. The tag accurately records the category, start and end time of all behavioral events, as well as the acoustic events that are synchronized with them.

Citation Information

Patent Citations

  • Smart text, travel and digital twin interaction system

    CN118822792A

  • Environment self-adaptive meta-universe scene sensing system and method

    CN119672209A