Dynamic environment and multi-event collaborative perception method based on multi-modal neural network
Patent Information
- Application Number
- CN202511682231.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-11-17
AI Technical Summary
这类方法虽然在静态或单一场景下具有一定效果,但在多源信息并存、环境因素复杂的情况下,其识别结果往往受到天气、光照、遮挡和噪声等条件的显著影响,难以实现不同类型数据的深度融合与全局感知
(1)通过对多源传感装置获取的可见光、红外、激光测距及雷达数据进行时间同步与空间配准,并在神经网络中引入置信度变化率修正机制,使系统能够根据环境变化动态调整各模态的特征权重,从而在雨雾、夜间、强光及振动干扰等复杂条件下保持高精度识别能力。(2)本发明通过计算相邻时间片间的模态相似度与置信度信息差异,建立跨时间片的特征关联关系,实现了多模态信息在时间维度上的连续融合,能够有效消除数据冗余与冲突,提升目标和事件识别的时效性与连贯性。(3)本方法不仅能够同时识别外部目标和内部安全事件,还可基于置信度差异和时序一致性执行加权融合与冲突消解,自动建立事件优先级关系,实现多事件间的关联判断与综合决策,大幅提高系统的综合响应效率。
Smart Images

Figure CN121562671B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method for dynamic environment and multi-event collaborative perception based on multimodal neural networks. Background Technology
[0002] Existing dynamic environmental perception and event recognition technologies mostly rely on single sensing modes or rule-based analysis methods. In fields such as shipping, autonomous driving, and industrial monitoring, independent devices such as radar, cameras, infrared sensors, or sonar are often used for target detection and status monitoring. While these methods are effective in static or single-scene conditions, their recognition results are often significantly affected by weather, lighting, occlusion, and noise when multiple sources of information coexist and environmental factors are complex. This makes it difficult to achieve deep fusion of different types of data and global perception.
[0003] However, existing technologies have several shortcomings when facing complex and dynamic environments. First, single-modal sensing devices have limited dimensions for extracting target features, resulting in a significant decrease in recognition capabilities under conditions such as rain, fog, and nighttime. Second, multimodal data fusion algorithms generally remain at the level of static weighting or simple splicing, lacking adaptive feature correction mechanisms for continuous temporal changes, making it difficult to accurately reflect the temporal correlation between events. Furthermore, existing methods mostly focus on the identification of single targets or single events, lacking unified analysis and collaborative judgment of the external environment and internal state, and are unable to simultaneously prioritize and provide early warning responses for multiple events.
[0004] Therefore, there is an urgent need for a technical solution that can achieve unified processing of multi-source data, automatic weight adjustment, and collaborative perception of multiple events in complex and dynamic environments, so as to improve the comprehensive understanding of multimodal information and the level of real-time response. Summary of the Invention
[0005] This application provides a dynamic environment and multi-event collaborative perception method based on multimodal neural networks to improve the real-time performance and accuracy in multi-source data fusion and security monitoring.
[0006] This application provides a dynamic environment and multi-event collaborative perception method based on multimodal neural networks, including: Multimodal data of the dynamic environment is collected based on multi-source sensing devices, and the multimodal data is synchronized in time and registered in space to generate a multimodal data sequence under a unified reference coordinate according to time slices. The multimodal data sequence is input into a multimodal deep neural network, the feature information of each modality is extracted, the modal similarity and confidence change rate of adjacent time slices are calculated, and the feature weights of each modality are corrected according to the confidence change rate to obtain the corrected feature information. Using the corrected feature information, external targets and internal events are identified respectively, generating corresponding candidate results and their confidence information, thus forming a candidate event set; Based on the differences in confidence information and temporal consistency in the candidate event set, weighted fusion and conflict resolution are performed to establish event priority relationships and obtain a collaborative event set containing event labels, spatial locations, and temporal ranges; Multi-event collaborative perception results are generated based on a set of collaborative events. These results include risk levels and corresponding early warning or control instructions.
[0007] The beneficial effects of the technical solution provided in this application include: (1) By performing time synchronization and spatial registration on visible light, infrared, laser ranging and radar data acquired by multi-source sensing devices, and introducing a confidence change rate correction mechanism in the neural network, the system can dynamically adjust the feature weights of each mode according to environmental changes, thereby maintaining high-precision recognition capability under complex conditions such as rain, fog, night, strong light and vibration interference. (2) By calculating the modal similarity and confidence information difference between adjacent time slices, this invention establishes a feature association relationship across time slices, realizing the continuous fusion of multimodal information in the time dimension, which can effectively eliminate data redundancy and conflict, and improve the timeliness and coherence of target and event recognition. (3) This method can not only identify external targets and internal security events at the same time, but also perform weighted fusion and conflict resolution based on confidence difference and temporal consistency, automatically establish event priority relationship, realize the association judgment and comprehensive decision-making between multiple events, and greatly improve the overall response efficiency of the system. Attached Figure Description
[0008] Figure 1 This is a flowchart of a dynamic environment and multi-event collaborative perception method based on a multimodal neural network provided in the first embodiment of this application. Detailed Implementation
[0009] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0010] The first embodiment of this application provides a method for dynamic environment and multi-event collaborative perception based on multimodal neural networks. Please refer to... Figure 1 This figure is a flowchart of the first embodiment of this application. The following is in conjunction with... Figure 1 The first embodiment of this application provides a detailed description of a dynamic environment and multi-event collaborative perception method based on a multimodal neural network.
[0011] Step S101: Collect multimodal data of the dynamic environment based on multi-source sensing devices, perform time synchronization and spatial registration of the multimodal data, and generate a multimodal data sequence under a unified reference coordinate according to time slices.
[0012] Multimodal data refers to data with different physical characteristics collected by various types of sensing devices in the same dynamic environment, including visible light images, infrared thermal images, laser ranging point clouds, radar direction finding information, and navigation position data. Because these data come from different sensors, their sampling frequencies, response delays, installation locations, and measurement angles vary. Therefore, time synchronization and spatial registration are necessary to ensure they correspond to a consistent physical scene under the same time reference and spatial coordinates. A time slice is a fixed-length time interval used to organize the data sets of each modality into a time-aligned set; the unified reference coordinates are a shared spatial benchmark for the entire system, used to describe the actual positional relationships corresponding to all modal data.
[0013] During implementation, time synchronization should be completed first. A high-precision clock from the system's main control unit should be selected as the unified time reference. The output clock time should be recorded, and the raw timestamps collected by each sensor should be compared with this reference. If the sampling time of a sensor lags behind the main clock, for example, by approximately 20 milliseconds, the system should record this offset and compensate for it before data input. For sensors with different sampling frequencies, interpolation or alignment strategies can be used: for example, when the update frequency of infrared images is lower than that of visible light images, the estimated image for the current time slice can be calculated by analyzing the data trends of adjacent moments between two frames of infrared data, thus maintaining the temporal correspondence of each mode. If a large time difference occurs, a delay buffer can be set to temporarily store high-frequency mode data before network input, until all mode data are aligned to the same time slice before input, ensuring precise temporal consistency.
[0014] Next, spatial registration is performed. Different sensors are installed at different positions and angles, with different coordinate origins and orientations. Therefore, it is necessary to establish a correspondence from their respective coordinate systems to a unified reference coordinate system. During implementation, several fixed spatial marker points should be selected during the system installation phase, such as three non-collinear structural feature points on the ship's deck. The positions of these marker points in each sensor's coordinate system are determined. By comparing the coordinates of the marker points observed by each sensor with the positions of the marker points in the predefined unified reference coordinate system, the attitude and position parameters of each sensor relative to the reference coordinates, i.e., spatial rotation and displacement, are obtained. In other words, the orientation and offset distance of each sensor relative to the unified coordinate system are determined. After registration is completed, any point or image pixel acquired by the sensors can be accurately mapped to the unified coordinate system according to this relationship, thus ensuring that the same target described by different modalities coincides in space.
[0015] After time synchronization and spatial registration are completed, data should be organized according to time slices. Each time slice represents the state at a certain moment or within a short period and should contain aligned data for all modalities within that time period. The length of the time slice should be selected according to the application scenario: in highly dynamic scenarios (such as ship collision avoidance monitoring), a time slice of about 100 milliseconds can be selected; in stable monitoring scenarios, it can be appropriately extended to 500 milliseconds or 1 second to reduce the amount of data without sacrificing timeliness. Each time slice should contain a multimodal data set under a unified reference coordinate system, including image frames, heatmap frames, point cloud data, radar direction finding, velocity vectors, and navigation position information, along with corresponding time labels.
[0016] Multimodal data have different numerical ranges and physical dimensions. For example, infrared thermal images use temperature grayscale values, visible light images use light intensity values, laser point clouds use distance values, radar information uses velocity and direction, and navigation data uses latitude, longitude, and heading angle. To facilitate subsequent neural network processing, these data need to be normalized to make different modalities numerically comparable. Normalization can be understood as mapping the value range of each data type to a unified interval, such as between 0 and 1. Specifically, the observed values are scaled proportionally using the upper and lower physical limits of each modality as a reference. For example, if the infrared temperature range is 0℃ to 100℃, then 50℃ is mapped to 0.5; if the maximum effective distance of laser ranging is 200 meters, then 100 meters is mapped to 0.5. In this way, data from all modalities can be received and processed by the neural network at a unified scale.
[0017] In complex environments, data from certain sensors may be noisy or incomplete. Visible light is prone to loss of clarity at night or in rain and fog, and laser ranging may produce holes under conditions of strong reflection or waves. Therefore, the system should attach a quality marker to each mode in each time slice to indicate the completeness and signal-to-noise ratio (SNR) of that mode's data. Modes with low SNR or severe data loss can have their influence weight reduced in subsequent processing to prevent them from misleading the overall identification. For locally missing data, it can be repaired through time-series interpolation or cross-modal compensation. For example, when a local area of a visible light image loses brightness, the temperature contrast of the corresponding area in an infrared image can be used to supplement the details; when there are gaps in the laser point cloud, point cloud information of that area in the preceding and following time slices can be referenced for interpolation to maintain spatial continuity. The compensation and repair process should adhere to the principle of not introducing false information and should be limited to reasonable filling based on reliable signals provided by adjacent or related modes.
[0018] Finally, the data, after time synchronization, spatial registration, normalization, and quality labeling, are organized into a multimodal data sequence in chronological order. Each time slice corresponds to a complete data packet, containing registered data for all modalities and additional information. The data packets should maintain temporal continuity to form a stable sequence input, enabling the subsequent neural network to identify time-varying patterns, such as target approach, heat source enhancement, or the gradual emergence of hazardous events. In a shipboard environment, the length of this sequence can be dynamically adjusted according to mission requirements, suitable for both short-term real-time early warning and long-term monitoring and trend analysis.
[0019] Furthermore, the process of acquiring multimodal data of the dynamic environment based on multi-source sensing devices, performing time synchronization and spatial registration on the multimodal data, and generating a multimodal data sequence under a unified reference coordinate system according to time slices includes: Environmental perception data is collected by multi-source sensing devices, and the sampling period of each sensing device is adjusted according to the environmental complexity parameter. The environmental complexity parameter is determined by a combination of the rate of change of light intensity, the amplitude of change of radar echo density, and the degree of gas concentration disturbance. When the environmental complexity parameter exceeds a preset threshold, the sampling frequency of each sensing device is increased to ensure the time continuity and data integrity in dynamic environments. The original data output by each sensor device is timestamped and compared. Using the unified reference clock of the main control system as a reference, the time delay of each sensor device is calculated and a time compensation curve is generated. The subsequent sampling triggering sequence of each sensor device is corrected according to the time compensation curve, so that the data of different modes correspond to the same sampling time under the unified time reference, thereby realizing the time synchronization of multimodal data. Spatial coordinate alignment is performed on the multimodal data that has been synchronized in time. Based on the difference in the response positions of multiple fixed reference points in the fields of view of different sensors, the spatial orientation and installation offset of each sensor are gradually corrected until the spatial position deviation of the same reference point in the unified coordinate system is lower than the set error range under the observation of all sensors, thereby ensuring the spatial consistency of multimodal data under the unified reference coordinate system. Based on the data that has been synchronized in time and registered in space, the multimodal data is organized into a continuous spatiotemporal sequence according to time slices. The length of the time slice is automatically determined according to the scene dynamic parameters, which are jointly determined by the rate of change of the target's motion speed and the intensity of external environmental disturbances. When the scene dynamic parameters exceed a preset threshold, the length of the time slice is automatically shortened to enhance the temporal resolution. Each data unit in each time slice contains unified reference coordinates, modality identifiers, data quality labels, and initial confidence values, thereby forming a standardized multimodal spatiotemporal data sequence.
[0020] In this embodiment, to ensure that the data collected by the multi-source sensing devices in a complex dynamic environment can achieve a unified correspondence in time and space, the entire process requires precise acquisition, synchronization, registration, and organization steps. Multimodal data refers to heterogeneous data collected by different types of sensing devices, including but not limited to image frames output by visible light cameras, infrared thermal imaging data, lidar point cloud data, millimeter-wave radar echo signals, environmental temperature and humidity sensing data, and gas concentration detection data. These data have significant differences in physical units, sampling frequencies, and spatial resolutions. If they are not processed uniformly, it will be difficult for subsequent neural network models to extract aligned feature information.
[0021] During the data acquisition phase, each sensor is activated collaboratively through the system control unit, and the sampling frequency is dynamically adjusted based on environmental complexity parameters during the acquisition process. Environmental complexity parameters are a comprehensive indicator measuring the rate of change and degree of disturbance in the scene, determined by three main factors: the rate of change of light intensity, the amplitude of change in radar echo density, and the degree of gas concentration disturbance. For example, when sunlight changes from sunny to cloudy, the rate of change of light intensity increases significantly in a short period; when the point cloud density detected per unit time in the radar echo fluctuates greatly, it indicates a rapid change in the number or location of targets within the field of view; and rapid fluctuations in gas concentration may indicate that the external environment is disturbed by wind. The system calculates a comprehensive environmental complexity value by weighted averaging of the above three factors and compares it with a preset threshold. When this value exceeds the threshold, the sampling frequency is increased proportionally, for example, from 10 Hz to 20 Hz, to ensure temporal continuity and data integrity in dynamic scenes.
[0022] During the time synchronization phase, the system uses the high-precision clock of the main control unit as a unified time reference and adds a timestamp to each sampling point. Due to communication delays and internal clock drift between different sensors, the time of directly acquired data often contains errors. To address this, this invention introduces a synchronization mechanism based on time deviation compensation. In the initial stage of acquisition, the time delay value relative to the main control clock is calculated by comparing the timestamp differences of the sampling data from each sensor, and a time compensation curve is generated accordingly. For example, when the radar signal arrives at an average time 3 milliseconds later than the camera frame time, this delay may increase to 5 milliseconds when wind speed increases. The system predicts the delay at the next moment based on this trend and automatically triggers the radar sampling signal in advance to synchronize it with the camera. This compensation process is updated in real time within each sampling period, thus maintaining sub-millisecond time consistency even in the presence of network latency, sensor jitter, or signal congestion.
[0023] During the spatial registration phase, the system first selects several fixed reference points within the sensing scene that can be simultaneously perceived by different sensing devices, such as building corners, dock railings, ship markings, or ground reflective markers. The system calculates the spatial response position differences of these reference points in different modal data. By comparing the spatial positions of the reference points in the image coordinate system, radar coordinate system, and laser point cloud coordinate system, the system progressively corrects the attitude angles (including pitch, yaw, and roll angles) and installation position offsets of each sensing device. Spatial registration is considered complete when the coordinate difference of all sensing devices at the same reference point is below a preset error range (e.g., less than 2 cm or 0.1 degree angular error). This ensures accurate spatial correspondence of all sensing data under a unified reference coordinate system, avoiding registration drift problems caused by installation deviations or vibrations.
[0024] After completing time synchronization and spatial registration, the system organizes all modal data into time slices. A time slice is a data segment with a fixed time length as its boundary, used to correspond and aggregate data from different modalities on the time axis. The time slice length is automatically determined by scene dynamics parameters, which take into account the rate of change of target motion speed and the intensity of external environmental disturbances. For example, when a ship is detected to be sailing smoothly, the rate of change of target speed is low, and the time slice length can be set to 500 milliseconds; when encountering strong winds and waves or berthing operations, the target motion changes drastically, and the system will automatically shorten the time slice to 100 milliseconds to ensure that the temporal resolution is not affected. The data unit within each time slice includes unified reference coordinate information (used to represent the spatial reference of each modal data), modality identifier (used to distinguish different types of sensor data), data quality label (used to indicate whether the data is interfered with or has lost frames), and initial confidence value (used to reflect the reliability of each modal data in that time slice).
[0025] For example, in a typical port scenario, when a ship approaches the dock, the image captured by the camera shows obvious changes in the edge of the target, while the density of the radar echo point cloud increases sharply, thus increasing the environmental complexity parameter. The system automatically increases the sampling frequency to 30 times per second. The master clock detects a 3.6-millisecond delay in the radar signal and automatically triggers sampling in advance, enabling it to complete data acquisition within the same time slice as the infrared camera. The spatial registration process corrects the attitude error of each sensor to within 0.05 degrees by identifying the position of the fixed lighthouse at the dock. The final time slice data sequence is in 200-millisecond units and contains multimodal information under a unified reference coordinate system, providing synchronous and spatially consistent input data for subsequent neural network feature extraction.
[0026] Through the implementation of the above process, this embodiment realizes the adaptive coordination of data acquisition by multi-source sensing devices in dynamic environments, enabling different types of data to be strictly aligned in both time and space dimensions, and the generated multimodal spatiotemporal data sequence has high precision, consistency and real-time performance.
[0027] Step S102: Input the multimodal data sequence into the multimodal deep neural network, extract the feature information of each modality, calculate the modal similarity and confidence change rate of adjacent time slices, and correct the feature weights of each modality according to the confidence change rate to obtain the corrected feature information.
[0028] In step S102, the multimodal data sequence generated in the previous step needs to be input into the multimodal deep neural network to perform feature extraction, temporal analysis, and weight correction on data from different sources. The goal of this step is to enable the system to automatically adjust the feature contributions of different modalities based on the temporal variation patterns and reliability of each modality's information when facing complex dynamic environments, so that the final feature representation is more stable, accurate, and interpretable.
[0029] Multimodal deep neural networks are neural network structures capable of simultaneously receiving and processing multiple types of input data. In this method, they are designed with multiple input channels, each corresponding to a different modality of data. Specifically, the visible light image channel is used to extract external visual morphological features, such as the shape, texture, and edges of the target; the infrared image channel is used to extract thermal distribution features, suitable for identifying objects at night or in occluded environments; the laser ranging point cloud channel is used to capture spatial geometry and distance changes; the radar direction finding channel is used to obtain relative velocity, azimuth, and trajectory information; and the navigation data channel is used to reflect the platform's own position, attitude, and motion trend. Each input channel extracts features through a set of convolutional layers and nonlinear activation layers, and the output is mapped to a feature representation of the same dimension for subsequent fusion.
[0030] When inputting multimodal data sequences, the first step is to encode the multimodal data for each time slice. The encoding process includes two parts: spatial feature encoding and temporal feature encoding. The purpose of spatial feature encoding is to extract the spatial structural features of each modality within a single time slice. For example, for visible light and infrared images, convolutional neural networks can be used to extract edge and brightness variation features of local regions; for laser ranging data, point cloud convolution or feature generation modules based on neighborhood distance statistics can be used to extract three-dimensional spatial distribution patterns; for radar information, one-dimensional convolution or Fourier transform layers can be used to analyze its echo intensity and direction changes; and navigation data can be mapped into attitude and displacement embedding features through fully connected layers. In this way, each modality can form a set of spatial feature vectors within each time slice.
[0031] The purpose of temporal feature encoding is to capture dynamic trends, that is, the evolution of each modality's features over time. Because the environment and target state change over time—for example, a ship bobbing in waves, smoke concentration changing, or a heat source gradually increasing—single-frame static features are often insufficient to accurately reflect the essence of an event; therefore, time series modeling is necessary. For this purpose, recurrent neural networks or temporal attention mechanisms can be used to establish correlations between consecutive time slices. For example, if a gradual increase in infrared intensity, a slight decrease in laser reflection distance, and a change in radar velocity vector from far away to near are detected in three consecutive time slices, the system can determine that the target may be approaching and heating up. Through such temporal analysis, the neural network can form dynamic feature representations, thus providing a foundation for subsequent similarity calculations and confidence adjustments.
[0032] After extracting spatial and temporal features, it is necessary to calculate the modal similarity between adjacent time slices. Modal similarity measures the magnitude and consistency of the same mode's variation across consecutive time slices, reflecting the stability and reliability of current observations. During calculation, the modal feature vectors of two adjacent time slices can be compared, for example, using the average absolute value of feature differences or the feature correlation coefficient. If the feature differences are small or the correlation is high, it indicates that the observations of that mode are stable and reliable; if the feature differences are large or the correlation decreases significantly, it indicates that the mode may be subject to interference or abnormal changes. For example, in clear conditions, visible light and infrared features should maintain high consistency over a short period; if radar signals change drastically due to wave interference, their similarity will decrease significantly. By comparing these similarities, a quantitative index reflecting the stability of each mode can be established.
[0033] After obtaining the modal similarity, the system also needs to calculate the confidence rate of change. Confidence represents the model's reliability assessment of the information represented by a certain modality feature, while the rate of change represents the fluctuation of that confidence over time. The calculation of confidence can be based on the stability of the feature distribution and feature quality assessment of the network output. For example, if an infrared image has large areas of overexposure or underexposure in consecutive time slices, the system will automatically reduce its confidence. The confidence rate of change can be measured by the difference in confidence between adjacent time slices. If the rate of change is small, it indicates that the modality has long-term stable performance, and its weight in feature fusion can be increased; if the rate of change is large, it indicates that it has been recently affected by environmental interference, and its weight should be appropriately reduced. For example, when a radar signal maintains a stable echo for several seconds and is consistent with the laser ranging result, its confidence rate of change is close to zero, and this mode can be considered a high-confidence mode; while when a visible light image suddenly blurs at night, its confidence rate of change is high, and the system should reduce its impact.
[0034] The process of adjusting the feature weights of each modality based on the confidence change rate is crucial to the entire process. Feature weights represent the importance of each modality in the overall feature representation, and the adjustment of these weights should dynamically reflect environmental conditions and modal performance. Specifically, the system assigns an initial weight to each modality, which can be set based on modality type and typical scenario experience; for example, visible light and infrared modalities have different priorities during the day and night. Subsequently, the system adjusts the initial weights based on the current confidence change rate: when the change rate is small, it indicates that the modality's output is stable and reliable, and the weight should be relatively increased; when the change rate is large, it indicates that the modality is severely disturbed, and the weight should be decreased. The magnitude of the adjustment should be determined based on the relative magnitude of the confidence change rate. For example, if the confidence change rate of a modality is only one-tenth of the average, its weight can be increased by about 10%; if the change rate is three times the average, its weight can be decreased by about 30%. This dynamic adjustment allows the network to automatically optimize the contribution ratio of different modalities according to real-time environmental changes, thereby improving the overall robustness of the perception results.
[0035] After completing the above corrections, the system obtains the corrected feature information. This corrected feature information is a comprehensive representation that retains the unique characteristics of each modality while reflecting the balance of credibility among different modalities in the current environment. It includes not only spatial visual, thermal, and geometric information but also incorporates dynamic trends and environmental adaptability over time. For example, in ship monitoring scenarios, if radar signals are stable and infrared thermal images are clear at night, while visible light images are severely affected by illumination, the system will automatically prioritize radar and infrared features after fusion, thus still accurately identifying obstacles or abnormal heat sources ahead. Conversely, in clear daytime environments, the weight of visible light features increases, and the system primarily relies on visual and laser ranging features to identify target outlines.
[0036] Through the processing in step S102, the system can establish a dynamic and self-correcting feature fusion mechanism, enabling the multimodal neural network to have environmental awareness, adaptability, and temporal robustness.
[0037] Furthermore, the step of inputting the multimodal data sequence into a multimodal deep neural network, extracting feature information of each modality, calculating the modal similarity and confidence change rate of adjacent time slices, and correcting the feature weights of each modality based on the confidence change rate to obtain corrected feature information includes: The multimodal deep neural network is composed of an input layer, a feature extraction layer, a time alignment layer, a confidence evaluation layer, and a feature fusion layer connected in sequence. The input layer is used to receive multimodal data sequences that have undergone time synchronization and spatial registration and to perform standardization processing to ensure that different modalities are consistent in resolution, value range and coordinate scale. The feature extraction layer includes parallel branches corresponding to the visible light mode, infrared mode, laser point cloud mode and radar mode, respectively. While sharing the main structure, each branch retains an independent parameter set to extract the texture, temperature gradient, spatial density and echo energy features of its own mode and output the corresponding feature vector. The time alignment layer indexes and compares the multimodal feature vectors of consecutive time slices, calculates the feature differences of the same spatial region in adjacent time slices and generates a time consistency index to reflect the stability of each mode in the time series. The confidence assessment layer calculates the confidence change rate based on the time consistency index and signal quality parameters, including signal integrity, noise ratio, and sampling stability. When time consistency decreases and signal quality deteriorates, the confidence change rate is negative; when time consistency is good and signal quality improves, the confidence change rate is positive. The feature fusion layer dynamically adjusts the feature weights of each mode based on the confidence change rate. When the confidence of the visible light mode decreases under low light conditions, its feature weight is reduced and the feature weight of the infrared mode is increased. When the confidence of the infrared mode decreases under high temperature background, the feature weights of the radar mode and the laser point cloud mode are increased. The corrected features are then normalized and superimposed to generate a comprehensive feature vector with corrected confidence. The comprehensive feature vector remains continuous in the time dimension and aligned in the spatial dimension, and serves as the input feature for subsequent external target recognition and internal event determination.
[0038] In this embodiment, a multimodal deep neural network serves as the core sensing unit of the system, used to achieve deep feature extraction, cross-modal dynamic alignment, adaptive confidence correction, and fusion feature generation from data from multiple sensing devices. The entire network structure sequentially includes an input layer, a feature extraction layer, a time alignment layer, a confidence evaluation layer, and a feature fusion layer. Each layer forms a strict data dependency relationship, and its inputs and outputs are logically interconnected, forming a continuous data flow link.
[0039] During the input phase, the system receives a multimodal data sequence generated from the preceding steps. Each time slice corresponds to a set of synchronized and registered modal data, including visible light image frames, infrared thermal image frames, lidar point cloud frames, and millimeter-wave radar echo frames. The main function of the input layer is to transform heterogeneous data into a tensor structure with a unified format. Due to significant differences in the output characteristics of different sensing devices—for example, visible light images are two-dimensional pixel matrices, infrared images are temperature distribution matrices, point clouds are three-dimensional coordinate sets, and radar signals are time-domain reflectance intensity sequences—the input layer needs to perform standardization and encoding operations. Visible light and infrared data are normalized to pixels, point cloud data is voxelized into dense matrices, and radar signals are transformed using time-frequency conversion to obtain a two-dimensional feature spectrum. All data are then adjusted to the same resolution and coordinate scale, forming four input tensors corresponding to the modalities, identified by time slice index numbers. The output of the input layer is a normalized tensor sequence differentiated by modality, with consistent input dimensions to facilitate parallel input to the next layer.
[0040] The feature extraction layer receives a normalized tensor sequence from the input layer and employs a multi-branch parallel structure, with each branch corresponding to a modal input. The visible light branch includes multi-level convolutional operations and pooling units to extract texture, edge, and spatial structure features, outputting a two-dimensional feature map. The infrared branch extracts heat source distribution features through a temperature gradient calculation unit, outputting a temperature response feature map. The laser point cloud branch extracts spatial density distribution and geometric contour features through three-dimensional convolution and neighborhood aggregation operations, outputting a three-dimensional spatial feature block. The radar branch extracts echo energy, target velocity, and motion direction features through temporal convolution and power spectrum analysis units, outputting a time-frequency feature spectrum. The output of each modal branch is converted into a feature vector of uniform dimension at the end of this layer and appended with a corresponding time-slice index label. The output of the feature extraction layer is four sets of time-calibrated modal feature vector sequences, all with consistent dimensions and strictly corresponding to the time slices.
[0041] The time alignment layer receives modal feature vectors from consecutive time slices, and its core task is to establish cross-modal consistency over time. This layer contains an index matching unit and a feature comparison unit. The index matching unit first pairs feature vectors from different modalities in chronological order based on the time slice number; the feature comparison unit calculates the modal feature differences between adjacent time slices. To avoid directly calculating feature values that cannot be directly compared between different modalities, this layer first independently calculates the rate of change within each modality, and then establishes a relative change ratio between modalities. For example, when the temperature center point position in the infrared feature is stable while the target distance in the radar feature changes slightly in a consecutive time slice, the system determines that the temporal consistency of the infrared mode is higher than that of the radar mode. The time alignment layer ultimately outputs a temporal consistency index for each modality in the current time slice, with a value ranging from 0 to 1; a higher value indicates stronger temporal stability. This output, along with the original feature vectors, is passed to the next layer.
[0042] The confidence assessment layer quantifies the reliability trends of each modality under the current environment. Its inputs include the time consistency index and signal quality parameters output from the time alignment layer. The signal quality parameters, provided by the front-end acquisition module, include signal integrity (the reciprocal of the data loss rate), noise ratio (the ratio of effective signal energy to background noise energy), and sampling stability (the inverse ratio of sampling period deviation). The confidence assessment unit generates an initial confidence value by weighted summation of the time consistency index and signal quality parameters, and obtains the confidence change rate by calculating the confidence difference between adjacent time slices. If time consistency decreases and the noise ratio deteriorates, the confidence change rate is negative; if time consistency improves and signal quality improves, the confidence change rate is positive. This layer uses a sliding window mechanism to smooth the confidence change rate, thereby offsetting the impact of instantaneous outliers. The confidence assessment layer outputs a confidence change rate vector for each modality, the value of which directly determines the feature weight adjustment ratio of the fusion layer.
[0043] The feature fusion layer is the output layer of the entire multimodal deep neural network. Its input includes modal feature vectors from the feature extraction layer and the confidence change rate from the confidence evaluation layer. The fusion layer first dynamically adjusts the feature weights of each modality based on the confidence change rate. The system maintains a set of variable weight parameters at this layer; when the confidence change rate of a certain modality is consistently negative, its weight decreases proportionally; when the confidence change rate of another modality is positive, its weight increases. After weight adjustment, the system normalizes the feature weights of all modalities to ensure the total weight sum equals 1, guaranteeing energy balance in the fusion process. Subsequently, a weighted superposition operation is performed, superimposing the features of each modality according to the corrected weights to generate a comprehensive feature vector. This comprehensive feature vector remains continuous in the time dimension, aligned in the spatial dimension, and has eliminated the influence of noisy modalities.
[0044] The data flow process of the entire network is strictly unidirectional from the input layer to the output layer. That is, the multimodal tensor sequence output from the input layer enters the feature extraction layer, and generates modal feature vectors through multi-branch convolution and feature compression. The time alignment layer receives these feature vectors, calculates the time consistency index, and outputs it. The confidence evaluation layer uses the time consistency index and signal quality parameters to generate the confidence change rate. The feature fusion layer performs weight correction and weighted superposition based on the confidence change rate and outputs the comprehensive feature vector after confidence correction.
[0045] For example, in a port environmental monitoring scenario, when the system receives visible light images, infrared images, point clouds, and radar signals, the input layer standardizes them and sends them to the network; the feature extraction layer extracts the edge, temperature, density, and echo features of each mode; the time alignment layer determines that the time consistency of the infrared and radar modes is higher than that of the visible light mode; the confidence assessment layer calculates that the confidence of infrared and radar increases while that of visible light decreases; the fusion layer then reduces the weight of visible light features and increases the weight of infrared and radar features, and the output comprehensive features accurately describe the dynamic risk areas and moving targets in the nighttime berthing scenario.
[0046] Through the above structural design, the multimodal deep neural network provided in this embodiment realizes a complete closed loop from input data preprocessing, feature extraction, time alignment, confidence quantification to feature fusion, enabling the system to adaptively adjust the modal contribution ratio in a multi-source dynamic environment, and realize a multi-event collaborative perception process with high confidence, strong interpretability and repeatability.
[0047] Step S103: Using the corrected feature information, identify external targets and internal events respectively, generate corresponding candidate results and their confidence information, and form a candidate event set.
[0048] In step S103, the feature information corrected in the previous stage is used to identify and classify external targets and internal events in the dynamic environment, and further generate corresponding candidate results and their confidence information, thus forming a candidate event set. The core of this step is to map the high-dimensional features after multimodal fusion back to specific physical meanings, that is, to recover understandable objects, events and their occurrence states from abstract feature representations. To achieve this goal, the system should clearly distinguish between the two concepts of "external targets" and "internal events" during the design phase. External targets refer to physical entities located in the space outside the monitoring range that have relative motion or interaction with the subject (such as ships, facilities, equipment, etc.), such as other ships, buoys, obstacles, floating objects or port facilities in navigation; internal events refer to abnormal situations occurring in the subject itself or its internal space, such as abnormal temperature, smoke generation, oil and gas leaks, equipment overload, structural deformation, intrusion behavior, etc.
[0049] After receiving the corrected multimodal feature information, the system first needs to identify external targets. External target identification typically relies on spatial distribution and motion patterns, determining targets by detecting the stable presence or regular changes in position, shape, or texture features across consecutive time slices. Visible light and infrared features are used to determine the target's shape and thermal characteristics, laser point clouds provide spatial distance and altitude information, and radar is used to measure relative velocity and azimuth. The system can first form several salient regions in the feature space using a convolution-based feature clustering method, with each region representing a possible external target. For example, if a region exhibits a consistent shape across consecutive frames in visible light features, a stable velocity direction in radar features, and a high temperature contrast in infrared features, then this region is highly likely to correspond to a ship in motion. To improve accuracy, the system can perform cross-validation between multimodal features. For instance, if a heat source detected by infrared has a stable spatial profile in laser ranging and a corresponding velocity signal exists in radar, then the heat source can be confirmed as a real target rather than background noise.
[0050] Meanwhile, internal event identification relies more on physical changes and status monitoring of local areas. The system can comprehensively analyze signals such as abnormal temperature distribution in infrared images, concentration trends from smoke detection sensors, gas composition changes from oil and gas detection modules, data fluctuations from structural strain sensors, and dynamic area detection from intrusion detection cameras. When the temperature in a certain area rises sharply in a short period of time, or the output signal of the smoke sensor continues to increase beyond the warning threshold, the system will mark this phenomenon as a candidate result for "temperature anomaly event" or "smoke generation event." If multiple abnormal signals from different sources are detected simultaneously in the same spatial area, such as temperature rise accompanied by smoke generation, the system will further infer a possible fire risk event. For oil and gas leak events, the system analyzes the rate of change of hydrocarbon concentration in the air and its relationship with wind direction and speed to determine whether the abnormal gas originates from a specific compartment or equipment area. For structural deformation or intrusion events, vibration or displacement sensors installed at key nodes can be used to determine abnormal forces or external force contact, thereby forming corresponding candidate event records.
[0051] After external target identification and internal event identification are completed, confidence information needs to be calculated for each identification result. Confidence is a quantitative evaluation value of the correctness of the identification result by the system, reflecting the degree of credibility of the event being identified as a specific category. The calculation of confidence comprehensively considers factors such as multimodal feature consistency, signal quality, temporal stability, and contextual rationality. Specifically, when multiple modalities give consistent classification results for the same object, the confidence is high; when the output results of each modality are inconsistent or conflicting, the confidence is low. For example, when radar, laser, and infrared simultaneously confirm the existence of a moving target ahead, and its position and speed information match, the confidence of the identification result can be determined as "high"; if the visible light image is blurry and the infrared signal is unstable, and only the radar detects the moving signal alone, the confidence is set to "low". In addition, the system also considers the impact of the signal environment on the identification accuracy. For example, under conditions such as low light at night, rainy or foggy weather, or strong light reflection, the confidence of some modalities will be automatically reduced by a preset attenuation coefficient to prevent environmental noise from causing misjudgment.
[0052] After candidate results are generated, the system needs to organize all identified external targets and internal events, along with their confidence levels, into a candidate event set. The candidate event set can be understood as a list of all events that the system considers "possible" for a given time slice. Each event in this set should contain at least three elements: event category (i.e., event type label), event spatial location (corresponding to the center point or range under a unified reference coordinate system), and confidence level value. For dynamic events, time range information should also be included to reflect the duration of the event. For example, if the system detects a heat source at the same location accompanied by rising smoke signals in five consecutive time slices, the time range of this event can be labeled as the corresponding duration. If the heat source disappears or the signal returns to normal in subsequent time slices, the system will automatically end the recording of that candidate event.
[0053] In practice, updating and maintaining the candidate event set is a dynamic process. Each time the system receives new time-slice data, it re-performs the identification, matching, and merging operations. If a newly identified event highly overlaps with existing events in the set in terms of spatial location and feature distribution, it is considered a continuation of the same event, and the system updates its confidence level and time range. If the new event is spatially independent or has different characteristics, it is added as a new candidate event. For example, when a ship detects multiple heat source movement trajectories during navigation, if one trajectory overlaps with the ship's radar heading, the system will consider it a continuous observation result from the same ship; while another heat source appearing on the starboard side of the hull and having different wave reflection characteristics may be separately identified as a floating object. For internal events, such as multiple temperature sensors detecting independent anomalies in different compartment areas, the system will record them separately as independent events to avoid confusing the sources.
[0054] Confidence updates can be smoothed based on temporal continuity to avoid false alarms caused by anomalies in a single frame. A common method is to incorporate a weighted average of the confidence scores from the previous time slice when calculating the new confidence score, making its change over time smoother. For example, if the confidence score of a fire candidate event was 0.8 in the previous moment and has risen to 0.9 in the current moment, the system can set the updated confidence score as the weighted average of the two scores based on a set smoothing coefficient to avoid abrupt changes. This mechanism can reflect the development trend of events, enabling the system to gradually increase its alertness in the early stages of an event, rather than triggering frequent false alarms due to short-term fluctuations.
[0055] Through the above process, step S103 realizes the identification mapping from abstract features to specific events and establishes a candidate event set that includes external targets and internal anomalies.
[0056] Furthermore, the process of using the modified feature information to identify external targets and internal events respectively, generating corresponding candidate results and their confidence information, and forming a candidate event set includes: Based on confidence-corrected multimodal feature information, a spatiotemporal fusion feature set is constructed under a unified reference coordinate system. Feature information from visible light, infrared, lidar, and radar modes is spatially mapped to form a multidimensional fusion feature block that can reflect scene texture, temperature distribution, spatial morphology, and energy reflection characteristics. Spatial partitioning and region similarity aggregation are performed on the multidimensional fusion feature block for each time slice to obtain a set of candidate target regions. Each candidate target region includes a feature center point, boundary contour, and temporal continuity label. Based on the temporal continuity label of the candidate target area, the consistency of the target trajectory between adjacent time slices is calculated, and external target candidates are output. The external target candidates include the direction of movement, the rate of change of velocity, and the position confidence interval. When the temporal consistency index is higher than the preset threshold and the trajectory is continuous, the external target type is determined as a valid candidate target. The external target candidates are compared and analyzed with the environmental state information to identify floating objects on the water surface and ship obstacles. Using the infrared temperature distribution, gas concentration features, and acoustic energy features contained in the fusion feature set as input, the internal event identification process is executed. The type and location of internal abnormal events are determined by deviation analysis from the normal operation state template library. Internal events include temperature abnormalities, smoke generation, oil and gas leaks, and equipment vibration exceeding limits. Candidate internal events are generated based on the deviation amplitude and duration. The joint confidence level calculation is performed on external target candidates and internal event candidates, taking into account feature completeness, temporal continuity, signal stability, and spatial correlation. Each event is assigned a confidence level, and event coupling relationships are established based on temporal overlap and spatial correlation. Events with high correlation are integrated into composite candidate events, and a candidate event set is output. Each item in the candidate event set includes event type, spatial location, time range, and confidence level, which are used in subsequent collaborative fusion and decision analysis stages to achieve integrated identification and dynamic management of external targets and internal events driven by multimodal features.
[0057] In this embodiment, using confidence-corrected multimodal feature information to identify external targets and internal events is the core step of the entire collaborative perception process. Its task is to transform feature signals in a complex multi-source environment into a set of candidate events that can be used for dynamic monitoring and decision-making.
[0058] First, the confidence-corrected multimodal feature information needs to be uniformly mapped to achieve fusion of different sensing modalities under the same spatial reference. Here, "multimodal feature information" refers to the set of feature vectors extracted by devices such as visible light sensors, infrared sensors, lidar, and millimeter-wave radar. The visible light modality reflects the texture and geometry of the scene, the infrared modality characterizes temperature gradients and thermal anomaly distributions, the lidar modality records spatial depth and point cloud density, and the millimeter-wave radar modality reflects energy reflection characteristics and the radial velocity of moving targets. The system registers the four modal features according to time slices using a unified reference coordinate system, ensuring a one-to-one correspondence between features of the same spatial region under different modalities. During fusion, the feature vectors of each modality are weighted and superimposed according to confidence levels to generate a multidimensional fused feature block reflecting the overall environmental state. Each spatial unit of this feature block has four attribute values: light intensity, temperature, depth, and energy reflection, which can be directly used for subsequent target recognition.
[0059] After obtaining the fused feature block, the system spatially divides it into time slices. Each time slice can be considered a dynamic snapshot, representing the environmental state at a certain moment. The system first divides the fused feature block into several spatial unit regions, and then calculates the similarity between these regions. Similarity refers to the degree of consistency of regions across multimodal feature dimensions. For example, when adjacent regions show little variation in optical texture and infrared thermal distribution, and are continuous in laser point cloud density, their similarity is high, and the system will merge them into a single candidate target region. To facilitate subsequent time tracking, each candidate target region generates a temporal continuity label, which records the positional correspondence of the target in consecutive time slices, used to determine whether it is a stable object or temporary noise.
[0060] Subsequently, the system calculates a consistency index for the target trajectory based on the time continuity label. Trajectory consistency reflects whether the target's motion trajectory is continuous and smooth across consecutive time slices. Specifically, the system judges consistency by comparing the change in the spatial center point position and the rate of change in velocity direction of candidate targets in adjacent time slices. When the target's positional shift is stable across multiple time slices, the rate of change in velocity is smooth, and there are no abrupt changes in direction, the trajectory consistency index is high, indicating that the target's motion is real and traceable. If the consistency index exceeds a preset threshold, the system marks the candidate region as a valid external target and outputs information such as its motion direction, rate of change in velocity, and position confidence interval. Such external targets may include ships, floating objects, navigation marks, or other obstacles. To further improve recognition accuracy, the system compares the external target features with an environmental state database, such as using water surface ripple characteristics and weather parameters to determine whether it is a real target, thereby eliminating false detections caused by light reflection, rain, fog, or other disturbances.
[0061] After completing external target identification, the system enters the internal event identification phase. Internal events refer to abnormal phenomena occurring within the monitored object's internal space, such as abnormal cabin temperature, oil and gas leaks, smoke generation, or excessive equipment vibration. In this phase, the system uses the infrared temperature distribution, gas concentration changes, and acoustic energy characteristics from the fused feature set as input, and determines the event type by performing deviation analysis with the normal operating state template. The template library is constructed from historical data under normal operating conditions, containing the characteristic range of each mode in a stable state. When real-time data continuously deviates from the template characteristic value exceeding a threshold in a certain mode for a certain duration, the system determines it as an anomaly. For example, in engine compartment monitoring, if the temperature distribution center of the infrared mode rises by more than 20%, and this change lasts for more than three seconds, the system marks that area as a temperature anomaly event; if simultaneously, the gas sensor shows an increase in hydrocarbon gas concentration, it is further identified as an oil and gas leak event. All identified internal events will generate event entries, including the anomaly type, location, duration, and deviation magnitude.
[0062] Next, the system performs a joint confidence calculation on external target candidates and internal event candidates. The confidence calculation is based not only on the characteristic stability of a single event but also on its spatial and temporal correlation with other events. To this end, the system introduces four core parameters: feature completeness, temporal continuity, signal stability, and spatial correlation. Feature completeness assesses the completeness of data across modalities; temporal continuity assesses the stability and persistence of an event across time slices; signal stability reflects fluctuations in sensor signal quality; and spatial correlation analyzes the degree of spatial overlap between events. The system assigns a confidence level to each event based on the weighted average of these four parameters, typically ranging from 0 to 1. For example, when an event has a feature completeness of 0.9, a temporal continuity of 0.8, a signal stability of 0.85, and a spatial correlation of 0.7, the overall confidence level, calculated using weighted ratios of 0.4, 0.3, 0.2, and 0.1, is 0.83. This confidence level can be used to determine whether the event should be included in the final candidate set.
[0063] After obtaining the confidence level results, the system further analyzes the spatiotemporal correlation between external targets and internal events to construct event coupling relationships. Event coupling refers to the temporal and spatial synchronicity between changes in the external environment and anomalies in the internal state. For example, if the system detects a high-speed approaching object (external target) outside the ship's hull and almost simultaneously detects a sharp increase in the vibration amplitude of the steering gear compartment (internal event), then a high degree of coupling between the two events can be considered. The system determines whether the two events are the same composite event by calculating the temporal overlap (the ratio of the intersection of event durations to the total duration) and spatial proximity (the ratio of the distance between the locations of the two events to the system's monitoring range). When the overlap exceeds 0.6 and the spatial proximity is less than 0.1, the system merges the two events into a single composite event entry.
[0064] Finally, the system integrates all the filtered external targets and internal events, outputting a candidate event set. Each entry in this set contains information such as event type, spatial location, time range, and confidence level, which can be directly used as input for subsequent collaborative fusion and decision analysis.
[0065] Step S104: Perform weighted fusion and conflict resolution based on the differences in confidence information and temporal consistency in the candidate event set, establish event priority relationships, and obtain a collaborative event set containing event labels, spatial locations, and time ranges.
[0066] In step S104, the candidate event set formed in the previous stage needs to be analyzed and integrated in depth to resolve issues of overlapping, conflicting, or redundant results that arise during the identification of multi-source data and multiple events. The core task of this step is to merge the identification results of multiple candidate events into a unified and reliable collaborative event set through a weighted fusion and conflict resolution mechanism, and further establish the priority relationships between events. This not only eliminates identification biases between different modalities or different time slices but also ensures that the system can judge and respond in a reasonable order when faced with multiple complex events occurring simultaneously.
[0067] In this process, the results in the candidate event set first need to be normalized and aligned. Since candidate events may originate from independent identification results of different time slices or different modalities, they may differ in spatial coordinates, temporal indexes, and confidence distributions. The system should first compare the spatial positions of all candidate events based on a unified reference coordinate system, considering events with a spatial distance less than a set threshold as different descriptions of the same physical region. For example, if the location of an obstacle detected by radar and the target point cloud identified by laser ranging differ by only tens of centimeters in space, they can be identified as the same object; if the bright spot area in a visible light image and the center of the heat source area in an infrared image are essentially coincident, they are considered different modalities reflecting the same event. For temporal alignment, the time slice should be used as the basic unit. If adjacent time slices contain event records of the same type, close location, and smooth confidence changes, they should be considered as a temporal continuation of the same event, rather than independent events.
[0068] After initial matching, weighted fusion needs to be performed based on the differences in confidence levels. Weighted fusion refers to calculating a comprehensive result based on the confidence values and modal weights of multiple candidate events corresponding to the same spatial region or time period, so that the modality with higher confidence has a greater impact on the final judgment. To achieve this, the system should first determine the fusion weight for each candidate event. The determination of the weight can comprehensively consider the event's confidence value, the stability of the data source modality, and temporal continuity. For example, when the confidence of infrared detection of a heat source is 0.9, while the confidence of visible light identification of the same area is only 0.6, the weight of the infrared modality can be increased in a nighttime environment to reflect its advantage under low-light conditions. Conversely, in a daytime high-brightness environment, the weight of the visible light modality can be increased accordingly. If the confidence of an event changes little over multiple consecutive time slices, it indicates high temporal stability and should also receive additional weight. This ensures that the fusion result is more consistent with the actual environmental conditions.
[0069] Weighted fusion calculations can be performed through a stepwise update process. The system establishes an event fusion record for each spatial region. When a new candidate event appears, its confidence level is compared with the existing record, and a combined confidence level is calculated based on the weight ratio of the two. For example, assuming an existing event had a combined confidence level of 0.8 in the previous time slice, while a similar event detected in the current time slice has a confidence level of 0.9, and the new data has twice the weight of the old data, the combined confidence level will be updated to a median value close to 0.87. This fusion process effectively filters out transient noise signals, enabling the system to have a smooth response to continuously occurring events.
[0070] In multi-source sensing, conflicts often arise between modes, meaning different modes give contradictory or inconsistent judgments about the same event. For example, a moving target might be detected in a radar signal, but no corresponding outline is found in a visible light image, or an infrared thermal signal might show a high-temperature area, but a laser point cloud might show no reflection point. Conflict resolution is necessary in these cases. The principle of conflict resolution is based on confidence level, modal reliability, and temporal consistency, eliminating or reducing the influence of unreliable information. Specifically, the system first checks the confidence level difference between conflicting events. If one confidence level is significantly higher than the other, the system retains the high-confidence event and ignores the low-confidence event. If the confidence levels are close, temporal consistency needs to be further assessed, i.e., whether the trends of change of the two modes are consistent across consecutive time slices. If one modal signal is stable across multiple consecutive time slices, while the other modality only appears momentarily within a single frame, the former is more likely to be a real event, and the latter is considered interference. If the difference remains unclear, the system records the conflicting event and adds it to the observation queue, continuing monitoring in subsequent time slices until sufficient temporal evidence is obtained before making a final judgment.
[0071] After weighted fusion and conflict resolution are completed, the system establishes a priority relationship for events based on their confidence level, spatial impact range, and potential risk level. Event priority is the basis for determining the processing order when multiple events occur simultaneously; it reflects the degree of impact of an event on the overall system safety or operational status. Typically, determining event priority considers the following factors: first, the level of confidence—events with higher confidence have higher priority; second, the severity of the event type—for example, safety-related events such as fires and collision risks have higher priority than general temperature fluctuations or equipment malfunctions; third, the timeliness of the event—events that occurred recently or are currently occurring are prioritized; and fourth, spatial relevance—events closer to the main body or with a larger impact range should have higher priority. The system can integrate these factors and generate an event priority list using a weighted sorting method. For example, when the confidence level of an external event "obstacles approaching" is 0.95, while the confidence level of an internal event "cabin temperature rise" is 0.85, the system will determine the former as a high-priority event and immediately generate navigation adjustment or collision avoidance commands, while the latter will subsequently trigger a warning or cabin ventilation control.
[0072] While establishing priority relationships, the system also needs to generate a collaborative event set containing event tags, spatial locations, and temporal ranges. Event tags are used to identify the category and nature of events, such as "target approaching," "smoke generation," "oil and gas leak," and "temperature anomaly." Spatial locations are represented using a unified reference coordinate system, specifying the exact location or area of the event. The temporal range reflects the duration of the event, starting from the time slice when it is first detected until the event signal disappears or stabilizes. By integrating spatial, temporal, and semantic information, the collaborative event set can comprehensively describe all important events in the current dynamic environment, providing a foundation for subsequent risk assessment and early warning.
[0073] This process is particularly important in shipboard or industrial monitoring scenarios. For example, when the system simultaneously detects three events—"moving target ahead," "abnormal increase in deck temperature," and "increased smoke concentration in the compartment"—weighted fusion will confirm that the moving target and temperature anomaly are both high-confidence events, while the smoke concentration increase has a relatively low confidence level. After conflict resolution, the system determines that the moving target and temperature anomaly are independent events, assigning them "Level 1" and "Level 2" priorities respectively, while the smoke event is temporarily placed under observation. Subsequently, the collaborative event set will record the category, spatial range, and time of occurrence of these three events, providing input for the subsequent integrated decision-making module to generate multi-event perception results.
[0074] Through the aforementioned mechanism, step S104 ensures that the system maintains the consistency and logic of the identification results under complex conditions of multimodal, multi-time-slice, and multi-event concurrency. By comprehensively analyzing confidence differences, temporal consistency, spatial overlap, and event severity, the system not only eliminates conflicts and redundancies but also provides clear and actionable structured event results for subsequent risk analysis and control.
[0075] Furthermore, the step of performing weighted fusion and conflict resolution based on the differences in confidence information and temporal consistency in the candidate event set to establish event priority relationships and obtain a collaborative event set containing event labels, spatial locations, and temporal ranges includes: The confidence information of the candidate event set is normalized and mapped to obtain the relative confidence ratio. The event type, spatial location and time slice index of the candidate event are used as the retrieval key to perform spatial aggregation within the same time slice to generate local event nodes. The local event nodes carry the weighted spatial center position, boundary contour, relative confidence ratio and the initial time range formed by its source time slice index. For local event nodes with a spatial overlap rate reaching a preset threshold and consistent event type within adjacent time slices, time extension fusion is performed. Multiple local event nodes are merged into intermediate event units according to time continuity rules. The time range of the intermediate event unit is defined by the earliest time slice index and the latest time slice index and includes all time slices in between. The relative confidence ratio is updated to a weighted average of the relative confidence ratios of the participating nodes, and its spatial center position and boundary contour are updated synchronously. The time continuity rules include considering time slice indices that are adjacent or separated by no more than a preset time slice interval as continuous. After obtaining multiple intermediate event units, conflict resolution is performed based on the overlap between spatial location and temporal range. When intermediate event units overlap spatially but do not overlap temporally, the intermediate event unit with the higher relative confidence ratio is retained. When intermediate event units overlap both spatially and temporally but have different event types, the dominant intermediate event unit is determined based on the order of occurrence, duration, and rate of change of feature intensity, and its derivative relationship with other intermediate event units is recorded. Intermediate event units that have undergone conflict resolution are semantically named and solidified as events according to event type. Events inherit the time range, spatial location and relative confidence ratio of the corresponding intermediate event units. The event priority index is calculated by weighting the event importance coefficient, relative confidence ratio and time urgency. The time urgency is determined by the growth trend of feature intensity in the most recent time slice and the length of the time range. Events are sorted from highest to lowest according to their event priority index and output as a collaborative event set. Each event in the collaborative event set contains an event label, spatial location, time range, and priority identifier.
[0076] In this embodiment, a collaborative event set is generated based on a candidate event set through confidence information normalization, temporal consistency analysis, weighted fusion, and conflict resolution. The events are then prioritized hierarchically. The entire process focuses on ensuring the continuity, consistency, and reliability of events in time and space, thereby achieving logical connections and priority distinctions among multiple events.
[0077] At the start of processing, the confidence information in the candidate event set is first normalized to a uniform scale. Confidence is a numerical value that measures the reliability of event identification and may be generated independently from data of different modalities, such as radar, infrared, laser point clouds, or visual signals. Since the confidence distribution range and sampling characteristics differ between modalities, direct comparison can lead to imbalance. Therefore, this invention uses a relative confidence ratio, which involves subtracting the minimum value within the monitoring period from the original confidence value of each event, and dividing by the difference between the maximum and minimum values, mapping it to the range of 0 to 1. After this processing, the event confidence values under different modalities can be compared and weighted on a uniform scale.
[0078] Within the same time slice, the system performs spatial aggregation on candidate events based on event type and spatial location. The goal of spatial aggregation is to identify redundant events within adjacent regions at the same time. For example, the same target may be detected simultaneously by different modalities, thus requiring spatial merging. The spatial aggregation process uses a spatial similarity threshold for judgment. If the Euclidean distance between the event center points is less than a set distance threshold and the event types are consistent, these events are merged into a single local event node. During merging, a confidence-weighted average is used for the spatial center position, meaning that events with higher confidence have a greater impact on the final position, thereby reducing interference from low-quality signals. Each local event node carries its own time slice index and stores the weighted spatial center, boundary contour, and mean confidence value; this information constitutes the basic description of the node.
[0079] In the temporal dimension, the system analyzes the evolutionary relationships of events between different time slices. When local event nodes in adjacent time slices have similar spatial locations (spatial overlap exceeding a threshold) and consistent event types, the system considers them to be temporal extensions of the same event and merges them into intermediate event units. The judgment of temporal extension is based on the "temporal continuity rule," that is, if the index difference between adjacent time slices does not exceed a preset gap threshold, they are considered continuous. After merging, the temporal range of the intermediate event unit is determined by the earliest and latest time slice indices, the spatial center and boundary contour are updated according to the weighted average of the spatial parameters of the participating nodes, and the confidence score is updated to the weighted average of all participating nodes. In this way, the system can integrate discrete instantaneous detection results into a complete temporal event trajectory.
[0080] Once multiple intermediate event units are obtained, the system enters the conflict resolution phase. Conflict resolution is used to address situations where multiple overlapping or contradictory events exist within the same spatial region. If two intermediate event units overlap spatially but not temporally, they are considered to belong to different event processes. The intermediate event unit with higher confidence is retained, while the one with lower confidence is discarded to avoid duplication. If two intermediate event units overlap both spatially and temporally but have different event types (e.g., simultaneous occurrence of "smoke" and "heat source anomaly"), the system will determine the dominant event and the derivative event based on the order of occurrence, duration, and rate of change of characteristic intensity. The dominant event is usually the one with a longer duration or a more significant characteristic growth rate, while the derivative event is marked as a subordinate state and is only used to analyze the causal relationship between events.
[0081] After conflict resolution, the system semantically names the retained intermediate event units, forming the final event set. The semantic naming process matches preset event type templates based on the event's characteristic patterns, such as "target movement," "oil and gas leak," and "abnormal structural vibration." Each event inherits the time range, spatial location, and confidence information of its source intermediate event unit. To prioritize events, the system calculates an event priority index by comprehensively considering the event's importance coefficient, relative confidence, and time urgency. The importance coefficient is determined by the event type; for example, "collision risk" has a higher weight than "slight temperature rise." Confidence reflects the reliability of the detection results; time urgency is assessed based on the event's time range length and the increasing trend of feature intensity within the most recent time slice. When the feature intensity of an event continuously increases in the most recent time slice, its time urgency is higher, thus increasing its priority.
[0082] Ultimately, the system sorts all events from highest to lowest priority, forming a collaborative event set. Each event entry includes an event label, spatial location, time range, and priority identifier.
[0083] Step S105: Generate multi-event collaborative perception results based on the collaborative event set. The multi-event collaborative perception results include risk levels and corresponding early warning or control instructions.
[0084] In step S105, the system generates the final multi-event collaborative perception result based on the collaborative event set formed in the previous step, in order to achieve quantitative assessment and response control of risk status in complex dynamic environments. The purpose of this step is to further structure the clearly defined event labels, spatial locations, temporal ranges, and priority information in the collaborative event set to form risk level judgments and control decisions that the system can execute. The multi-event collaborative perception result refers to the overall security assessment output generated by the system after comprehensively analyzing the spatial correlation, temporal overlap, and potential causal relationships between multiple events. This output includes the interaction between events, risk level classification, and corresponding automatic response instructions or early warning prompts.
[0085] Before starting the analysis, the system needs to structure all events in the collaborative event set to have a unified descriptive format. Each event record includes an event tag (e.g., obstacle approach, smoke generation, temperature anomaly, oil and gas leak, etc.), spatial coordinate range, time period, confidence level, and priority information. For spatial coordinates, the relative position of the event and the monitored subject (e.g., vessel or facility) should be calculated based on a unified reference coordinate system; for time information, the start time and duration of the event should be included for subsequent determination of whether there is temporal overlap or continuity. During the processing, the system automatically filters out events with too low confidence or too short duration that are judged as misidentified, thereby ensuring that the perception results are based only on reliable data.
[0086] Building upon this foundation, it is necessary to calculate the spatial and temporal correlations between events. Spatial correlation refers to whether different events have overlapping or close areas in space. This process can be achieved by calculating the minimum distance between event boundaries. When the spatial center distance between two events is less than a preset threshold, or their boundary ranges overlap, the system considers them to be spatially correlated. For example, when an "obstacle approaching" event is detected externally, and an "abnormal hull structural strain" event is simultaneously detected internally, and both are located in the same side area and their time periods overlap, it can be inferred that these two events may have a physical causal relationship, i.e., the external collision caused the internal structural anomaly. Temporal correlation is used to determine whether different events have synchronicity or sequential dependence in the time dimension. If the end time of one event is very close to the start time of another event, or their durations overlap, the system will mark them as time-related events. For example, when "abnormal temperature rise" is detected followed by "smoke generation," the system can identify that these two events belong to different stages of the same potential fire process. Spatial and temporal correlation analysis enables the system to transform isolated single events into logically related event groups, thereby constructing a collaborative perception model among multiple events.
[0087] After obtaining the spatial and temporal relationships of events, the system needs to conduct risk assessment and classification of the events. Risk level refers to the severity of the impact an event may have on the safety or operational status of the entity, and is the core basis for the system's decision-making and response. Risk levels are typically divided into several tiers, such as low risk, medium risk, high risk, and extremely high risk. The system can comprehensively consider the following factors for quantitative risk assessment: First, the confidence level of the event; the higher the confidence level, the greater the likelihood that the event is considered a real risk. Second, the hazard of the event type; different types of events have different potential hazards; for example, events like "fire," "explosion," and "collision" have a higher weight than "temperature fluctuations" or "minor leaks." Third, the spatial proximity of the events; events closer to the entity and covering a larger area have a higher risk. Fourth, the temporal duration and evolution trend; events with a long duration and gradually increasing confidence are more likely to evolve into high-risk events. Fifth, the synergistic effect between events; if multiple medium-risk events overlap spatially or occur consecutively in time, their overall risk level can be increased. When calculating risk levels, the system can employ a weighted cumulative method, assigning different coefficients to confidence level, type weight, spatial proximity, and temporal duration for comprehensive calculation. For example, if the confidence level of "temperature anomaly" is 0.8, the type weight is 0.7, the duration exceeds a set threshold, and it spatially overlaps with the "smoke generation" event, the system can assess its overall risk level as "high risk." Those skilled in the art can flexibly implement the risk calculation logic based on these definitions and weighting rules.
[0088] After completing the risk assessment, the system needs to generate corresponding warning or control commands. Warning commands are used to alert operators or higher-level systems to potential risks in the current environment, while control commands are actions automatically triggered by the system in high-risk or emergency situations. Warning commands are usually issued in the form of sound, light, text, or remote communication signals. For example, when an "obstacle approaching" is detected, the system can display "Danger ahead, please slow down" on the control panel display; when "smoke generation" is detected, it can activate the audible and visual alarms and mark the warning area. Control commands directly act on equipment or system actuators. For example, when a "high fire risk" is determined, the system can automatically close the compartment ventilation valves, activate the fire extinguishing device, or cut off the power; when a "collision risk" is determined, it can automatically issue an evasive steering or deceleration command. These commands are generated and executed according to event priority, ensuring that high-priority risks are handled immediately, while low-priority events are delayed or handled through manual confirmation.
[0089] Furthermore, when generating perception results, the system must also consider the evolutionary characteristics of events over time. Events in dynamic environments often do not occur or end instantaneously, but rather develop gradually as the environment changes. For example, a rise in temperature may evolve into a combustion event after several minutes, or a continuous increase in the relative speed of an external target may indicate an impending collision. Therefore, the system should update the status and risk level of events in real time when generating collaborative perception results. To this end, a sliding time window strategy can be adopted to continuously analyze event data from several recent time slices. When the system detects an upward trend in the confidence level or hazard parameter of an event, it will upgrade the event from "medium risk" to "high risk" and recalculate the overall risk index of the related event group. Conversely, when the event signal gradually weakens or disappears, the system will mark it as "deactivated" and automatically remove it from the perception results to avoid long-term false alarms.
[0090] In practical applications, such as ship safety monitoring scenarios, the effectiveness of this step is particularly evident. When an external multimodal perception system simultaneously detects three events—"moving obstacle ahead," "rapidly decreasing distance to the starboard radar target," and "increased strain on the starboard side of the hull"—the system can determine through collaborative analysis that these events have a clear spatial and temporal correlation, inferring a potential starboard collision risk. At this point, the system will assess the risk level of this combined event as "extremely high risk" and immediately generate a two-level response: the first level is an immediate audible and visual alarm and voice prompt; the second level is an automatic control command to initiate emergency deceleration or yaw maneuvers to avoid a collision. Simultaneously, the system can also transmit the event information to a shore-based monitoring center via a communication network for remote synchronous early warning. If the event is mitigated within the following seconds, such as a change in the obstacle's course or successful avoidance by the ship, the system will adjust the risk level to "medium risk" or "low risk" in real time and record the entire event data for subsequent analysis.
[0091] Through the above process, step S105 realizes a closed loop from event recognition to decision execution, enabling the system to have autonomous judgment and response capabilities.
[0092] The following is a specific example. This embodiment takes ship navigation and safety monitoring as a typical application scenario and fully illustrates the entire process from data acquisition to risk decision control, covering all implementation steps from S101 to S105.
[0093] For example, this method is applied to a medium-sized intelligent ship equipped with a multimodal perception system to achieve automatic perception and early warning functions under complex sea conditions. The ship carries a multi-source sensing device, including a visible light camera and infrared thermal imager mounted on the bow, laser ranging units arranged on both sides, a top-mounted millimeter-wave radar, and temperature sensors, gas concentration sensors, strain sensors, and cabin smoke sensors embedded in the hull structure. It also includes inertial navigation and GPS positioning devices to provide attitude and position references. All sensing devices are connected to a central processing unit via a high-speed communication bus for unified acquisition, synchronization, and processing of sensing data.
[0094] In the initial stage of system operation, step S101 is executed, namely multimodal data acquisition and temporal-spatial alignment. During nighttime navigation, all sensors simultaneously sample at preset frequencies. The visible light camera acquires images at a rate of 30 frames per second, while the infrared thermal imager outputs thermal radiation distribution at a rate of 20 frames per second. The laser rangefinder scans the 3D point cloud 10 times per second, and the millimeter-wave radar outputs target distance and velocity information at 5 Hz. The navigation system provides latitude, longitude, and attitude angle data updated every second. The system's main control unit adds a unified timestamp to all sensor data using a high-precision clock and aligns all modal data in units with a minimum time slice length of 100 milliseconds. Within the same time slice, if the sampling frequency of a certain mode is low, data from nearby times or interpolation methods are used to fill in the gaps, ensuring temporal consistency. In the spatial registration stage, the system transforms the coordinate system of each sensor to a unified reference coordinate system with the ship's center as the origin using calibration parameters established during the installation stage, thus aligning all modal data in space. For example, heat sources in infrared images, reflection points in laser point clouds, and azimuth information in radar can all be mapped to the same physical location. After time synchronization and spatial registration are completed, the system organizes the data in time slice order to form a multimodal data sequence containing visible light images, infrared thermal maps, point cloud data, radar velocity fields, and navigation pose information.
[0095] After the multimodal data sequence is generated, the system proceeds to step S102, where the data is input into a multimodal deep neural network for feature extraction and weight correction. The network input includes five modalities, each processed through an independent feature extraction branch. Visible light and infrared images undergo convolutional layers to extract texture and temperature gradient features; laser point clouds undergo sparse convolutional layers to obtain spatial contour features; radar data undergoes temporal convolutional layers to extract target velocity patterns; and navigation information is converted into position and attitude embedding features through fully connected layers. All features then enter the temporal coding module to analyze the trends of each modality over time. The system calculates the modal similarity and confidence rate of change between adjacent time slices to determine the reliability of each modality in the current environment. For example, when sea fog worsens at night, the clarity of visible light images decreases, and the confidence rate of change increases. Therefore, the system automatically reduces the weight of the visible light modality while increasing the weights of the infrared and radar modalities. The corrected feature information better reflects the current actual environment, making the results of subsequent recognition processes more stable.
[0096] Next, step S103 is executed, using the corrected feature information to identify external targets and internal events, and generating a candidate event set. At this point, the system detects a high-temperature area in the infrared image, located approximately 30 meters ahead of the starboard side of the hull. Simultaneously, radar shows a moving target approaching at a speed of 3 meters per second, and laser point cloud data shows dense, regularly shaped reflection points in the area. After multimodal consistency assessment, the system determines the target to be an external object and generates a "proximity of an external target" event candidate with a confidence level of 0.93. At the same time, the internal temperature sensor detects a 4-degree Celsius increase in temperature in the mid-starboard section of the starboard side over the past 20 seconds, and the smoke sensor signal shows a slight increase. The system generates a "local temperature anomaly" event candidate with a confidence level of 0.76. The event tags, spatial locations, time slice ranges, and confidence values of both candidate events are recorded and added to the candidate event set.
[0097] In step S104, the system performs weighted fusion and conflict resolution on the candidate event set to establish event priority relationships. First, the system compares the spatial coordinates and temporal overlap of the events, determining that the "external target approaching" and "local temperature anomaly" events have no spatial overlap and are independent events. Then, the system analyzes the confidence distribution and temporal consistency of each event, performing weighted fusion on the detection results in adjacent time slices to smooth fluctuations. For example, the confidence of the "external target approaching" event remains stable above 0.9 in five consecutive time slices, confirming its extremely high reliability. For the "local temperature anomaly" event, although the confidence fluctuates slightly, the continuous upward trend is obvious, and the system uses temporal consistency analysis to weightedly increase its confidence to 0.8. Based on the event type and confidence, the system determines that external collision risk events have a higher priority than cabin temperature events and generates a collaborative event set, where each event has an event tag, spatial location, temporal range, and priority level.
[0098] Once the coordinated event set is generated, the system proceeds to step S105, where risk assessment and control command generation are performed on all events. First, the system calculates the risk level based on the event type and confidence level. External target approach events are assessed as "high risk" due to high confidence, high target speed, and rapidly decreasing distance; abnormal cabin temperature events are assessed as "medium risk." Subsequently, the system generates different response strategies based on the risk level: for high-risk external events, the system immediately triggers a Level 1 response, including issuing an audible and visual alarm on the bridge, displaying "Obstacle approaching to the right front, please avoid immediately," and automatically generating a collision avoidance control command, instructing the autopilot system to deflect the course 3 degrees to the left while simultaneously reducing speed to a safe level; for medium-risk internal events, the system triggers a Level 2 warning, displaying "Starboard cabin temperature rising, it is recommended to check equipment status" on the monitoring interface, and instructing the ventilation system to open forced ventilation in the cabin. At this time, the system continues to monitor event changes; if the external target gradually moves away or the cabin temperature returns to normal, the risk level is automatically reduced and the corresponding control command is released. Conversely, if the temperature inside the cabin continues to rise and the smoke concentration increases sharply, the system will automatically upgrade the event level to "high risk," activate the automatic fire suppression system, and report to the shore-based monitoring center.
[0099] Throughout the process, the system updates the status of multiple events in real time and records complete time-series data, ensuring that the generation, fusion, and processing of events form a self-consistent closed loop. Ultimately, the system's output of multi-event collaborative perception results includes not only all risk levels and corresponding control actions in the current environment, but also event development trend analysis results, which can provide a basis for subsequent safety assessments, maintenance, and accident tracing.
[0100] Furthermore, the generation of multi-event collaborative perception results based on the collaborative event set includes risk levels and corresponding early warning or control instructions, including: For each event in the collaborative event set, risk characteristics are extracted according to event type, event priority index and time range, and event risk factors are calculated. The event risk factors are obtained by weighting event confidence, time urgency and spatial hazard range. Among them, time urgency is calculated based on the event intensity growth rate in the most recent time slice, and spatial hazard range is determined by the overlapping area between the event impact area and the critical safety boundary of the monitoring platform. The obtained event risk factors are input into the risk comprehensive assessment process. In the risk comprehensive assessment process, the global risk degree is calculated according to the spatiotemporal coupling relationship between events. When multiple events overlap in time and are spatially related, they are regarded as a composite risk scenario. The superposition correction is made according to the impact propagation weight between events to reflect the cumulative effect of risk. The collaborative event set is classified based on the global risk level. Events with a global risk level exceeding a set threshold are marked as high-risk events, and events with a global risk level below the set threshold but with an upward trend are marked as potential risk events. Risk labeling results containing risk level, time warning window and spatial impact area are generated. Based on the risk labeling results, early warning or control instructions are generated. When a high-risk event is located within the safety boundary, an immediate early warning signal is generated and early warning information including the event location, risk type, and emergency handling suggestions is output. When the time warning window of a potential risk event is shorter than the preset response period, control instructions are automatically triggered to adjust monitoring parameters or perform local avoidance behavior. All early warning information and control commands are integrated into a multi-event collaborative perception result output, which includes risk level, event label, spatial location, time range and corresponding response strategy, to realize real-time risk assessment and intelligent decision linkage of multiple events in complex dynamic environments.
[0101] In this embodiment, in order to achieve risk quantification assessment and proactive response control for multiple events, the system uses a set of collaborative events as the input data source, extracts event features step by step, calculates risk indicators, performs global risk assessment, and generates executable early warning or control instructions.
[0102] At the start of processing, the system first extracts risk features from each event in the collaborative event set. The "collaborative event set" refers to the event dataset obtained after the previous stage of temporal consistency analysis, conflict resolution, and prioritization. Each event includes information such as event type, spatial location, time range, and event priority index. The system establishes a corresponding risk feature template for each event type (e.g., abnormal temperature, oil and gas leak, structural vibration, approach of external obstacles) to identify its potential threat level in the current scenario. During risk feature extraction, the system focuses on calculating the "event risk factor," a core indicator used to quantitatively represent the risk intensity of a single event. The event risk factor is obtained by weighting three components: event confidence, time urgency, and spatial hazard range. Event confidence reflects the reliability of event identification and is output by a multimodal neural network in the early identification stage. Time urgency measures the rate of change of the event, specifically by calculating the intensity growth rate of the event over the three most recent consecutive time slices. For example, if the oil and gas concentration rises from 20 ppm to 40 ppm and then to 80 ppm, the system calculates its growth rate as an average of +100% per time slice, indicating high time urgency. Spatial hazard range represents the coverage ratio of the event's affected area relative to the safety boundary of the monitoring platform. For example, if the oil and gas leak area overlaps with the monitoring area by 50%, the hazard range value is 0.5. The system obtains the final event risk factor value by weighting the above three indicators with fixed weights (e.g., 0.4, 0.3, 0.3), which is used for subsequent evaluation.
[0103] After obtaining the risk factors for each event, the system inputs them into a comprehensive risk assessment process. In this process, the system first analyzes the spatiotemporal coupling relationships between events. Spatiotemporal coupling describes the interactive effects of multiple events in time and space. For example, temperature anomalies and smoke generation events often have a causal relationship. The system quantifies the degree of coupling by calculating the overlap rate in the time dimension and the ratio of their spatial overlap area. If the time overlap rate of two events exceeds 60% and the spatial overlap area ratio exceeds 40%, the system classifies it as a "compound risk scenario." For such compound risk scenarios, the system introduces an "impact propagation weight" for risk superposition correction, which involves multiplying the risk factor of the main event by the propagation weight and superimposing it onto the risk value of related events. For example, when the risk factor of an oil and gas leak event is 0.8 and the propagation weight is 0.6, its correction amount for adjacent smoke events is 0.48, thus reflecting the cumulative effect and cascading impact of the risk. Through coupling analysis and propagation correction of all related events, the system obtains the "global risk level" of the current environment, which reflects the overall risk level of all events under the intertwined effects of time and space.
[0104] After obtaining the global risk level, the system classifies the set of collaborative events to distinguish different levels of risk. The classification process is based on a preset threshold model, for example, setting 0.7 as the high-risk threshold and 0.4 as the potential-risk threshold. When the global risk level of an event is higher than 0.7, the event is marked as a high-risk event; when the global risk level is lower than 0.7 but shows a continuous upward trend over the past three time slices, it is marked as a potential-risk event; if it is lower than 0.4 and the trend is stable, it is classified as a low-risk event. Simultaneously with classification, the system generates a "risk labeling result" for each event, including the risk level, time warning window, and spatial impact area. The time warning window refers to the time period from the current time slice of the event until it is expected to reach the high-risk threshold. For example, if the smoke concentration is expected to reach a dangerous level in 3 minutes under the current upward trend, the time warning window is 3 minutes. The spatial impact area is determined by superimposing the event location and diffusion rate, providing the system with a basis for predicting the risk diffusion range.
[0105] After risk labeling is completed, the system enters the early warning and control command generation phase. If an event is marked as high-risk and located within the safety boundary (i.e., the shortest distance between the event center and the platform's critical structure is less than the safety threshold), the system immediately generates an early warning signal. The early warning information includes the event location, risk type, risk level, and emergency handling recommendations, such as "High-concentration oil and gas leak detected in the starboard near-shore detection area; please immediately initiate the isolation valve closure procedure." Simultaneously, for potential risk events, the system determines whether to trigger preventative control based on the time-based early warning window. When the early warning window is less than a preset response period (e.g., 60 seconds), the system automatically outputs control commands, such as "Reduce the sensor sampling interval" or "Adjust the camera angle to focus on the abnormal area." These control commands directly act on the execution layer of the monitoring equipment, achieving automatic response and proactive avoidance.
[0106] Ultimately, the system integrates all generated early warning information and control commands in a structured manner to form a complete "multi-event collaborative perception result." This result is output in the form of a data table, with each record containing an event label, risk level, spatial location, time range, and corresponding response strategy. This output is not only provided for display on the monitoring terminal but can also be invoked by external decision-making systems to coordinate and execute further safety protection measures, such as audible and visual alarms, unattended automatic emergency handling, or dynamic adjustment of the navigation path.
[0107] A second embodiment of this application provides an electronic device, the electronic device comprising: processor; The memory is used to store a program, which, when read and executed by the processor, executes the dynamic environment and multi-event collaborative perception method based on a multimodal neural network provided in the first embodiment of this application.
[0108] The third embodiment of this application provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it executes a dynamic environment and multi-event collaborative perception method based on a multimodal neural network provided in the first embodiment of this application.
[0109] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A dynamic environment and multi-event collaborative perception method based on multimodal neural networks, characterized in that, include: Multimodal data of the dynamic environment is collected based on multi-source sensing devices, and the multimodal data is synchronized in time and registered in space to generate a multimodal data sequence under a unified reference coordinate according to time slices. The multimodal data sequence is input into a multimodal deep neural network, the feature information of each modality is extracted, the modal similarity and confidence change rate of adjacent time slices are calculated, and the feature weights of each modality are corrected according to the confidence change rate to obtain the corrected feature information. Using the corrected feature information, external targets and internal events are identified respectively, generating corresponding candidate results and their confidence information, thus forming a candidate event set; Based on the differences in confidence information and temporal consistency in the candidate event set, weighted fusion and conflict resolution are performed to establish event priority relationships and obtain a collaborative event set containing event labels, spatial locations, and temporal ranges; Multi-event collaborative perception results are generated based on a set of collaborative events. These results include risk levels and corresponding early warning or control instructions.
2. The dynamic environment and multi-event collaborative perception method based on multimodal neural networks according to claim 1, characterized in that, The process of acquiring multimodal data of the dynamic environment based on multi-source sensing devices, performing time synchronization and spatial registration on the multimodal data, and generating a multimodal data sequence under a unified reference coordinate system according to time slices includes: Environmental perception data is collected by multi-source sensing devices, and the sampling period of each sensing device is adjusted according to the environmental complexity parameter. The environmental complexity parameter is determined by a combination of the rate of change of light intensity, the amplitude of change of radar echo density, and the degree of gas concentration disturbance. When the environmental complexity parameter exceeds a preset threshold, the sampling frequency of each sensing device is increased to ensure the time continuity and data integrity in dynamic environments. The original data output by each sensor device is timestamped and compared. Using the unified reference clock of the main control system as a reference, the time delay of each sensor device is calculated and a time compensation curve is generated. The subsequent sampling triggering sequence of each sensor device is corrected according to the time compensation curve, so that the data of different modes correspond to the same sampling time under the unified time reference, thereby realizing the time synchronization of multimodal data. Spatial coordinate alignment is performed on the multimodal data that has been synchronized in time. Based on the difference in the response positions of multiple fixed reference points in the fields of view of different sensors, the spatial orientation and installation offset of each sensor are gradually corrected until the spatial position deviation of the same reference point in the unified coordinate system is lower than the set error range under the observation of all sensors, thereby ensuring the spatial consistency of multimodal data under the unified reference coordinate system. Based on the data that has been synchronized in time and registered in space, the multimodal data is organized into a continuous spatiotemporal sequence according to time slices. The length of the time slice is automatically determined according to the scene dynamic parameters, which are jointly determined by the rate of change of the target's motion speed and the intensity of external environmental disturbances. When the scene dynamic parameters exceed a preset threshold, the length of the time slice is automatically shortened to enhance the temporal resolution. Each data unit in each time slice contains unified reference coordinates, modality identifiers, data quality labels, and initial confidence values, thereby forming a standardized multimodal spatiotemporal data sequence.
3. The dynamic environment and multi-event collaborative perception method based on multimodal neural networks according to claim 1, characterized in that, The process involves inputting the multimodal data sequence into a multimodal deep neural network, extracting feature information for each modality, calculating the modal similarity and confidence change rate between adjacent time slices, and adjusting the feature weights of each modality based on the confidence change rate to obtain adjusted feature information, including: The multimodal deep neural network is composed of an input layer, a feature extraction layer, a time alignment layer, a confidence evaluation layer, and a feature fusion layer connected in sequence. The input layer is used to receive multimodal data sequences that have undergone time synchronization and spatial registration and to perform standardization processing to ensure that different modalities are consistent in resolution, value range and coordinate scale. The feature extraction layer includes parallel branches corresponding to the visible light mode, infrared mode, laser point cloud mode and radar mode, respectively. While sharing the main structure, each branch retains an independent parameter set to extract the texture, temperature gradient, spatial density and echo energy features of its own mode and output the corresponding feature vector. The time alignment layer indexes and compares the multimodal feature vectors of consecutive time slices, calculates the feature differences of the same spatial region in adjacent time slices and generates a time consistency index to reflect the stability of each mode in the time series. The confidence assessment layer calculates the confidence change rate based on the time consistency index and signal quality parameters, including signal integrity, noise ratio, and sampling stability. When time consistency decreases and signal quality deteriorates, the confidence change rate is negative; when time consistency is good and signal quality improves, the confidence change rate is positive. The feature fusion layer dynamically adjusts the feature weights of each mode based on the confidence change rate. When the confidence of the visible light mode decreases under low light conditions, its feature weight is reduced and the feature weight of the infrared mode is increased. When the confidence of the infrared mode decreases under high temperature background, the feature weights of the radar mode and the laser point cloud mode are increased. The corrected features are then normalized and superimposed to generate a comprehensive feature vector with corrected confidence. The comprehensive feature vector remains continuous in the time dimension and aligned in the spatial dimension, and serves as the input feature for subsequent external target recognition and internal event determination.
4. The dynamic environment and multi-event collaborative perception method based on multimodal neural networks according to claim 1, characterized in that, The process of using modified feature information to identify external targets and internal events respectively, generating corresponding candidate results and their confidence information, and forming a candidate event set includes: Based on confidence-corrected multimodal feature information, a spatiotemporal fusion feature set is constructed under a unified reference coordinate system. Feature information from visible light, infrared, lidar, and radar modes is spatially mapped to form a multidimensional fusion feature block that can reflect scene texture, temperature distribution, spatial morphology, and energy reflection characteristics. Spatial partitioning and region similarity aggregation are performed on the multidimensional fusion feature block for each time slice to obtain a set of candidate target regions. Each candidate target region includes a feature center point, boundary contour, and temporal continuity label. Based on the temporal continuity label of the candidate target area, the consistency of the target trajectory between adjacent time slices is calculated, and external target candidates are output. The external target candidates include the direction of movement, the rate of change of velocity, and the position confidence interval. When the temporal consistency index is higher than the preset threshold and the trajectory is continuous, the external target type is determined as a valid candidate target. The external target candidates are compared and analyzed with the environmental state information to identify floating objects on the water surface and ship obstacles. Using the infrared temperature distribution, gas concentration features, and acoustic energy features contained in the fusion feature set as input, the internal event identification process is executed. The type and location of internal abnormal events are determined by deviation analysis from the normal operation state template library. Internal events include temperature abnormalities, smoke generation, oil and gas leaks, and equipment vibration exceeding limits. Candidate internal events are generated based on the deviation amplitude and duration. The joint confidence level calculation is performed on external target candidates and internal event candidates, taking into account feature completeness, temporal continuity, signal stability, and spatial correlation. Each event is assigned a confidence level, and event coupling relationships are established based on temporal overlap and spatial correlation. Events with high correlation are integrated into composite candidate events, and a candidate event set is output. Each item in the candidate event set includes event type, spatial location, time range, and confidence level, which are used in subsequent collaborative fusion and decision analysis stages to achieve integrated identification and dynamic management of external targets and internal events driven by multimodal features.
5. The dynamic environment and multi-event collaborative perception method based on multimodal neural networks according to claim 1, characterized in that, The step involves performing weighted fusion and conflict resolution based on the differences in confidence information and temporal consistency within the candidate event set, establishing event priority relationships, and obtaining a collaborative event set containing event labels, spatial locations, and temporal ranges, including: The confidence information of the candidate event set is normalized and mapped to obtain the relative confidence ratio. The event type, spatial location and time slice index of the candidate event are used as the retrieval key to perform spatial aggregation within the same time slice to generate local event nodes. The local event nodes carry the weighted spatial center position, boundary contour, relative confidence ratio and the initial time range formed by its source time slice index. For local event nodes with a spatial overlap rate reaching a preset threshold and consistent event type within adjacent time slices, time extension fusion is performed. Multiple local event nodes are merged into intermediate event units according to time continuity rules. The time range of the intermediate event unit is defined by the earliest time slice index and the latest time slice index and includes all time slices in between. The relative confidence ratio is updated to a weighted average of the relative confidence ratios of the participating nodes, and its spatial center position and boundary contour are updated synchronously. The time continuity rules include: time slice indices that are adjacent or whose interval does not exceed the preset time slice gap are considered continuous. After obtaining multiple intermediate event units, conflict resolution is performed based on the overlap between spatial location and temporal range. When intermediate event units overlap spatially but do not overlap temporally, the intermediate event unit with the higher relative confidence ratio is retained. When intermediate event units overlap both spatially and temporally but have different event types, the dominant intermediate event unit is determined based on the order of occurrence, duration, and rate of change of feature intensity, and its derivative relationship with other intermediate event units is recorded. Intermediate event units that have undergone conflict resolution are semantically named and solidified as events according to event type. Events inherit the time range, spatial location and relative confidence ratio of the corresponding intermediate event units. The event priority index is calculated by weighting the event importance coefficient, relative confidence ratio and time urgency. The time urgency is determined by the growth trend of feature intensity in the most recent time slice and the length of the time range. Events are sorted from highest to lowest according to their event priority index and output as a collaborative event set. Each event in the collaborative event set contains an event label, spatial location, time range, and priority identifier.
6. The dynamic environment and multi-event collaborative perception method based on multimodal neural networks according to claim 1, characterized in that, The multi-event collaborative sensing result is generated based on the collaborative event set. The multi-event collaborative sensing result includes a risk level and a corresponding early warning or control instruction, including: For each event in the collaborative event set, risk characteristics are extracted according to event type, event priority index and time range, and event risk factors are calculated. The event risk factors are obtained by weighting event confidence, time urgency and spatial hazard range. Among them, time urgency is calculated based on the event intensity growth rate in the most recent time slice, and spatial hazard range is determined by the overlapping area between the event impact area and the critical safety boundary of the monitoring platform. The obtained event risk factors are input into the risk comprehensive assessment process. In the risk comprehensive assessment process, the global risk degree is calculated according to the spatiotemporal coupling relationship between events. When multiple events overlap in time and are spatially related, they are regarded as a composite risk scenario. The superposition correction is made according to the impact propagation weight between events to reflect the cumulative effect of risk. The collaborative event set is classified based on the global risk level. Events with a global risk level exceeding a set threshold are marked as high-risk events, and events with a global risk level below the set threshold but with an upward trend are marked as potential risk events. Risk labeling results containing risk level, time warning window and spatial impact area are generated. Based on the risk labeling results, early warning or control instructions are generated. When a high-risk event is located within the safety boundary, an immediate early warning signal is generated and early warning information including the event location, risk type, and emergency handling suggestions is output. When the time warning window of a potential risk event is shorter than the preset response period, control instructions are automatically triggered to adjust monitoring parameters or perform local avoidance behavior. All early warning information and control commands are integrated into a multi-event collaborative perception result output, which includes risk level, event label, spatial location, time range and corresponding response strategy, to realize real-time risk assessment and intelligent decision linkage of multiple events in complex dynamic environments.
Citation Information
Patent Citations
Accident early warning analysis method and system based on driving data
CN119920076A
Road area environment safety evaluation method and system in highway engineering construction period
CN120355239A