Cross-modal data real-time acquisition and analysis system based on multi-sensor fusion

By adopting high-precision clock synchronization technology and cross-modal feature mapping model in multi-sensor fusion system, the problems of low synchronization accuracy and high system complexity in real-time acquisition and analysis of cross-modal data are solved, and high-precision and real-time data fusion analysis are achieved.

CN120217299APending Publication Date: 2025-06-27ANOTHER ME (BEIJING) VIRTUAL TECH DEV CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510357713.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing multi-sensor fusion system has problems such as low data synchronization accuracy, high fusion algorithm complexity, and poor system scalability in real-time acquisition and analysis of cross-modal data.

Method used

Multiple types of sensors are used to synchronize data acquisition with high-precision clock synchronization technology, and deep fusion of data is carried out through cross-modal feature mapping models. Deep learning algorithms are used to extract the features of multimodal data to realize real-time data analysis and decision-making.

Benefits of technology

It realizes high-precision synchronous acquisition of multiple modal data, improves the time consistency of data and the accuracy of fusion analysis, reduces the problems of system complexity and poor scalability, and is suitable for application scenarios with high requirements for real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217299A_ABST
    Figure CN120217299A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal data real-time acquisition and analysis system based on multi-sensor fusion, and relates to the technical field of data acquisition and analysis, comprising: acquiring cross-modal data in real time through various types of sensors; data synchronous acquisition is carried out among various types of sensors through a high-precision clock synchronization technology; preprocessing the synchronously collected original data; constructing a cross-modal feature mapping model according to the preprocessed data, mapping cross-modal data features to the same feature space, and performing deep fusion of the data; and performing real-time analysis and decision according to the fused data. Through a high-precision clock synchronization technology, synchronous acquisition of various modal data is realized, and the time consistency of the data is improved; data fusion is carried out by adopting a deep learning algorithm, so that features of multi-modal data can be effectively extracted, deep fusion of the data is realized, and the accuracy and efficiency of data fusion are improved; the system can realize real-time acquisition and analysis of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data acquisition and analysis, and more specifically, to a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion. Background Art

[0002] With the rapid development of technologies such as the Internet of Things and artificial intelligence, real-time acquisition and analysis of multi-modal data are required in various application scenarios, such as visual and radar data in autonomous driving, temperature and vibration data in industrial equipment monitoring, etc. Traditional single sensors or simple sensor combinations can no longer meet the requirements for data accuracy, real-time performance, and stability in complex scenarios. Multi-sensor fusion technology can effectively improve the accuracy and reliability of data by integrating data from different sensors. However, existing multi-sensor fusion systems still have many deficiencies in the real-time acquisition and analysis of cross-modal data, such as low data synchronization accuracy, high complexity of fusion algorithms, and poor system scalability.

[0003] Therefore, how to propose a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion to achieve high-precision synchronous acquisition and efficient fusion analysis of multi-modal data is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the present invention provides a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion to achieve high-precision synchronous acquisition and efficient fusion analysis of multi-modal data. To achieve the above object, the present invention adopts the following technical solutions:

[0005] A real-time cross-modal data acquisition and analysis system based on multi-sensor fusion, comprising:

[0006] Real-time acquisition of cross-modal data through various types of sensors;

[0007] Data synchronization acquisition between various types of sensors through high-precision clock synchronization technology;

[0008] Preprocessing the synchronously acquired original data;

[0009] Construct a cross-modal feature mapping model based on the preprocessed data, map the cross-modal data features to the same feature space, and perform deep fusion of the data;

[0010] Perform real-time analysis and decision-making based on the fused data.

[0011] Optionally, the various types of sensors include: image sensors, lidar, position sensors, speed sensors, and acceleration sensors.

[0012] Optionally, the preprocessing includes data cleaning, filtering, and normalization operations to remove invalid data and noise.

[0013] Optionally, the data synchronization acquisition between the multiple types of sensors through high-precision clock synchronization technology includes:

[0014] The master node periodically sends time frames and management frames to the sensor slave nodes. The management frame is sent after the time frame, and the start symbol of the management frame is sent at the start of the whole second.

[0015] The sensor slave node receives the data frame sent by the master node, and then determines whether the data frame is a time frame. If it is a time frame, the sensor slave node synchronizes its full time scale time.

[0016] The sensor slave node determines whether the received data frame is a management frame. If it is a management frame, the sensor slave node performs second pulse synchronization and sends a feedback message to the master node according to the type of the received management frame. Otherwise, it discards the data frame, and the sensor slave node operates independently and normally and waits for the arrival of the next frame of data.

[0017] Optionally, it further includes that the master node receives the feedback message sent by the sensor slave node, modifies and stores the corresponding parameters according to the received feedback message. If the master node does not receive the feedback message of a certain sensor slave node, it determines that the corresponding sensor slave node is faulty.

[0018] Optionally, the construction of the cross-modal feature mapping model according to the preprocessed data, mapping the cross-modal data features to the same feature space, and performing deep fusion of the data includes:

[0019] Perform an embedding operation on the preprocessed cross-modal data, use a cross-modal alignment network to map the cross-modal data to the same embedding space, and introduce a self-attention mechanism to weight the features in the embedding space, and fuse the weighted multi-modal data to obtain cross-modal fusion features.

[0020] Optionally, after using the cross-modal alignment network to map the multi-modal data to the same embedding space, it further includes calculating the covariance matrix obtained by improving the covariance descriptor, performing a compact representation of the cross-modal data, and generating the input data for fusing the multi-modal data.

[0021] Optionally, the obtaining of the cross-modal fusion features includes:

[0022] According to the data features extracted from the image sensor, lidar, position sensor, speed sensor, or acceleration sensor, use the cross-modal alignment network f CA Map the data of different modalities to the same embedding space, which is expressed as follows:

[0023] L avg (x i )=f CA (l t (x i ),l i (x i ),l a (x i ),l e (x i ),l r (x i ));

[0024] Among them, x i represents image data, radar data, position data, velocity data, or acceleration data, and f CA represents a cross-modal alignment network. L avg (x i ) represents the feature vector mapped to the same embedding space. l t , l i , l a、 l e and l r respectively represent the embedding vectors of image data, radar data, position data, velocity data, and acceleration data;

[0025] The covariance matrix obtained by improving the covariance descriptor is used to compactly represent the cross-modal data:

[0026]

[0027] The self-attention mechanism is used to weight the features in the embedding space, which is expressed as follows:

[0028] L att (x i )=(W att ·L avg (x i ))·U E ;

[0029] Among them, W att represents the self-attention weight matrix obtained through training, U E represents the covariance matrix, and L att represents the feature vector after being processed by the self-attention mechanism;

[0030] The weighted feature vectors of the image data, radar data, position data, velocity data, and acceleration data after being processed by the self-attention mechanism are fused to generate a comprehensive multi-modal feature representation, which is expressed as follows: l finally (x i )=g finally (l att(x i ));

[0031] Among them, l finally (x i ) represents the finally fused multi-modal feature vector, and g finally represents the fusion function.

[0032] Optionally, the real-time analysis and decision-making based on the fused data includes using the random forest algorithm for obstacle avoidance:

[0033] Taking the obstacle avoidance decision as the dependent variable and the fused data as the independent variable, and inputting them into the random forest model;

[0034] Setting the number of random forest trees K, the training ratio, using MSE as the node splitting evaluation criterion, setting it to sampling with replacement, setting the maximum depth of the tree to a, the maximum number of leaf nodes to b, and the threshold of node division impurity to c;

[0035] After running the predicted random forest model, adjust the parameters to retrain the random forest model according to whether the observation evaluation criterion is reasonable;

[0036] Input the fused data to be predicted into the random forest model to obtain the obstacle avoidance decision prediction result.

[0037] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a cross-modal data real-time acquisition and analysis system based on multi-sensor fusion, having the following beneficial effects:

[0038] The present invention proposes a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion, including: real-time acquisition of cross-modal data through various types of sensors; data synchronization acquisition among various types of sensors through high-precision clock synchronization technology; preprocessing of the synchronously acquired raw data; constructing a cross-modal feature mapping model based on the preprocessed data to map the cross-modal data features to the same feature space for in-depth data fusion; and performing real-time analysis and decision-making based on the fused data. Through the high-precision clock synchronization technology, the present invention realizes the synchronous acquisition of multi-modal data, improves the time consistency of the data, and provides a reliable data basis for subsequent data fusion analysis; adopts deep learning algorithms for data fusion, can effectively extract the features of multi-modal data, and realizes in-depth data fusion, improving the accuracy and efficiency of data fusion; the system can realize real-time data acquisition and analysis, with a fast response speed, and is applicable to application scenarios with high requirements for real-time performance; after feature extraction, the raw data effectively compresses the data volume, and collaboratively shares the data in the form of features, enriching the data sources of each unit, greatly reducing the bandwidth requirements of the channel, further reducing the delay, and further improving the accuracy, meeting the requirements of vehicle obstacle avoidance for the real-time performance of intelligent algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.

[0040] Figure 1 It is a structural framework diagram of a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] An embodiment of the present invention discloses a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion, as Figure 1 shown, including:

[0043] Real-time acquisition of cross-modal data through various types of sensors;

[0044] Data synchronization acquisition is carried out among multiple types of sensors through high-precision clock synchronization technology;

[0045] Preprocess the raw data collected synchronously;

[0046] Construct a cross-modal feature mapping model based on the preprocessed data, map the cross-modal data features to the same feature space, and perform in-depth data fusion;

[0047] Perform real-time analysis and decision-making based on the fused data.

[0048] Further, the multiple types of sensors include: image sensors, lidar, position sensors, speed sensors, and acceleration sensors.

[0049] Further, the preprocessing includes data cleaning, filtering, and normalization operations to remove invalid data and noise.

[0050] Further, the data synchronization acquisition among the multiple types of sensors through high-precision clock synchronization technology includes:

[0051] The master node circularly sends time frames and management frames to the sensor slave nodes. The management frame is sent after the time frame, and the starting code element of the management frame starts to be sent at the start of the whole second;

[0052] The sensor slave node receives the data frame sent by the master node, and then judges whether the data frame is a time frame. If it is a time frame, the sensor slave node synchronizes its full time scale time;

[0053] The sensor slave node judges whether the received data frame is a management frame. If it is a management frame, the sensor slave node performs second pulse synchronization and sends a return message to the master node according to the type of the received management frame. Otherwise, it discards the data frame, and the sensor slave node operates independently and normally and waits for the arrival of the next frame of data.

[0054] Further, it also includes that the master node receives the return message sent by the sensor slave node, modifies and stores the corresponding parameters according to the received return message. If the master node does not receive the return message of a certain sensor slave node, it determines that the corresponding sensor slave node is faulty.

[0055] In the specific implementation, before the sensor slave node synchronizes its full time scale time, it includes:

[0056] First, compare the time quality of the received time frame with its own time quality. If the time quality level of the time frame is higher, the sensor secondary node synchronizes its full-time scale time according to the full-time scale time information of the received time frame. If the time quality level of the sensor secondary node itself is higher, discard the time frame, and the sensor secondary node operates independently and normally and waits for the arrival of the next frame of data.

[0057] The sensor secondary node determines whether the data frame is a management frame based on whether there are three consecutive P bits starting from the preamble of the received data frame. If there are three consecutive P bits, namely P0, Pr, and Pt, it is determined as a management frame. Then, compare the time quality of the received management frame with its own time quality. If the time quality level of the management frame is higher, synchronize the self-second pulse of the sensor secondary node. If the time quality of the sensor secondary node itself is higher, do not synchronize. Then, determine the type of the management frame according to the frame type information segment of the data frame. If the frame type is 0x1 to 0x2, it is determined as a sensor secondary node configuration frame. Then, the sensor secondary node sends a feedback message including the frame type, the address of the sensor secondary node, and the check information to the master node.

[0058] Further, constructing a cross-modal feature mapping model according to the preprocessed data, mapping the cross-modal data features to the same feature space, and performing deep fusion of the data includes:

[0059] Perform an embedding operation on the preprocessed cross-modal data, use a cross-modal alignment network to map the cross-modal data to the same embedding space, and introduce a self-attention mechanism to weight the features in the embedding space, and fuse the weighted multi-modal data to obtain cross-modal fusion features.

[0060] Specifically, for speed data and acceleration data, use the BERT model for vector embedding; for image data and radar data, use the ResNet50 model for vector embedding; for position data, use the VGGish model for vector embedding.

[0061] Further, after using the cross-modal alignment network to map the multi-modal data to the same embedding space, it also includes compactly representing the cross-modal data by calculating the covariance matrix obtained by improving the covariance descriptor, and generating the input data for fusing the multi-modal data.

[0062] Further, obtaining the cross-modal fusion features includes:

[0063] According to the data features extracted from the image sensor, lidar, position sensor, speed sensor, or acceleration sensor, use the cross-modal alignment network f CA Map the data of different modalities to the same embedding space, which is expressed as follows:

[0064] L avg (x i ) = f CA (l t (x i ), l i (x i ), l a (x i ), l e (x i ), l r (x i ));

[0065] Among them, x i represents image data, radar data, position data, velocity data, or acceleration data, f CA represents a cross-modal alignment network, and L avg (x i ) represents the feature vector mapped to the same embedding space, and l t , l i , l a , l e , and l r respectively represent the embedding vectors of image data, radar data, position data, velocity data, and acceleration data;

[0066] The covariance matrix calculated by improving the covariance descriptor is used to compactly represent the cross-modal data:

[0067] Among them, β represents the point set average.

[0068] The self-attention mechanism is used to weight the features in the embedding space, which is expressed as follows:

[0069] L att (x i ) = (W att ·L avg (x i ))·U E ;

[0070] Among them, W att represents the self-attention weight matrix obtained through training, U E represents the covariance matrix, and L att represents the feature vector after being processed by the self-attention mechanism;

[0071] The weighted feature vectors of the image data, radar data, position data, velocity data, and acceleration data after being processed by the self-attention mechanism are fused to generate a comprehensive multi-modal feature representation, which is expressed as follows: l finally (xi ) = g finally (l att (x i ));

[0072] Among them, l finally (x i ) represents the finally fused multi-modal feature vector, and g finally represents the fusion function.

[0073] In the specific implementation manner, it is also necessary to optimize the loss of the cross-modal feature mapping model:

[0074] Adding a regularization term to improve the generalization ability and robustness of the model is expressed as follows:

[0075]

[0076] Among them, R1(W) represents the regularization term of the multi-modal data, ∥ Watt ∥ F represents the Frobenius norm, L o (x i ) represents the original feature vector, and λ1 and λ2 are the weight coefficients of the regularization term respectively;

[0077] Then the overall loss function L1 of the multi-modal data is expressed as follows:

[0078] L1 = l align + R1(W);

[0079] Among them, l align represents the alignment loss, which is used to measure the effect of feature alignment between different modalities.

[0080] In the specific implementation manner, the covariance matrix obtained by improving the covariance descriptor is used to compactly represent the cross-modal data, where:

[0081] The improved covariance descriptor is obtained in the following way: extending the standard covariance calculation from the Euclidean space to the complete inner product space of the symmetric positive definite manifold; characterizing the covariance structure according to the arccosine kernel that satisfies the target condition, and performing mean centering processing based on the positive definite manifold matrix; where, the target condition is: any semi-positive definite function can be used as a kernel function; obtaining the parameters of different-order arccosine kernels by using supervised kernel alignment learning, and obtaining the improved covariance descriptor through different-order arccosine kernels and their corresponding parameters.

[0082] In the specific implementation manner, the first key step in the compact representation is to convert the modal data into a feature description form. Theoretically, the research results of existing feature descriptors are very rich and diverse. Representative ones include: Harris, Sobel, Prewitt, Canny, Brisk, LBP, SIFT, SURF, DoG, LoG, and HoG, etc. These methods have their own unique advantages in carrying spatial domain information, but they are not suitable for representing multi-modal heterogeneous data because such data has obvious temporal continuity. The covariance matrix calculated using covariance descriptors (CovDs) not only has the ability to describe individual data but also is suitable for representing data sets including sequences. In addition, this algorithm can also alleviate the heterogeneous differences of multi-modal data to a certain extent.

[0083] Specifically, in the CovDs theory, for the object I to be processed, the feature extraction process can be expressed as:

[0084] F(x,y) = δ(I,x,y);

[0085] Here, δ can be understood as any mapping relationship (i.e., the feature extraction descriptor). When the representation region is satisfied, the covariance matrix of the feature point set within the region can be calculated in the following manner:

[0086] where β represents the average value of the point set.

[0087] Furthermore, the real-time analysis and decision-making based on the fused data include obstacle avoidance using the random forest algorithm:

[0088] Taking the obstacle avoidance decision as the dependent variable and the fused data as the independent variable, input them into the random forest model;

[0089] Set the number of random forest trees K, the training ratio, use MSE as the node splitting evaluation criterion, set it to sampling with replacement, set the maximum depth of the tree to a, the maximum number of leaf nodes to b, and the threshold of node division impurity to c;

[0090] After running the predicted random forest model, adjust the parameters to retrain the random forest model according to whether the observation evaluation criterion is reasonable;

[0091] Input the fused data to be predicted into the random forest model to obtain the predicted result of the obstacle avoidance decision.

[0092] In the specific implementation manner, the real-time analysis and decision-making including obstacle avoidance using the random forest algorithm specifically include:

[0093] S1: Provide a random forest F including K decision trees, where k = 1, 2... N; provide several strings B i , where i = 1, 2... N; and perform the following training steps:

[0094] S1.1: Given several sample data tables A i , where i = 1, 2... N;

[0095] S1.2: Randomly select a group of X p tuples from table A p , and randomly select a group of X q tuples from table A p that may match the said table A q . Pair the said X p with the said X q to form a sample S, where: p ∈ i, q ∈ i, p ≠ q;

[0096] S1.3: Check the schema of the said table A p and the said table A q , create a set of features, and use the features to convert the tuple pairs in the sample S into feature vectors;

[0097] S1.4: Use the feature vectors in S1.3 to train the random forest F;

[0098] S2: Perform a pruning step, and the pruning step includes:

[0099] S2.1: Extract m decision trees Q1, Q2... Q m from the k decision trees, and use each of the Q1, Q2... Q m to execute each of the said strings B i , to obtain outputs C1, C2... C m , where m is the minimum number of decision trees required for correct parsing, and the correct parsing means that the random forest F correctly parses the said string B i as an entity;

[0100] S2.2: Establish a set I = C1 ∩ C2 ∩... ∩ C m ;

[0101] S3: Perform a verification step, and the verification step includes:

[0102] S3.1: Establish a set J = (C1 ∪ C2 ∪... ∪ C m ) \ (C1 ∩ C2 ∩... ∩ C m );

[0103] S3.2: Extract n decision trees G1, G2... G from the random forest F n , and use the G1, G2... G n to execute the set J to generate sets K1, K2... K n , and among them

[0104]

[0105] S4: The random forest F outputs the entity resolution result as I ∪ K1 ∪ K2 ∪... ∪ K n ;

[0106] S5: Perform real-time obstacle avoidance according to the output entity resolution result.

[0107] In the specific implementation, taking the autonomous driving scenario as an example, the system can collect multi-modal data such as image data, lidar data, position data, acceleration data, and speed data around the vehicle in real time. Through the high-precision synchronous collection of the data collection module, the consistency of each modal data in time is ensured. The data preprocessing module performs operations such as cleaning and filtering on the collected raw data to remove invalid data and noise. The data fusion module uses deep learning algorithms to extract and fuse features from the preprocessed data, constructs a cross-modal feature mapping model, maps the data features of different modalities to the same feature space, and realizes the deep fusion of data. The analysis and decision-making module recognizes pedestrians, vehicles, obstacles, etc. in the road environment in real time according to the fused data, and makes corresponding driving decisions according to the analysis results, such as decelerating, avoiding, changing lanes, etc. The system of the present invention can effectively improve the perception ability and decision-making accuracy of autonomous driving vehicles in complex environments, and has important application value and broad market prospects.

[0108] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0109] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A real-time cross-modal data acquisition and analysis system based on multi-sensor fusion, characterized in that: include: Collect cross-modal data in real time through multiple types of sensors; Data is collected synchronously between various types of sensors through high-precision clock synchronization technology; Preprocess the raw data collected synchronously; A cross-modal feature mapping model is constructed based on the preprocessed data, and the cross-modal data features are mapped to the same feature space for deep data fusion. Conduct real-time analysis and decision-making based on the integrated data.

2. According to claim 1, a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion is characterized in that: The multiple types of sensors include: image sensors, laser radars, position sensors, speed sensors and acceleration sensors.

3. According to claim 1, a real-time cross-modal data acquisition and analysis system based on multi-sensor fusion is characterized in that: The preprocessing includes data cleaning, filtering, and normalization operations to remove invalid data and noise.

4. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 1 is characterized in that: The multiple types of sensors use high-precision clock synchronization technology to synchronously collect data, including: The master node sends time frames and management frames to the sensor secondary nodes in a cycle. The management frame is sent after the time frame, and the starting code element of the management frame is sent at the beginning of the whole second. The sensor secondary node receives the data frame sent by the master node, and then determines whether the data frame is a time frame. If it is a time frame, the sensor secondary node synchronizes its own full time scale time; The sensor secondary node determines whether the received data frame is a management frame. If it is a management frame, the sensor secondary node performs second pulse synchronization and sends a feedback message to the main node according to the type of management frame received. Otherwise, the data frame is discarded and the sensor secondary node operates independently and normally and waits for the arrival of the next frame of data.

5. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 4 is characterized in that: It also includes that the main node receives the feedback message sent by the sensor secondary node, and modifies and stores the corresponding parameters according to the received feedback message. If the main node does not receive the feedback message from a certain sensor secondary node, it determines that the corresponding sensor secondary node is faulty.

6. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 1 is characterized in that: The step of constructing a cross-modal feature mapping model based on the preprocessed data, mapping the cross-modal data features to the same feature space, and performing deep fusion of the data includes: The preprocessed cross-modal data is embedded, and the cross-modal alignment network is used to map the cross-modal data to the same embedding space. The self-attention mechanism is introduced to weight the features in the embedding space, and the weighted multimodal data is fused to obtain cross-modal fusion features.

7. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 6 is characterized in that: After using the cross-modal alignment network to map the multimodal data to the same embedding space, the method also includes compactly representing the cross-modal data using the covariance matrix calculated by improving the covariance descriptor to generate input data for fusion of the multimodal data.

8. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 7 is characterized in that: The cross-modal fusion features obtained include: Use a cross-modal alignment network f based on the data features extracted from image sensors, lidar, position sensors, velocity sensors, or accelerometers CA Mapping data of different modalities to the same embedding space is expressed as follows: L avg (x i )=f CA (l t (x i ),l i (x i ),l a (x i ),l e (x i ),l r (x i )); Among them, x i represents image data or radar data or position data or velocity data or acceleration data, f CA represents the cross-modal alignment network, L avg (x i ) represents the feature vector mapped to the same embedding space, l t , l i , l a、 l e and l r The embedding vectors representing image data, radar data, position data, velocity data, and acceleration data respectively; The cross-modal data is compactly represented by improving the covariance matrix calculated by the covariance descriptor: The self-attention mechanism is used to weight the features in the embedding space, which can be expressed as follows: L att (x i )=(W att ·L avg (x i ))·U E ; Among them, W att represents the self-attention weight matrix obtained through training, U E represents the covariance matrix, L att Represents the feature vector after processing by the self-attention mechanism; The weighted feature vectors of image data, radar data, position data, velocity data, and acceleration data processed by the self-attention mechanism are fused to generate a comprehensive multimodal feature representation, which is expressed as follows: finally (x i ) = g finally (l att (x i )); Among them, l finally (x i ) represents the final fused multimodal feature vector, g finally Represents the fusion function.

9. The real-time cross-modal data acquisition and analysis system based on multi-sensor fusion according to claim 1, characterized in that: The real-time analysis and decision-making based on the fused data includes using the random forest algorithm for obstacle avoidance: The obstacle avoidance decision is used as the dependent variable and the fused data is used as the independent variable, which is input into the random forest model; Set the number of random forest trees K, the training ratio, use MSE as the node splitting evaluation criterion, set sampling with replacement, set the maximum depth of the tree to a, set the maximum number of leaf nodes to b, and set the node partition impurity threshold to c; After the predicted random forest model is run, the parameters are adjusted and the random forest model is retrained according to whether the observed evaluation criteria are reasonable; The fused data to be predicted is input into the random forest model to obtain the obstacle avoidance decision prediction result.

Citation Information

Cited By

  • Landslide prediction method based on improved random forest in underground engineering scene

    CN119862488A

  • Epimedium koreanum planting area identification method based on multi-source remote sensing image fusion

    CN120544053A

  • Multi-mode adaptive industrial instrument intelligent control system based on deep learning

    CN120630717A

  • Distributed environment multi-mode intelligent monitoring system

    CN121309635A

  • A distributed environmental multimodal intelligent monitoring system

    CN121309635B