Sensor fusion and incremental learning based detection method and system in adverse weather

CN117333846BActive Publication Date: 2026-09-29UNIV OF SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311474818.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-03
Publication Date
2026-09-29
Estimated Expiration
2043-11-03

AI Technical Summary

Technical Problem

但是该专利文献并没有将传感器与恶劣天气结合,无法提高精度和鲁棒性

Benefits of technology

[0078]1)本发明提出了一种较完善的传感器融合和增量学习技术,即通过时空对齐规则处理多传感器数据,再按照数据级融合规则进行融合,之后用增量学习的规则对数据进行存储训练,最终分析得到恶劣天气下的目标检测结果,不仅提高了恶劣天气下的目标检测精度,还具有强鲁棒性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117333846B_ABST
    Figure CN117333846B_ABST
Patent Text Reader

Abstract

The application discloses a detection method and system based on sensor fusion and incremental learning under severe weather, comprising the following steps: a radar and a visual sensor respectively collect data of a driving environment to obtain driving environment information; the two kinds of sensor data obtained are processed according to a space-time alignment rule; bottom features of the visual sensor data are extracted and data-level fusion is performed with the radar sensor data; the fused data is selectively stored according to different weather conditions, training is performed according to an incremental learning rule, and a target detection result is obtained. The application effectively improves the target detection precision and robustness of a vehicle under severe weather, provides a reliable basis for a next-level task, and thus the purpose of improving driving safety is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, specifically to a detection method and system based on sensor fusion and incremental learning under adverse weather conditions, and more particularly to a target detection method and system based on sensor fusion and incremental learning technology under adverse weather conditions. Background Technology

[0002] Intelligent driving is of great significance to the future development of the automotive industry and represents the development level of my country's intelligent mobility industry. Specifically, it involves comprehensive analysis of data obtained from sensors combined with high-precision maps, autonomously making decisions and estimating vehicle status in different driving scenarios, and purposefully completing tracking control and collaborative control to reduce the probability of various road accidents caused by human factors, achieve effective and safe intelligent driving, and improve my country's industrial development level.

[0003] With the development of artificial intelligence technology, vehicle target detection has become a key research focus. Deep learning has played a significant role in this field, processing sensor data offline or online to enable intelligent vehicles to accurately identify objects in the driving environment, providing a foundation for subsequent tasks such as obstacle avoidance and path planning. However, when existing technologies are applied to target detection in adverse weather conditions, accuracy and robustness decrease, hindering safe driving and becoming one of the key issues in intelligent driving technology research.

[0004] Patent document CN110008843B discloses a vehicle target joint cognition method and system based on point cloud and image data. It includes a data-level joint module, a deep learning target detection module, and a joint cognition module. The data-level joint module acquires 3D point cloud data and image data, and fuses them. The fused data is then aggregated in the deep learning target detection module for feature-level detection and recognition, outputting the detection results. The joint cognition module uses evidence theory to judge the feature-level fusion detection results and the data-level fusion detection results, obtaining a confidence assignment as the output. However, this patent document does not integrate sensors with adverse weather conditions, failing to improve accuracy and robustness. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a detection method and system based on sensor fusion and incremental learning under severe weather conditions.

[0006] A detection method based on sensor fusion and incremental learning under severe weather conditions, provided by the present invention, includes the following steps:

[0007] Data acquisition steps: Data on the driving environment is acquired using radar sensors and vision sensors respectively, resulting in radar sensor data and vision sensor data.

[0008] Processing steps: The radar sensor and vision sensor data are processed according to the spatiotemporal alignment rules to obtain time-synchronized radar sensor and vision sensor data, spatially aligned binocular vision sensor data, and spatially aligned radar sensor and vision sensor data.

[0009] Fusion steps: Extract the low-level features of the visual sensor data and perform data-level fusion with the radar sensor data to obtain the data information stream;

[0010] Detection steps: Depending on the weather conditions, the fused data from the fusion step is selectively stored, and trained according to the rules of incremental learning to obtain the target detection result.

[0011] Preferably, the spatiotemporal alignment rule in the processing step includes:

[0012] At the radar sensor's t d Before and after the timestamp, take the visual sensor timestamp t with a time threshold of Δt. c The selection rules are as follows

[0013] |t c -t d |≤Δt

[0014] To obtain time-synchronized data from two types of sensors;

[0015] The coordinates of the corner points on the vision sensor calibration board in the world coordinate system are p. w After rigid body transformation, the visual sensor images with initially non-parallel optical centers are corrected by the same corner point p. w The imaging point falls at the same height on the left and right visual sensor images, achieving spatial alignment of the visual sensors. The visual sensor image undergoes rigid body transformation to convert the visual sensor coordinate system to the radar sensor coordinate system, thus achieving spatial alignment between the radar sensor and the visual sensor.

[0016] Preferably, the data-level fusion in the fusion step includes the following steps:

[0017] Integration steps: Read the images from the vision sensors, convert the image data into a data tensor with 3 channels, connect it after filtering, and integrate the data from the left and right vision sensors;

[0018] Extraction step: The integrated data from the integration step is processed by a convolutional network and normalization to extract the low-level features of the visual sensor image and retain detailed data information;

[0019] Normalization: The radar data tensor is normalized to preserve its original characteristics;

[0020] Data fusion steps: The raw data features from radar and visual sensors belonging to different distributions are fused by connecting the filter to the visual sensor image tensor after feature extraction, resulting in a normalized data information stream for different weather conditions.

[0021] Preferably, the incremental learning training in the detection step includes the following steps:

[0022] Segmentation steps: Weather data belonging to different domains are continuously segmented; among them, data stream batches near the segmentation points of different weather data streams include data from both the left and right sides;

[0023] Input steps: The image data to be continuously segmented is fed into the model for training in the form of tensors. During the training process, M tensors are stored in a fixed quantity.

[0024] Training steps: The input data stream is processed by a style metric to measure the similarity between two tensors. If the similarity is high, the new tensor with high similarity replaces the old tensor with the lowest similarity in storage. If the similarity is low, for tensors whose style differs too much from the style of the existing domains, pseudo-domain detection is performed using the random forest algorithm and the sparse random projection algorithm. Tensors that fail the detection are stored separately, and when a certain number are accumulated, they are added to M storage locations. At the same time, the same number of tensors are removed from the M storage locations to maintain a balance between storage and the number of data from different domains.

[0025] Preferably, the spatial alignment process includes:

[0026] Binocular vision sensor correction: The rotation matrix R and translation matrix T of the right visual sensor relative to the left visual sensor are estimated based on the rotation and translation matrices of the left and right visual sensors.

[0027] The rotation matrix can be calculated using the following formula:

[0028]

[0029] In the formula, R l R r This is the rotation matrix of the left and right vision sensors relative to the world coordinate system, where the subscripts l and r represent the left and right sides respectively, and the superscript T represents the matrix transpose, and satisfies the following conditions: The superscript -1 indicates the inverse of the matrix;

[0030] The translation matrix can be calculated using the following formula:

[0031] T = T r -RT l

[0032] In the formula, T l T rIt is the translation matrix of the left and right visual sensors relative to the world coordinate system;

[0033] Corner point p on the left calibration plate w Its coordinates (U,V) are transformed from the world coordinate system to the left visual sensor coordinate system through rigid body transformation, resulting in the coordinates (X,V) of the point in the left visual sensor coordinate system. l ,Y l Z l );

[0034] coordinates (X) l ,Y l Z l The calculation of ) can be obtained from the following formula:

[0035]

[0036] In the formula, R l T l These are the rotation matrix and translation matrix of the left visual sensor coordinate system relative to the world coordinate system, respectively.

[0037] Coordinates in the left visual sensor coordinate system (X) l ,Y l Z l After rigid body transformation, perspective projection, and radial transformation, the pixel coordinates (u′) of the point are obtained by converting the image to the pixel coordinate system of the image obtained from the right-side visual sensor. r ,v′ r ), that is, corner point p w The projected coordinates of the coordinates in the pixel coordinate system;

[0038] The projected coordinates can be calculated using the following formula:

[0039]

[0040] In the formula, K r It is the intrinsic parameter matrix of the right-side vision sensor;

[0041] Take n calibration board corner points from m images, optimize the objective function, and use the LM algorithm to optimize all the matrix parameters to minimize the projection error of the corner points;

[0042] The objective function can be calculated using the following formula:

[0043]

[0044] In the formula, (u r ,v r () represents the actual position of the corner point in the pixel coordinate system of the image obtained by the right-side visual sensor;

[0045] For the left-side visual sensor, in the spatial alignment of the radar sensor and the visual sensor, according to its rotation matrix R relative to the radar sensor... lra Translation matrix T lra Achieve coordinate system transformation, converting points in the left-side visual sensor coordinate system to the radar sensor coordinate system;

[0046] The coordinate system transformation can be obtained from the following formula:

[0047]

[0048] In the formula, (X l ,Y l Z l (X) represents the coordinates in the left visual sensor coordinate system. ra ,Y ra Z ra ) represents the coordinates in the corresponding radar sensor coordinate system.

[0049] Preferably, the data fusion step includes:

[0050] Two data tensors f from the left and right vision sensors l f r After filtering and concatenation, a single visual sensor feature map is obtained. This visual sensor feature map is then subjected to 1×1 convolution for channel fusion, integrating information from different channels. Finally, it undergoes 3×3 convolution to change the number of channels, resulting in the final visual sensor feature map f. c ;

[0051] The feature map of a visual sensor can be calculated using the following formula:

[0052]

[0053] In the formula, The filter connections are represented by W1 and W2, which represent the convolutions of the first and second layers of the network, respectively, and BN1 and BN2 represent the batch normalizations of the first and second layers of the network, respectively.

[0054] Radar data tensor f ra After normalization, the visual sensor feature map f is compared with the normalized version. c Data-level fusion is achieved through filter concatenation to obtain feature map f;

[0055] The feature map can be calculated using the following formula:

[0056]

[0057] In the formula, F N This represents a normalization function that scales the feature map elements to 0-1.

[0058] The normalization function can be calculated using the following formula:

[0059]

[0060] In the formula, a = f - min(f), where f is a two-dimensional tensor.

[0061] Preferably, the training steps include:

[0062] The three-dimensional feature map f is flattened to obtain the two-dimensional feature map x, while the number of channels remains unchanged. The Gram matrix G is obtained from the two-dimensional feature map x.

[0063] The Gram matrix can be calculated using the following formula:

[0064] G = x T ×x

[0065] In the formula, × represents matrix multiplication, and T represents matrix transpose;

[0066] The Gram matrix G is used to measure the difference between two feature maps and is denoted by d;

[0067] The difference can be calculated using the following formula:

[0068]

[0069] In the formula, l represents the l-th layer network. Let F represent the i-th and j-th elements of two Gram matrices G and A, where G and A are Gram matrices calculated from two different feature maps. F = h × w represents the size of the Gram matrix, where h and w are the length and width of the feature map, respectively.

[0070] This invention also provides a detection system based on sensor fusion and incremental learning under severe weather conditions, comprising:

[0071] Data Acquisition Module: Collects data on the driving environment using radar and vision sensors respectively, obtaining radar and vision sensor data;

[0072] Processing module: Processes the radar sensor and vision sensor data according to the spatiotemporal alignment rules to obtain time-synchronized radar sensor and vision sensor data, spatially aligned binocular vision sensor data, and spatially aligned radar sensor and vision sensor data.

[0073] Fusion module: Extracts low-level features from visual sensor data and fuses them with radar sensor data at the data level to obtain a data information stream;

[0074] Detection module: Depending on the weather conditions, the data fused in the fusion module is selectively stored, trained according to the rules of incremental learning, and the target detection result is obtained.

[0075] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described detection method based on sensor fusion and incremental learning under severe weather conditions.

[0076] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the above-described detection method based on sensor fusion and incremental learning under adverse weather conditions.

[0077] Compared with the prior art, the present invention has the following beneficial effects:

[0078] 1) This invention proposes a relatively complete sensor fusion and incremental learning technology, which processes multi-sensor data through spatiotemporal alignment rules, then fuses it according to data-level fusion rules, and then uses incremental learning rules to store and train the data. Finally, the target detection results under severe weather conditions are obtained through analysis. This not only improves the target detection accuracy under severe weather conditions, but also has strong robustness.

[0079] 2) This invention proposes a highly accurate and robust target detection method, which reduces the probability of missed or incorrect target detection when vehicles are driving in adverse weather conditions, avoids safety accidents caused by incorrect target detection, and improves the driving safety of intelligent vehicles and road traffic efficiency. Attached Figure Description

[0080] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0081] Figure 1 This is a flowchart illustrating a detection method based on sensor fusion and incremental learning under severe weather conditions. Detailed Implementation

[0082] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0083] like Figure 1As shown, this embodiment discloses a detection method based on sensor fusion and incremental learning under severe weather conditions, including the following steps:

[0084] Radar and vision sensors collect data on the driving environment to obtain driving environment information;

[0085] The two types of sensor data obtained above are processed according to the spatiotemporal alignment rules;

[0086] Extract low-level features from visual sensor data and perform data-level fusion with radar sensor data;

[0087] Depending on the weather conditions, the fused data is selectively stored and trained according to the rules of incremental learning to obtain target detection results.

[0088] It should be noted that the onboard sensors of the main vehicle include binocular vision sensors and radar sensors.

[0089] The main vehicle inspection process is as follows:

[0090] In adverse weather conditions, the main vehicle captures driving environment information through onboard sensors, analyzes the environmental information using sensor fusion and incremental learning techniques, identifies targets to be detected, and outputs detection results. It then uses sensor fusion and incremental learning techniques again to analyze and identify the next target to be detected in the driving environment, outputting detection results. The main vehicle continuously repeats these steps to achieve target detection in adverse weather conditions.

[0091] Specifically, the radar and vision sensors respectively collect data on the driving environment to obtain driving environment information, including:

[0092] In severe weather, the radar sensors of the main vehicle are at t d High-resolution range-azimuth images are acquired continuously using a binocular vision sensor at t c Color images are collected continuously to form the main vehicle's driving environment information data, resulting in radar sensor images, left visual sensor images, and right visual sensor images, which together form a dataset.

[0093] Specifically, the process of processing the two types of sensor data obtained above according to spatiotemporal alignment rules includes:

[0094] Data from radar and binocular vision sensors are synchronized according to timestamp intervals. One-to-one matching sensor data is selected, and the vision sensor is used for correction to achieve spatial alignment of the same object.

[0095] It should be noted that radar and binocular vision sensors have different sampling frequencies, resulting in mismatched raw driving environment data. The rules for selecting time-matched data from different sensors are as follows:

[0096] |t c -t d |≤Δt

[0097] In the formula, Δt is the time threshold. Based on the radar sensor timestamp, visual sensor images that exceed this time threshold will be discarded, thereby obtaining one-to-one matched sensor data.

[0098] It should be further explained that in the spatial calibration alignment rule, n calibration board corner points are taken from m images, (u r ,v r (u) represents the actual position of the corner point in the pixel coordinate system of the right-side visual sensor. ′ r ,v r ′ If is its estimated value, then the projection error is calculated as follows:

[0099]

[0100] The estimated value of the corner projection (u) ′ r ,v r ′ The intrinsic parameter matrix K of the right-side visual sensor r The rotation matrix R and translation matrix T of the right visual sensor relative to the left visual sensor, and the rotation matrix R of the left visual sensor coordinate system relative to the world coordinate system. l Translation matrix T l The rotation matrix R and translation matrix T are determined by the rotation matrices R and T of the left and right visual sensors relative to the world coordinate system. l R r Translation matrix T l T r We can obtain:

[0101]

[0102] To solve for the matrix parameters, the LM (Levenberg-Marquardt) algorithm is used to optimize the projection error, minimizing the projection error at the corner points.

[0103] Specifically, the extraction of low-level features from visual sensor data and its data-level fusion with radar sensor data includes:

[0104] The dataset is categorized by weather and fed into the network, which continuously splits the dataset into fixed batch sizes, such that batches near the dividing line contain data on different weather conditions on both sides.

[0105] Visual sensor images are processed via bilinear interpolation xs =(x d +0.5)×w x -0.5, converted to radar sensor Figure 1 An image of the same size. x The scale factor is determined by the size of the original visual sensor image and the size of the radar sensor image, x. d These are the pixels in the original image.

[0106] Two data tensors f from the left and right vision sensors l f r A single visual sensor feature map is obtained after filtering and concatenation. To integrate features from the two visual sensor images, convolution and batch normalization were applied to f. c First, channel fusion is performed using a 1×1 convolution W1 to integrate information from different channels. Then, a batch normalization BN1 layer is applied to obtain f. c :

[0107] f c =BN1[W1(f c )]

[0108] Then, after a 3×3 convolution W2 and a batch normalization BN2 layer to change the number of channels, it is finally passed through a normalization function.

[0109]

[0110] Scaling the feature map elements to 0-1 yields the visual sensor feature map f before fusion. c :

[0111] f c =F N (BN2[W2(f c )])

[0112] During this process, the size of the visual sensor feature map remains unchanged, only its number of channels is changed, in order to preserve the original information to the greatest extent possible.

[0113] Radar data tensor f ra After normalization f ra =F N (f ra ), and the above f c The feature map is obtained after data-level fusion through filtering and concatenation.

[0114] Specifically, depending on the weather conditions, the fused data is selectively stored, trained according to incremental learning rules, and the resulting object detection results include:

[0115] Feature map of each batch Transforming a three-dimensional tensor into a two-dimensional tensor Calculate the Gram matrix G = x in this way T ×x. Assume the L-th layer network has Gram matrices G for two different feature maps. l ={g l} ij and A l ={a l} ij The distinct styles of the feature maps can be obtained by measuring their differences. The difference can be calculated using the following formula:

[0116]

[0117] With a fixed storage size M, the input data to the network is continuously stored until the storage capacity is reached. If the storage capacity is full, the feature map in storage that is most similar to the new feature map style is replaced by calculating the difference in feature map styles. The discovery of new styles is achieved by random forest and sparse random projection algorithms to ensure that feature maps of different weather conditions are included in memory.

[0118] Each round of training data includes new data and historical data randomly selected from storage. The memory is continuously updated during the training process and participates in the next round of training, so as to continuously strengthen historical memory while learning new knowledge.

[0119] The present invention also provides a detection system based on sensor fusion and incremental learning under severe weather conditions. The detection system based on sensor fusion and incremental learning under severe weather conditions can be implemented by executing the process steps of the detection method based on sensor fusion and incremental learning under severe weather conditions. That is, those skilled in the art can understand the detection method based on sensor fusion and incremental learning under severe weather conditions as a preferred embodiment of the detection system based on sensor fusion and incremental learning under severe weather conditions.

[0120] A detection system based on sensor fusion and incremental learning for adverse weather conditions includes: an acquisition module: radar sensors and visual sensors respectively acquire data of the driving environment to obtain driving environment information; a processing module: processing the two types of sensor data obtained in the acquisition module according to spatiotemporal alignment rules; a fusion module: extracting low-level features from the visual sensor data and performing data-level fusion with the radar sensor data; and a detection module: selectively storing the fused data according to different weather conditions, training it according to incremental learning rules, and obtaining target detection results.

[0121] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0122] In the description of this application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0123] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A detection method based on sensor fusion and incremental learning under severe weather conditions, characterized in that, Includes the following steps: Data acquisition steps: Data on the driving environment is acquired using radar sensors and vision sensors respectively, resulting in radar sensor data and vision sensor data. Processing steps: The radar sensor and vision sensor data are processed according to the spatiotemporal alignment rules to obtain time-synchronized radar sensor and vision sensor data, spatially aligned binocular vision sensor data, and spatially aligned radar sensor and vision sensor data. Fusion steps: Extract the low-level features of the visual sensor data and perform data-level fusion with the radar sensor data to obtain the data information stream; Detection steps: Depending on the weather conditions, selectively store the fused data from the fusion step, train it according to the rules of incremental learning, and obtain the target detection result; The incremental learning training in the detection step includes the following steps: Segmentation steps: Weather data belonging to different domains are continuously segmented; among them, data stream batches near the segmentation points of different weather data streams include data from both the left and right sides; Input steps: The image data to be continuously segmented is fed into the model for training in the form of tensors. During the training process, M tensors are stored in a fixed quantity. Training steps: The input data stream is processed by a style metric to measure the similarity between two tensors. If the similarity is high, the new tensor with high similarity replaces the old tensor with the lowest similarity in storage. If the similarity is low, for tensors whose style differs too much from the style of the existing domain, pseudo-domain detection is performed using the random forest algorithm and the sparse random projection algorithm. Tensors that fail the detection are stored separately, and when a certain number are accumulated, they are added to M storage locations. At the same time, the same number of tensors are removed from the M storage locations to maintain the balance between storage and the number of data from different domains. The data-level fusion in the fusion step includes the following steps: Two data tensors from the left and right visual sensors After filtering and concatenation, a single visual sensor feature map is obtained. This visual sensor feature map is then subjected to 1×1 convolution for channel fusion, integrating information from different channels. Finally, it undergoes 3×3 convolution to change the number of channels, resulting in the final visual sensor feature map. ; The feature map of a visual sensor can be calculated using the following formula: In the formula, Indicates filter connection, These represent the convolutions of the first and second layers of the network, respectively. These represent the batch normalization of the first and second layers of the network, respectively. Radar data tensor Normalized visual sensor feature maps Data-level fusion is achieved through filter concatenation to obtain feature maps. ; The feature map can be calculated using the following formula: In the formula, This represents a normalization function that scales the feature map elements to 0-1. The normalization function can be calculated using the following formula: In the formula, , It is a two-dimensional tensor.

2. The detection method based on sensor fusion and incremental learning under severe weather conditions according to claim 1, characterized in that, The spatiotemporal alignment rules in the processing steps include: In radar sensors Before and after the timestamp, take the time threshold as follows: Visual sensor timestamps The selection rules are as follows To obtain time-synchronized data from two types of sensors; The coordinates of the corner points on the vision sensor calibration board in the world coordinate system are: After rigid body transformation, the visual sensor images with initially non-parallel optical centers are corrected to the same corner point. The imaging point falls at the same height on the left and right visual sensor images, achieving spatial alignment of the visual sensors. The visual sensor image undergoes rigid body transformation to convert the visual sensor coordinate system to the radar sensor coordinate system, thus achieving spatial alignment between the radar sensor and the visual sensor.

3. The detection method based on sensor fusion and incremental learning under severe weather conditions according to claim 2, characterized in that, The spatial alignment process includes: Binocular vision sensor correction estimates the rotation matrix of the right visual sensor relative to the left visual sensor based on the rotation and translation matrices of the left and right visual sensors. Translation matrix ; The rotation matrix can be calculated using the following formula: In the formula, It is the rotation matrix of the left and right visual sensors relative to the world coordinate system, subscript Indicates left and right sides respectively, superscript Let represent the matrix transpose, and satisfy . The superscript -1 indicates the inverse of the matrix; The translation matrix can be calculated using the following formula: In the formula, It is the translation matrix of the left and right visual sensors relative to the world coordinate system; Corner of the left calibration plate Its coordinates After rigid body transformation, the coordinates of the point are obtained from the world coordinate system to the left visual sensor coordinate system. ; coordinate The calculation can be obtained from the following formula: In the formula, These are the rotation matrix and translation matrix of the left visual sensor coordinate system relative to the world coordinate system, respectively. Coordinates in the left visual sensor coordinate system After rigid body transformation, perspective projection, and radial transformation, the pixel coordinates of the point are obtained by converting the image to the pixel coordinate system of the image obtained from the right-side visual sensor. Corner The projected coordinates of the coordinates in the pixel coordinate system; The projected coordinates can be calculated using the following formula: In the formula, It is the intrinsic parameter matrix of the right-side vision sensor; Take n calibration board corner points from m images, optimize the objective function, and use the LM algorithm to optimize all the matrix parameters to minimize the projection error of the corner points; The objective function can be calculated using the following formula: In the formula, This represents the actual position of the corner point in the pixel coordinate system of the image obtained from the right-side visual sensor; For the left-side vision sensor, in the spatial alignment of the radar sensor and the vision sensor, based on its rotation matrix relative to the radar sensor... Translation matrix Achieve coordinate system transformation, converting points in the left-side visual sensor coordinate system to the radar sensor coordinate system; The coordinate system transformation can be obtained from the following formula: In the formula, These are the coordinates in the left-hand visual sensor coordinate system. These are the coordinates in the corresponding radar sensor coordinate system.

4. The detection method based on sensor fusion and incremental learning under severe weather conditions according to claim 1, characterized in that, The training steps include: 3D feature map Flattened to obtain a two-dimensional feature map The number of channels remains unchanged, based on the two-dimensional feature map. Obtain the Gram matrix ; The Gram matrix can be calculated using the following formula: In the formula, This represents matrix multiplication, and T represents matrix transpose. Gram matrix It is used to measure the difference between two feature maps, and is denoted by d; The difference can be calculated using the following formula: In the formula, Indicates the first Layered network, Let the first digit of two Gram matrices G and A be the first digit of the second digit of the third digit of i, j There are 12 elements, G and A are the Gram matrices calculated from two different feature maps, and F = h × w represents the size of the Gram matrix, where h and w are the length and width of the feature map, respectively.

5. A detection system based on sensor fusion and incremental learning for severe weather conditions, characterized in that, include: Data Acquisition Module: Collects data on the driving environment using radar and vision sensors respectively, obtaining radar and vision sensor data; Processing module: Processes the radar sensor and vision sensor data according to the spatiotemporal alignment rules to obtain time-synchronized radar sensor and vision sensor data, spatially aligned binocular vision sensor data, and spatially aligned radar sensor and vision sensor data. Fusion module: Extracts low-level features from visual sensor data and fuses them with radar sensor data at the data level to obtain a data information stream; Detection module: Depending on the weather conditions, the data fused in the fusion module is selectively stored, trained according to the rules of incremental learning, and the target detection result is obtained; The incremental learning training process in the detection module is as follows: Weather data belonging to different domains are continuously segmented; among them, data stream batches near the segmentation point of different weather data streams contain data from both the left and right sides; The continuously segmented image data is fed into the model for training in the form of tensors, and M tensors are stored in a fixed quantity during the training process. The input data stream is processed by a style metric to measure the similarity between two tensors. If the similarity is high, the new tensor with high similarity replaces the old tensor with the lowest similarity in storage. If the similarity is low, for tensors whose style differs too much from the style of the existing domain, pseudo-domain detection is performed using the random forest algorithm and the sparse random projection algorithm. Tensors that fail the detection are stored separately and added to M storage units when a certain number are accumulated. At the same time, the same number of tensors are removed from the M storage units to maintain the balance between storage and the number of data from different domains. The data-level fusion process in the fusion module is as follows: Two data tensors from the left and right visual sensors After filtering and concatenation, a single visual sensor feature map is obtained. This visual sensor feature map is then subjected to 1×1 convolution for channel fusion, integrating information from different channels. Finally, it undergoes 3×3 convolution to change the number of channels, resulting in the final visual sensor feature map. ; The feature map of a visual sensor can be calculated using the following formula: In the formula, Indicates filter connection, These represent the convolutions of the first and second layers of the network, respectively. These represent the batch normalization of the first and second layers of the network, respectively. Radar data tensor Normalized visual sensor feature maps Data-level fusion is achieved through filter concatenation to obtain feature maps. ; The feature map can be calculated using the following formula: In the formula, This represents a normalization function that scales the feature map elements to 0-1. The normalization function can be calculated using the following formula: In the formula, , It is a two-dimensional tensor.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the detection method based on sensor fusion and incremental learning under severe weather conditions as described in any one of claims 1 to 4.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the detection method based on sensor fusion and incremental learning under severe weather conditions as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • A joint vehicle target cognition method and system based on point cloud and image data

    CN110008843B