A multi-dimensional pixel fusion method for environmental perception and a storage medium
By using a multi-dimensional pixel fusion method, frame synchronization signals and hard-wired trigger signals are used to drive multiple sensors to acquire data and fuse features. This solves the limitations of single-sensor perception systems, achieves temporal alignment and spatial registration of sensor data, and improves the environmental perception accuracy of autonomous driving systems.
Patent Information
- Application Number
- CN202411168504.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-08-23
AI Technical Summary
Existing single-sensor perception systems cannot accurately identify and understand the environment around a vehicle, and therefore have limitations.
A multi-dimensional pixel fusion method is adopted, which uses frame synchronization signal and hard-line trigger signal to drive visual sensor, infrared sensor and millimeter-wave radar to collect time-aligned raw environmental data, perform feature extraction and spatial coordinate alignment, construct weight matrix for feature fusion, and obtain multi-dimensional perception data.
It achieves temporal alignment and spatial registration of data from different sensors, improves the accuracy of obstacle recognition, provides multimodal accurate perception information, and enhances the environmental understanding ability of autonomous driving systems.
Smart Images

Figure CN119223299B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of driving assistance, in particular to a multi-dimensional pixel fusion method for environment perception and a storage medium. BACKGROUND
[0002] Advanced Driving Assistance System is to use various sensors (millimeter wave radar, laser radar, single\double camera and satellite navigation) installed on the car to sense the environment around the car at any time during driving, collect data, identify, detect and track static and dynamic objects, and combine navigation map data to perform system calculation and analysis, so as to let the driver aware of the possible danger in advance, effectively increase the comfort and safety of car driving. In recent years, the ADAS market has grown rapidly. Originally, such systems were limited to the high-end market, but now they are entering the mid-end market. At the same time, many low-tech applications are more common in entry-level passenger vehicles, and improved new sensor technology is also creating new opportunities and strategies for system deployment.
[0003] Therefore, the perception system of the autonomous vehicle needs to accurately identify and understand the environment around the vehicle. Currently, the perception system of a single sensor has limitations, for example: the camera has high resolution but is greatly affected by light, the radar can penetrate bad weather but has low resolution, and the infrared camera is suitable for night use but has limited detail capture capability. SUMMARY
[0004] The present application provides a multi-dimensional pixel fusion method for environment perception and a storage medium, which solves the technical problem that the existing single sensor perception system cannot accurately identify and understand the environment around the vehicle.
[0005] To solve the above technical problems, the present application provides a multi-dimensional pixel fusion method for environment perception, comprising:
[0006] Based on the frame synchronization signal and the hard-wired trigger signal, the sensor is driven to collect environmental information, and time sequence aligned original environmental data is obtained, wherein the sensor includes a visual sensor, an infrared sensor and a millimeter wave radar;
[0007] The original environmental data is preprocessed, and based on the processed original environmental data, feature extraction is performed to obtain corresponding visual features, infrared features and dynamic features;
[0008] The space coordinates are aligned to obtain a spatially registered FOV field of view overlap region;
[0009] A weight matrix is constructed, and then the visual features, infrared features and dynamic features of the FOV field of view overlap area are fused to obtain multi-dimensional perception data.
[0010] In a further embodiment, the environment information is collected based on the frame synchronization signal and the hard-wired trigger signal, and the original environment data is obtained in time sequence alignment, including:
[0011] The control signal is simultaneously issued to the visual deserializer, the infrared deserializer and the millimeter wave radar;
[0012] The visual sensor is controlled based on the frame synchronization signal, the visual data is obtained from the visual camera serializer, and the visual data is deserialized by the visual deserializer to obtain the visual data;
[0013] The infrared sensor is controlled based on the frame synchronization signal, the infrared data is obtained from the infrared camera serializer, and the infrared data is deserialized by the infrared deserializer to obtain the infrared data;
[0014] The millimeter wave radar is driven based on the hard-wired trigger signal to collect data and obtain point cloud data;
[0015] The original environment data includes visual data, infrared data and point cloud data.
[0016] In a further embodiment, the feature extraction based on the processed original environment data includes:
[0017] The visual features are extracted from the visual data based on a deep learning model, the infrared features are extracted from the infrared data based on a deep learning model, and the dynamic features are extracted from the point cloud data based on an optical flow method or a key point detection algorithm.
[0018] In a further embodiment, the spatial coordinate alignment processing is performed to obtain the spatially registered FOV field of view overlap area, including:
[0019] A plane coordinate system is established with the FOV center point of the millimeter wave radar as the origin of the coordinate system;
[0020] The coordinate offset values of the visual FOV center point, the infrared FOV center point and the origin of the coordinate system are calculated;
[0021] The coordinate offset values and the sensor resolution are obtained with the plane coordinate system as a reference, and the first coordinate group, the second coordinate group and the third coordinate group corresponding to the radar field of view FOV range, the visual FOV range and the infrared FOV range are obtained respectively;
[0022] Based on the first coordinate group, the second coordinate group, the third coordinate group, a fourth coordinate group of a FOV field of view overlap area is calculated, so that the coordinates of the pixel points in the FOV field of view overlap area are aligned; the FOV field of view overlap area is a field of view overlap area of a radar field of view FOV range, a visual FOV range, and an infrared FOV range.
[0023] In further embodiments, assuming that the resolution of the millimeter wave radar is U*V, the resolution of the visual sensor is M*N, and the resolution of the infrared sensor is R*T, then:
[0024] The fourth coordinate group of the FOV field of view overlap area is as follows:
[0025]
[0026] In the formula, Axy, Bxy, Cxy, and Dxy are the coordinates of the four vertices of the FOV field of view overlap area; (m, n) is the coordinate offset value of the visual FOV center point from the origin of the coordinate system; and (p, q) is the coordinate offset value of the infrared FOV center point from the origin of the coordinate system.
[0027] In further embodiments, the constructing of the weight matrix comprises:
[0028] Obtaining obstacle data of the FOV field of view overlap area and establishing a probability distribution model;
[0029] According to the probability distribution estimated by the probability distribution model, the information entropy of the visible light camera, the 4D millimeter wave radar, and the infrared camera is calculated, respectively;
[0030] Normalization processing is performed to convert the information entropy into a confidence weight coefficient;
[0031] All the confidence weight coefficients corresponding to the visual sensor, the millimeter wave radar, and the infrared sensor are integrated to obtain a weight matrix.
[0032] In further embodiments, the calculation formula of the information entropy is as follows:
[0033]
[0034] In the formula, H(X) is the information entropy of a random variable X, p(x i ) is the probability of X taking the i-th value, and n is the total number of possible values.
[0035] The calculation formula of the confidence weight coefficient is as follows:
[0036]
[0037] In the formula, w i is the weight coefficient of the i-th sensor, and H(Xi ) is the information entropy of the i-th sensor, and ∈ is a positive number.
[0038] In a further embodiment, the feature fusion of the visual features, infrared features and dynamic features of the FOV field of view overlap region obtains multi-dimensional perception data, including:
[0039] Based on the FOV field of view overlap region, obstacle pixel matrices, point cloud pixel matrices, and thermal radiation pixel matrices corresponding to the visual features, infrared features, and dynamic features are respectively obtained;
[0040] The obstacle pixel matrices, point cloud pixel matrices, and thermal radiation pixel matrices are subjected to feature splicing processing to obtain a multi-dimensional pixel matrix;
[0041] According to the multi-dimensional pixel matrix and the weight matrix, multi-dimensional pixel fusion is performed to obtain multi-dimensional perception data;
[0042] The obstacle data includes color channel values (R, G, B) of an object image collected by a visual sensor, distance L, speed S, azimuth R, and height signal h of the object collected by a millimeter wave radar, and thermal radiation H of the object collected by an infrared sensor.
[0043] In a further embodiment, the multi-dimensional pixel matrix is as follows:
[0044]
[0045] The fusion formula of the multi-dimensional perception data is as follows:
[0046]
[0047] wherein,
[0048]
[0049] In the formula, W is a weight matrix, are the confidence weight coefficients of the visual sensor, the millimeter wave radar, and the infrared sensor, respectively; P Matrix_A , P Matrix_B , P Matrix_C represent obstacle attribute features detected by the visual sensor, obstacle attribute features detected by the millimeter wave radar, and obstacle attribute features detected by the infrared sensor, respectively.
[0050] The application further provides a storage medium having a computer program stored thereon, the computer program being used to implement the multi-dimensional pixel fusion method for environment perception.
[0051] The application has the following advantages:
[0052] (1) The control signals are simultaneously sent to the visual deserializer, the infrared deserializer and the millimeter wave radar, the frame synchronization signals and the hard-wire trigger signals are used to simultaneously drive the visual sensor, the infrared sensor and the millimeter wave radar to collect data, the data is collected at the same time, the timing alignment between the sensors is realized, and a fusion basis for the multi-dimensional data fusion is provided.
[0053] (2) Based on the particularity of the FOV center, the coordinate offset values of the visual FOV center point, the infrared FOV center point and the coordinate system origin (the FOV center point of the millimeter wave radar) are calculated to determine the offset range of the entire field of view, and then the positions of the visual FOV center point and the infrared FOV center point relative to the plane coordinate system are determined based on the coordinate offset values, the data collected by the three sensors is fused in the same coordinate system, the spatial alignment is realized, different types of sensors and different FOV sensors can find a multi-dimensional pixel field of view overlapping area through software algorithm, and the coordinate space in the area is aligned.
[0054] (3) Based on the characteristics that the obstacle data collected by different sensors are different in type, a probability distribution model is established according to real-time data, then the information entropy is calculated and converted into a confidence weight coefficient, the confidence weight coefficient of each sensor is updated in real time, the actual adaptation degree of the obstacle recognition coefficient is ensured, and the recognition accuracy is improved.
[0055] (4) Based on the different types of obstacle data collected by different sensors, the detection data of each sensor is "unified in coordinates and aligned in timing", the image and radar data are completed in real time "spatial and temporal alignment and synchronization" at the pixel level and are output in the format of "multi-dimensional pixels" through splicing and fusion. The multi-modal accurate perception information of the target and the environment is provided for the automatic driving system: including the comprehensive perception of the visual data (light and shade, texture, color, etc.), 4D millimeter wave radar data (distance, speed, direction, height, etc.) and infrared radiation data (texture, temperature, etc.) of the target and the environment perceived by the sensor, so as to realize the accurate identification of the target obstacle in various driving scenes and improve the driving safety. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1is a workflow diagram of a multi-dimensional pixel fusion method for environment perception provided by an embodiment of the present application;
[0057] Figure 2 is a detection domain space mapping relationship diagram of a multi-sensor combination provided by an embodiment of the present application;
[0058] Figure 3 is a synchronous acquisition diagram of an FPGA provided by an embodiment of the present application;
[0059] Figure 4 is a sensor installation layout diagram provided by an embodiment of the present application;
[0060] Figure 5 is a sensor coordinate alignment diagram provided by an embodiment of the present application;
[0061] Figure 6 is a feature stitching diagram provided by an embodiment of the present application. DETAILED DESCRIPTION
[0062] The embodiments of the present application will be described in detail below with reference to the accompanying drawings. The embodiments are presented only for the purpose of illustration and should not be understood as limiting the present application. The accompanying drawings are used for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present application, because many changes can be made to the present application without departing from the spirit and scope thereof.
[0063] Embodiment 1
[0064] A multi-dimensional pixel fusion method for environment perception provided by an embodiment of the present application is as shown in the figure, in the present embodiment, comprising: Figures 1-6
[0065] S1, driving the sensor to collect environment information based on the frame synchronization signal and the hard-wire trigger signal, obtaining time sequence aligned original environment data, comprising:
[0066] At the same time, the control signal is issued to the visual deserializer, the infrared deserializer and the millimeter wave radar;
[0067] The visual sensor is controlled to expose based on the frame synchronization signal, the visual data is obtained from the visual camera serializer, and the visual data is deserialized by the visual deserializer to obtain the visual data;
[0068] The infrared sensor is controlled to expose based on the frame synchronization signal, the infrared data is obtained from the infrared camera serializer, and the infrared data is deserialized by the infrared deserializer to obtain the infrared data;
[0069] The millimeter wave radar is driven to collect data based on the hard-wire trigger signal, and point cloud data is obtained;
[0070] The sensors include a visual sensor (for example, a front-view 8M camera), an infrared sensor (that is, an infrared camera), and a millimeter wave radar; and the raw environment data includes visual data, infrared data, and point cloud data.
[0071] In this embodiment, a 4D millimeter wave radar is preferably used to collect point cloud data, the 4D millimeter wave radar collects millimeter wave radar data through a hard-wire trigger signal issued by an SOC, and inputs the point cloud data to the SOC through an Ethernet for further data processing.
[0072] In this embodiment, time synchronization of the sensors is performed through hard-wire and image frame exposure control to align the time sequences of the sensors. Referring to Figure 3 , the specific steps are as follows:
[0073] When the system is started, an exposure control center of an FPGA (field programmable gate array) issues a control signal to a deserializer 1 (that is, a visual deserializer) through Ctrl0, and the deserializer 1 feeds the control signal to an exposure control pin of a sensor 1 (that is, a visual sensor) through a GMSL link reverse channel for exposure control.
[0074] Similarly, the exposure control center of the FPGA issues a control signal to a deserializer 2 (that is, an infrared deserializer) through Ctrl1, and the deserializer 2 feeds the control signal to an exposure control pin of a sensor 2 for exposure control.
[0075] The time synchronization of the 4D millimeter wave radar is performed through hard-wire control, and the FPGA control center controls radar data collection through a Ctrl2 control signal.
[0076] S2, pre-processing the raw environment data, and performing feature extraction based on the processed raw environment data to obtain corresponding visual features, infrared features, and dynamic features;
[0077] Specifically, the visual features are extracted from the visual data based on a deep learning model, the infrared features are extracted from the infrared data based on a deep learning model, and the dynamic features are extracted from the point cloud data based on an optical flow method or a key point detection algorithm. In the specific implementation process, the feature extraction algorithm can be selected according to actual needs, and the present embodiment is not limited in this regard.
[0078] S3, performing spatial coordinate alignment processing to obtain a spatially registered FOV field of view overlap region, including:
[0079] S31, establishing a plane coordinate system (that is, O1-XY plane) with a FOV center point of the millimeter wave radar as an origin of the coordinate system;
[0080] S32, calculating coordinate offset values of the visual FOV center point, the infrared FOV center point, and the origin of the coordinate system.
[0081] S33, taking the planar coordinate system as a reference, obtaining the coordinate offset value and the sensor resolution, and obtaining a first coordinate group, a second coordinate group and a third coordinate group corresponding to a radar field of view (FOV) range, a visual FOV range and an infrared FOV range respectively;
[0082] The coordinate alignment in the prior art is usually based on a spatial coordinate system, as shown in Figure 4 , O1 is the FOV1 origin of the 4D millimeter wave radar; O2 is the center point of the FOV2 of the visible light camera; and O3 is the center point of the FOV3 of the infrared camera. A multi-dimensional pixel sensor pixel coordinate system O1-XYZ is established with O1 as the coordinate origin.
[0083] Among them, the coordinate offset of the FOV2 center point of the visible light camera from O1 is (m, n, 0); and the coordinate offset of the FOV3 center point of the infrared camera from O1 is (p, q, 0).
[0084] For this, in the embodiment, it is assumed that the resolution of the millimeter wave radar is U*V, the resolution of the visual sensor is M*N, and the resolution of the infrared sensor is R*T, then:
[0085] Referring to Figure 5 , taking the field of view FOV coordinate of the 4D millimeter wave radar as a reference, it is assumed that on the O1-XY plane, the first coordinate group of the FOV range of the 4D millimeter wave radar is P1 xy , the second coordinate group of the FOV range of the visual sensor is P2 xy , and the third coordinate group of the FOV range of the infrared sensor is P3 xy , so that:
[0086]
[0087]
[0088] Among them, Pi xy (x, y are integers) is the pixel coordinate of the i-th sensor on the O1-XY plane, and satisfies the following relationship:
[0089] x≤max{U,M,R}
[0090] y≤max{V,N,T}
[0091] S34, based on the first coordinate group, the second coordinate group and the third coordinate group, calculating a fourth coordinate group corresponding to a FOV field of view overlap area, so that the coordinate space of the pixel points in the FOV field of view overlap area is aligned; the FOV field of view overlap area is a field of view overlap area of the radar field of view FOV range, the visual FOV range and the infrared FOV range (such as Figure 5The rectangular pixel region surrounded by the ABCD.
[0092] In the embodiment, the fourth coordinate group of the FOV field of view overlap region is as follows:
[0093]
[0094] In the formula, Axy, Bxy, Cxy, Dxy are the coordinates of the four vertices of the FOV field of view overlap region; (m, n) is the coordinate offset value of the visual FOV center point and the origin of the coordinate system; (p, q) is the coordinate offset value of the infrared FOV center point and the origin of the coordinate system.
[0095] In the embodiment, steps S1 and S3 synchronize and register the original environment data collected by the visual sensor, the infrared sensor and the millimeter wave radar in time and space to ensure the consistency of the data.
[0096] S4, constructing a weight matrix, and then performing feature fusion on the visual features, infrared features and dynamic features of the FOV field of view overlap region to obtain multi-dimensional perception data.
[0097] In the embodiment, the construction of the weight matrix includes:
[0098] A1, obtaining the obstacle data of the FOV field of view overlap region, and establishing a probability distribution model;
[0099] Defining a data model: defining a unified data model to represent data from different sensors. For example, Bayesian estimation and the like are used to estimate the probability density function of the data.
[0100] For example, if image data is processed, each pixel point can be represented as a multi-dimensional vector, including color information and credibility obtained from different sensors.
[0101] A2, calculating the information entropy of the visible light camera, the 4D millimeter wave radar and the infrared camera respectively according to the probability distribution estimated by the probability distribution model;
[0102] The calculation formula of the information entropy is as follows:
[0103]
[0104] In the formula, H(X) is the information entropy of the random variable X, p(x i ) is the probability of X taking the i-th value, and n is the total number of possible values.
[0105] A3, performing normalization processing to convert the information entropy into a confidence weight coefficient;
[0106] The calculation formula of the confidence weight coefficient is as follows:
[0107]
[0108] wherein w i is the weight coefficient of the i-th sensor, H(X i ) is the information entropy of the i-th sensor, and ∈ is a positive number. The ∈ is a very small positive number to ensure numerical stability.
[0109] The confidence of each sensor is normalized to the same scale for comparison and fusion. The maximum normalization or other appropriate methods can be selected as needed.
[0110] A4, integrate all the confidence weight coefficients of the corresponding visual sensor, millimeter wave radar, and infrared sensor to obtain a weight matrix.
[0111] Let the confidence weight coefficients of the visual sensor, millimeter wave radar, and infrared sensor be α, β, and γ, respectively. The weight matrix is:
[0112]
[0113] wherein,
[0114]
[0115] wherein, are the confidence weights of the visual sensor, millimeter wave radar, and infrared sensor, respectively.
[0116] In this embodiment, the obstacle data includes color channel values (R, G, B) of the object image collected by the visual sensor, distance L, speed S, azimuth R, and height signal h of the object collected by the millimeter wave radar, and heat radiation H of the object collected by the infrared sensor. In other embodiments, the visual sensor can collect YUV or RAW format image signals according to actual needs.
[0117] In this embodiment, the feature fusion of the visual features, infrared features, and dynamic features of the FOV field of view overlapping area to obtain multi-dimensional perception data includes:
[0118] B1, based on the FOV field of view overlapping area, respectively acquiring obstacle pixel matrices corresponding to the visual features, infrared features, and dynamic features, point cloud pixel matrices, and heat radiation pixel matrices, as follows:
[0119]
[0120]
[0121] wherein P RGBrepresents the obstacle pixel matrix detected by the visible light camera; P hLSR represents the point cloud pixel matrix detected by the 4D millimeter wave radar. H represents the thermal radiation pixel matrix detected by the infrared camera.
[0122] B2, the obstacle pixel matrix, the point cloud pixel matrix, and the thermal radiation pixel matrix are processed by feature splicing to obtain a multi-dimensional pixel matrix.
[0123] Referring to Figure 6 For each pixel or point, the attributes and corresponding weights in different sensor data are fused. If the data is an image, the color and depth of each pixel can be weighted and averaged. The multi-dimensional pixel matrix corresponding to the field of view overlap area ABCD is as follows:
[0124]
[0125] B3, according to the multi-dimensional pixel matrix and the weight matrix, multi-dimensional pixel fusion is performed to obtain multi-dimensional perception data;
[0126] The fusion formula of the multi-dimensional perception data is as follows:
[0127]
[0128] wherein,
[0129]
[0130] In the formula, P Matrix_A , P Matrix_B , P Matrix_C respectively represent the obstacle attribute features detected by the visual sensor, the obstacle attribute features detected by the millimeter wave radar, and the obstacle attribute features detected by the infrared sensor.
[0131] In this embodiment, it also includes processing boundaries and outliers.
[0132] Specifically, in the data fusion process, boundary problems or outliers may be encountered, which need to be processed. Including but not limited to outlier processing algorithms such as interpolation, filtering, etc.
[0133] In this embodiment, it also includes optimization and post-processing.
[0134] Specifically, the fused data may need further optimization and post-processing to improve its quality and usability. Including but not limited to steps such as denoising, enhancement, correction, etc.
[0135] In this embodiment, the obtained multi-dimensional perception data can be used for target identification and tracking.
[0136] Apply target detection and tracking algorithm on the fused image to identify and track the objects around the vehicle. The multi-dimensional pixel matrix can solve the following typical single sensor cannot cover the corner case:
[0137] Driving at night with high beam of the vehicle in the opposite lane causes the surrounding cyclists to be unable to be accurately perceived, thereby bringing driving risk.
[0138] Accurate perception of the front vehicle target in front of the vehicle on the highway can perceive the intersection information in front, thereby reserving more response time for system decision and execution.
[0139] At the same time, the perception results of different sensors in the multi-dimensional pixel matrix can be set with different confidence levels for different weather conditions and visual field environments, so as to accurately identify the target obstacles in various driving scenes and improve driving safety.
[0140] Embodiment 2
[0141] The embodiment of the application also provides a storage medium having a computer program stored thereon, the computer program being used to implement the multi-dimensional pixel fusion method for environment perception.
[0142] The embodiment of the application has the following beneficial effects:
[0143] (1) The control signals are simultaneously sent to the visual deserializer, the infrared deserializer and the millimeter wave radar, the frame synchronization signal and the hard-wire trigger signal are used to simultaneously drive the visual sensor, the infrared sensor and the millimeter wave radar to collect data, the data is collected at the same time, the timing alignment between the sensors is realized, and the fusion basis for the multi-dimensional data fusion is provided.
[0144] (2) Based on the particularity of the FOV center, the coordinate offset values of the visual FOV center point, the infrared FOV center point and the coordinate system origin (the FOV center point of the millimeter wave radar) are calculated to determine the offset range of the entire visual field, and then the positions of the visual FOV center point and the infrared FOV center point relative to the plane coordinate system are determined based on the coordinate offset values, the data collected by the three sensors is fused in the same coordinate system, the spatial alignment is realized, different types of sensors and different FOV visual fields can find a multi-dimensional pixel visual field overlapping area through software algorithm, and the coordinate space in the area is aligned.
[0145] (3) Based on the characteristics of different types of obstacle data collected by different sensors, a probability distribution model is established according to real-time data, and then information entropy is calculated and converted into confidence weight coefficient, real-time updating the confidence weight coefficient of each sensor, ensuring the actual fitness of the coefficient of obstacle recognition, to improve the recognition accuracy.
[0146] (4) Based on the different types of obstacle data collected by different sensors, the detection data of each sensor is "coordinate unified, time sequence aligned", and the image and radar data are completed pixel-level real-time "spatial-temporal alignment and synchronization" and output in "multi-dimensional pixel" format through splicing and fusion. Provide multi-modal accurate perception information of target and environment for the automatic driving system: including the comprehensive perception of visual data (light and shade, texture, color, etc.), 4D millimeter wave radar data (target distance, speed, direction, height, etc.) and infrared radiation data (texture, temperature, etc.) of the sensor to the target and environment, so as to realize accurate identification of target obstacles in various driving scenes and improve driving safety.
[0147] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application shall be equivalent replacement methods, and all shall be included in the protection scope of the present application.
Claims
1. A multi-dimensional pixel fusion method for environmental perception, characterized in that, include: The sensors are driven by frame synchronization signals and hard-wired trigger signals to collect environmental information and obtain time-aligned raw environmental data. The sensors include a vision sensor, an infrared sensor, and a millimeter-wave radar. The original environmental data is preprocessed, and feature extraction is performed based on the processed original environmental data to obtain the corresponding visual features, infrared features, and dynamic features. Spatial coordinate alignment is performed to obtain the FOV overlap area of spatial registration; A weight matrix is constructed, and then the visual features, infrared features, and dynamic features of the overlapping area of the field of view (FOV) are fused to obtain multidimensional perception data. The process of performing spatial coordinate alignment to obtain the spatially registered FOV (Field of View) overlap region includes: A planar coordinate system is established with the center point of the FOV of the millimeter-wave radar as the origin of the coordinate system; Calculate the coordinate offset values between the visual FOV center point, the infrared FOV center point, and the origin of the coordinate system; Using a planar coordinate system as a reference, the coordinate offset value and sensor resolution are obtained, and the first coordinate group, the second coordinate group, and the third coordinate group corresponding to the radar field of view (FOV) range, visual FOV range, and infrared FOV range are obtained respectively. Based on the first coordinate group, the second coordinate group, and the third coordinate group, a fourth coordinate group is calculated for the corresponding FOV field of view overlap area, so that the coordinate space of the pixels in the FOV field of view overlap area is aligned; the FOV field of view overlap area is the field of view overlap area of radar field of view FOV range, visual FOV range, and infrared FOV range.
2. The multi-dimensional pixel fusion method for environmental perception as described in claim 1, characterized in that, The process of acquiring environmental information based on frame synchronization signals and hard-wired trigger signals to obtain time-aligned raw environmental data includes: Simultaneously, control signals are sent to the vision deserializer, infrared deserializer, and millimeter-wave radar; Exposure control of the visual sensor is performed based on the frame synchronization signal. Visual data is obtained from the visual camera serializer and deserialized by the visual deserializer to obtain visual data. Exposure control of the infrared sensor is performed based on the frame synchronization signal. Infrared data is obtained from the infrared camera serializer and deserialized by the infrared deserializer to obtain infrared data. Data acquisition is performed using a millimeter-wave radar driven by a hard-wired trigger signal to obtain point cloud data. The raw environmental data includes visual data, infrared data, and point cloud data.
3. The multi-dimensional pixel fusion method for environmental perception as described in claim 2, characterized in that, The feature extraction based on the processed original environmental data includes: Visual features are extracted from the visual data based on a deep learning model; infrared features are extracted from the infrared data based on a deep learning model; and dynamic features are extracted from the point cloud data based on optical flow or key point detection algorithms.
4. The multi-dimensional pixel fusion method for environmental perception as described in claim 1, characterized in that, Let the resolution of the millimeter-wave radar be U*V, the resolution of the visual sensor be M*N, and the resolution of the infrared sensor be R*T, then: The fourth coordinate group of the overlapping area of the field of view (FOV) is as follows: In the formula, Axy, Bxy, Cxy, and Dxy are the coordinates of the four vertices of the overlapping area of the FOV field of view; m is the coordinate offset of the visual FOV center point from the origin of the coordinate system on the X-axis; and (p,q) is the coordinate offset of the infrared FOV center point from the origin of the coordinate system.
5. The multi-dimensional pixel fusion method for environmental perception as described in claim 1, characterized in that, The construction of the weight matrix includes: Obstacle data in the overlapping area of the field of view (FOV) is acquired, and a probability distribution model is established. Based on the probability distribution estimated by the probability distribution model, calculate the information entropy of the visible light camera, millimeter-wave radar and infrared camera respectively; Normalization is performed to convert the information entropy into confidence weight coefficients; By integrating all the confidence weight coefficients of the corresponding visual sensors, millimeter-wave radar, and infrared sensors, a weight matrix is obtained.
6. The multi-dimensional pixel fusion method for environmental perception as described in claim 5, characterized in that, The formula for calculating the information entropy is as follows; In the formula, H(X) is the information entropy of the random variable X. Let X be the probability of taking the i-th value, and n be the total number of possible values; The formula for calculating the confidence weight coefficient is as follows: In the formula, Let be the weighting coefficient of the i-th sensor. Let i be the information entropy of the i-th sensor. It is a positive number.
7. The multi-dimensional pixel fusion method for environmental perception as described in claim 5, characterized in that, The feature fusion of visual features, infrared features, and dynamic features of the overlapping area of the field of view (FOV) to obtain multidimensional perception data includes: Based on the FOV field of view overlap area, the obstacle pixel matrix, point cloud pixel matrix, and thermal radiation pixel matrix corresponding to the visual features, infrared features, and dynamic features are obtained respectively. The pixels of the obstacle pixel matrix, point cloud pixel matrix, and thermal radiation pixel matrix are spliced together to obtain a multi-dimensional pixel matrix. Based on the multidimensional pixel matrix and the weight matrix, multidimensional pixel fusion is performed to obtain multidimensional perceptual data; The obstacle data includes: color channel values of the object image acquired by the visual sensor (… The distance L, velocity S, azimuth R, and altitude h of the object are collected by the millimeter-wave radar, and the thermal radiation H of the object is collected by the infrared sensor.
8. The multi-dimensional pixel fusion method for environmental perception as described in claim 7, characterized in that, The multi-dimensional pixel matrix is as follows: The fusion formula for the multidimensional sensing data is as follows: in, In the formula, W is the weight matrix. These are the confidence weighting coefficients for visual sensors, millimeter-wave radar, and infrared sensors, respectively. , , These represent the obstacle attribute characteristics detected by the visual sensor, the obstacle attribute characteristics detected by the millimeter-wave radar, and the obstacle attribute characteristics detected by the infrared sensor, respectively.
9. A storage medium having a computer program stored thereon, characterized in that: The computer program is used to implement a multi-dimensional pixel fusion method for environmental perception as described in any one of claims 1-8.
Citation Information
Patent Citations
Method for improving target detection capability by multi-sensor deep fusion
CN108663677A
Millimeter wave radar and camera combined calibration method
CN114119771A