Millimeter-Wave Radar-Based Visual Fusion Perception Method and System

By using deep learning technology to extract and combine perception analysis of obstacle images in blind spot monitoring methods, the problem of poor recognition effect when obstacles are blocked is solved, and more accurate obstacle recognition and improved driving safety are achieved.

CN119131748BActive Publication Date: 2025-06-10BEIJING ZHONGCHENG KANGFU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411156103.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-06-10
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

The existing blind spot monitoring method based on the perception of millimeter wave radar and vision fusion fails to effectively deal with the situation where obstacles are partially obstructed by other objects, resulting in limited recognition effects.

Method used

The computer vision technology based on deep learning is used to divide the foreground and background of the obstacle images collected by the camera, extract the foreground and background features of the obstacle images respectively, and conduct joint perception analysis of their foreground and background features to fully understand the possible occlusion relationship in traffic scenes, thereby more accurately identifying the obstacle types.

Benefits of technology

It effectively enhances the ability to identify obstacles in complex environments and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119131748B_ABST
    Figure CN119131748B_ABST
Patent Text Reader

Abstract

This application relates to the field of visual fusion perception technology. Specifically, it discloses a visual fusion perception method and system based on a millimeter-wave radar. It acquires obstacle images collected by a camera, and uses computer vision technology based on deep learning to divide the foreground and background of the obstacle images collected by the camera, respectively extract the foreground features and background features of the obstacle images, and through joint perception analysis of its foreground features and background features, fully understand the possible occlusion relationships in the traffic scene, so as to more accurately identify the obstacle types, effectively enhance the ability to identify obstacles in complex environments, and thus improve the safety of obstacle avoidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of visual fusion perception technology, and more specifically, to a method and system for visual fusion perception based on millimeter-wave radar. Background Art

[0002] Millimeter-wave radar is a radar that operates in the millimeter-wave band. Since the wavelength of millimeter waves is between microwaves and centimeter waves, it can distinguish and identify very small targets and can simultaneously identify multiple targets; it has imaging capabilities, small size, good mobility and concealment, and is now widely used in fields such as intelligent driving, intelligent navigation, and intelligent healthcare; and with the development of machine vision technology, the intelligent perception technology of millimeter-wave radar and machine vision fusion has also developed rapidly.

[0003] For example, the invention patent with the publication number CN111060904A proposes a blind spot monitoring method based on millimeter-wave and visual fusion perception. By constructing a multi-modal information fusion model and combining the perception advantages of millimeter-wave radar and camera, it improves the detection and recognition ability of objects in the driver's line-of-sight blind spot. However, in actual traffic scenarios, there may be situations where obstacles are partially blocked by other objects, and this complex situation is not considered in the above blind spot monitoring method, resulting in limited recognition effects of its visual sensors.

[0004] Therefore, an optimized visual fusion perception method is desired. Summary of the Invention

[0005] This application provides a method and system for visual fusion perception based on millimeter-wave radar, which can fully understand the possible occlusion relationships in traffic scenarios, thereby more accurately identifying the types of obstacles.

[0006] In a first aspect, a vision fusion perception method based on millimeter-wave radar is provided, including: calibrating the millimeter-wave radar and the vision sensor respectively, and then performing joint calibration and extrinsic parameter calibration of the two sensors; effectively determining the target based on the millimeter-wave radar; effectively identifying obstacles based on the machine vision sensor; building a fusion model based on the millimeter-wave radar and machine vision; and adopting different alarm methods and information prompt methods according to different types of obstacles. Among them, effectively identifying obstacles based on the machine vision sensor includes: acquiring an obstacle image collected by a camera; respectively extracting the background feature and the foreground feature of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map; respectively performing self-correlation feature enhancement on the obstacle background feature map and the obstacle foreground feature map to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map; inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; and determining the type label of the obstacle based on the foreground-background joint perception feature map.

[0007] In a possible implementation manner, respectively extracting the background feature and the foreground feature of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map includes: inputting the obstacle image into a foreground-background division module to obtain an obstacle background partial image and an obstacle foreground partial image; and respectively performing image feature extraction on the obstacle background partial image and the obstacle foreground partial image to obtain the obstacle background feature map and the obstacle foreground feature map.

[0008] In a possible implementation manner, respectively performing image feature extraction on the obstacle background partial image and the obstacle foreground partial image to obtain the obstacle background feature map and the obstacle foreground feature map includes: inputting the obstacle background partial image into a background feature extractor based on a first dilated convolutional neural network model to obtain the obstacle background feature map; and inputting the obstacle foreground partial image into a foreground feature extractor based on a second dilated convolutional neural network model to obtain the obstacle foreground feature map.

[0009] In a possible implementation manner, respectively performing self-correlation feature enhancement on the obstacle background feature map and the obstacle foreground feature map to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map includes: respectively inputting the obstacle background feature map and the obstacle foreground feature map into a feature space structure consistency self-attention cross-channel enhancement module to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map.

[0010] In a possible implementation, inputting the obstacle background feature map and the obstacle foreground feature map into a feature space structure consistency self-attention cross-channel enhancement module respectively to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map includes: performing layer normalization on the obstacle background feature map to obtain a normalized obstacle background feature map; performing point convolution processing on the normalized obstacle background feature map to obtain an obstacle background channel context correlation representation feature map; performing convolutional encoding on the obstacle background channel context correlation representation feature map to obtain an obstacle background spatial context correlation representation feature map; performing channel-space global interaction attention fusion on the obstacle background channel context correlation representation feature map and the obstacle background spatial context correlation representation feature map to obtain the enhanced obstacle background feature map.

[0011] In a possible implementation, performing channel-space global interaction attention fusion on the obstacle background channel context correlation representation feature map and the obstacle background spatial context correlation representation feature map to obtain the enhanced obstacle background feature map includes: copying the obstacle background spatial context correlation representation feature map to obtain a backup obstacle background spatial context correlation representation feature map; reshaping the feature shapes of the obstacle background channel context correlation representation feature map, the obstacle background spatial context correlation representation feature map, and the backup obstacle background spatial context correlation representation feature map to obtain an obstacle background channel context correlation representation feature matrix, an obstacle background spatial context correlation representation feature matrix, and a backup obstacle background spatial context correlation representation feature matrix; calculating the cross-channel cross-covariance matrix between the obstacle background channel context correlation representation feature matrix and the obstacle background spatial context correlation representation feature matrix; activating the cross-channel cross-covariance matrix using the Softmax function to obtain an obstacle background feature global interaction attention matrix; calculating the product between the backup obstacle background spatial context correlation representation feature matrix and the obstacle background feature global interaction attention matrix to obtain an attention-enhanced obstacle background feature representation matrix; reshaping the feature shape of the attention-enhanced obstacle background feature representation matrix to obtain the enhanced obstacle background feature map.

[0012] In a possible implementation, inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map includes: reshaping the feature shapes of the enhanced obstacle background feature map and the enhanced obstacle foreground feature map to obtain an enhanced obstacle background feature vector and an enhanced obstacle foreground feature vector; performing correlation encoding on the enhanced obstacle background feature vector and the enhanced obstacle foreground feature vector to obtain a foreground-background dependency matrix; performing linear interpolation on the foreground-background dependency matrix to obtain a dimension-adjusted foreground-background dependency matrix; multiplying the dimension-adjusted foreground-background dependency matrix by the enhanced obstacle foreground feature vector to obtain a foreground-background joint perception feature vector; and reshaping the feature shape of the foreground-background joint perception feature vector to obtain the foreground-background joint perception feature map.

[0013] In a possible implementation, based on the foreground-background joint perception feature map, determining the type label of the obstacle includes: inputting the foreground-background joint perception feature map into an obstacle recognizer based on a classifier to obtain an identification result, and the identification result is used to represent the type label of the obstacle.

[0014] In a second aspect, a vision fusion perception system based on a millimeter-wave radar is provided, including: an obstacle image acquisition module for acquiring an obstacle image collected by a camera; a foreground-background feature extraction module for respectively extracting the background feature and the foreground feature of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map; a self-correlation feature enhancement module for respectively performing self-correlation feature enhancement on the obstacle background feature map and the obstacle foreground feature map to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map; a foreground-background joint perception module for inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; and an obstacle type determination module for determining the type label of the obstacle based on the foreground-background joint perception feature map.

[0015] In a possible implementation, the feature extraction module includes: a foreground-background partitioning unit for inputting the obstacle image into a foreground-background partitioning module to obtain an obstacle background partial image and an obstacle foreground partial image; and an image feature extraction unit for respectively performing image feature extraction on the obstacle background partial image and the obstacle foreground partial image to obtain the obstacle background feature map and the obstacle foreground feature map.

[0016] In a third aspect, a chip is provided, which includes an input / output interface, at least one processor, at least one memory, and a bus. The at least one memory is used to store instructions, and the at least one processor is used to call the instructions in the at least one memory to execute the method in the first aspect.

[0017] In a fourth aspect, a computer-readable medium is provided for storing a computer program, where the computer program includes the method for executing the above-mentioned first aspect.

[0018] In a fifth aspect, a computer program product including instructions is provided. When a computer runs the instructions of the computer program product, the computer executes the method in the above-mentioned second aspect.

[0019] The present application has at least the following technical effects: Compared with the prior art, a vision fusion perception method and system based on a millimeter-wave radar provided by the present application acquires an obstacle image collected by a camera, and uses computer vision technology based on deep learning to divide the foreground and background of the obstacle image collected by the camera, respectively extract the foreground features and background features of the obstacle image, and through joint perception analysis of its foreground features and background features, fully understand the possible occlusion relationships in the traffic scene, so as to more accurately identify the type of obstacle, which can effectively enhance the ability to identify obstacles in a complex environment, thereby improving driving safety. Description of the Drawings

[0020] Figure 1 It is a schematic flowchart of the effective recognition of obstacles based on a machine vision sensor in the vision fusion perception method based on a millimeter-wave radar according to an embodiment of the present application.

[0021] Figure 2 It is a schematic data flow diagram of the effective recognition of obstacles based on a machine vision sensor in the vision fusion perception method based on a millimeter-wave radar according to an embodiment of the present application.

[0022] Figure 3 It is a schematic flowchart of respectively extracting the background features and foreground features of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map in the vision fusion perception method based on a millimeter-wave radar according to an embodiment of the present application.

[0023] Figure 4 It is a schematic flowchart of respectively inputting the obstacle background feature map and the obstacle foreground feature map into a feature space structure consistency self-attention cross-channel enhancement module to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map in the vision fusion perception method based on a millimeter-wave radar according to an embodiment of the present application.

[0024] Figure 5In the method for vision fusion perception based on millimeter-wave radar according to the embodiments of the present application, a schematic flowchart for obtaining the enhanced obstacle background feature map by performing channel-space global interaction attention fusion on the obstacle background channel context association representation feature map and the obstacle background spatial context association representation feature map.

[0025] Figure 6 In the method for vision fusion perception based on millimeter-wave radar according to the embodiments of the present application, a schematic flowchart for inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map.

[0026] Figure 7 A schematic block diagram of a vision fusion perception system based on millimeter-wave radar according to the embodiments of the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts also belong to the scope of protection of the present application.

[0028] Considering that in actual traffic scenarios, there may be situations where obstacles are partially blocked by other objects, and such complex situations are not considered in the above blind area monitoring method, resulting in limited recognition effects of its vision sensors. To address the above technical problems, the technical concept of the present application is to use computer vision technology based on deep learning to divide the foreground and background of the obstacle images collected by the camera, extract the foreground features and background features of the obstacle images respectively, and through joint perception and analysis of their foreground features and background features, fully understand the possible occlusion relationships in the traffic scenario, so as to more accurately identify the obstacle types, effectively enhance the recognition ability of obstacles in complex environments, and thus improve driving safety.

[0029] Based on this, the present application provides a vision fusion perception method based on millimeter-wave radar, including: respectively calibrating the millimeter-wave radar and the vision sensor, and then performing joint calibration and external parameter calibration of the two sensors; effectively determining the target based on the millimeter-wave radar; effectively identifying the obstacle based on the machine vision sensor; building a fusion model of the millimeter-wave radar and the machine vision; and adopting different alarm methods and information prompt methods according to different types of obstacles. In the technical solution of the present application, Figure 1 In the method for vision fusion perception based on millimeter-wave radar according to the embodiments of the present application, a schematic flowchart for effectively identifying the obstacle based on the machine vision sensor. Figure 2In the method for vision fusion perception based on millimeter-wave radar according to the embodiments of the present application, it is a schematic diagram of data flow based on the effective recognition of obstacles by a machine vision sensor. As Figure 1 and Figure 2 shown, based on the effective recognition of obstacles by a machine vision sensor, it includes: S110, obtaining an obstacle image collected by a camera; S120, respectively extracting the background feature and the foreground feature of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map; S130, respectively performing self-correlation feature enhancement on the obstacle background feature map and the obstacle foreground feature map to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map; S140, inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; S150, based on the foreground-background joint perception feature map, determining the type label of the obstacle.

[0030] In the above method for vision fusion perception based on millimeter-wave radar, in S110, an obstacle image collected by a camera is obtained. Specifically, as a vision sensor, the camera can provide rich visual information such as color, texture, and shape. This information is crucial for understanding the appearance characteristics of obstacles and helps to more accurately identify and classify obstacles. By obtaining the obstacle image, further computer vision techniques such as deep learning can be used to process the image to identify the obstacles in the image, even if they may be partially blocked by other objects. By analyzing the image collected by the camera, the system can more accurately identify the type of obstacle. This is very important for improving driving safety and realizing functions of intelligent transportation systems such as collision avoidance.

[0031] In the above method for vision fusion perception based on millimeter-wave radar, in S120, the background feature and the foreground feature of the obstacle image are respectively extracted to obtain an obstacle background feature map and an obstacle foreground feature map. Specifically, in a traffic scene, an obstacle may be partially or completely blocked by other objects. By respectively extracting the foreground and background features, the details and environmental context information of the obstacle can be captured more precisely. The foreground feature usually contains direct information of the obstacle, while the background feature provides information about the environment where the obstacle is located.

[0032] Optionally, in an embodiment of the present application, Figure 3 In the method for vision fusion perception based on millimeter-wave radar according to the embodiments of the present application, it is a schematic flowchart of respectively extracting the background feature and the foreground feature of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map. As Figure 3As shown, the background features and foreground features of the obstacle image are respectively extracted to obtain an obstacle background feature map and an obstacle foreground feature map, including: S121, inputting the obstacle image into a foreground-background division module to obtain an obstacle background partial image and an obstacle foreground partial image; S122, respectively performing image feature extraction on the obstacle background partial image and the obstacle foreground partial image to obtain the obstacle background feature map and the obstacle foreground feature map.

[0033] In the above-mentioned vision fusion perception method based on a millimeter-wave radar, in S121, the obstacle image is input into a foreground-background division module to obtain an obstacle background partial image and an obstacle foreground partial image. Specifically, by dividing the obstacle image into a foreground part and a background part, the edge information of the obstacle can be highlighted, which helps to identify the loss of boundary information caused by occlusion, so as to better analyze the possible occlusion relationship and provide more environmental context information for understanding the relative position and shape of the obstacle. At the same time, by clearly distinguishing the foreground and the background, the interference of the other part of the information can also be removed in the subsequent feature extraction process, improving the discrimination of the feature representation.

[0034] In the above-mentioned vision fusion perception method based on a millimeter-wave radar, in S122, image feature extraction is respectively performed on the obstacle background partial image and the obstacle foreground partial image to obtain the obstacle background feature map and the obstacle foreground feature map. Specifically, in order to effectively capture the feature representations of the obstacle background partial image and the obstacle foreground partial image, the present application uses a first dilated convolutional neural network model and a second dilated convolutional neural network model to respectively perform image feature extraction on the obstacle background partial image and the obstacle foreground partial image, so as to utilize the characteristics of dilated convolution to effectively capture the long-range context information in the image, while maintaining a high feature resolution, improving the extraction accuracy of background features and foreground features, enhancing the expression ability of features, and thus generating an obstacle background feature map and an obstacle foreground feature map.

[0035] Optionally, in an embodiment of the present application, performing image feature extraction on the obstacle background partial image and the obstacle foreground partial image respectively to obtain the obstacle background feature map and the obstacle foreground feature map includes: inputting the obstacle background partial image into a background feature extractor based on a first dilated convolutional neural network model to obtain the obstacle background feature map; inputting the obstacle foreground partial image into a foreground feature extractor based on a second dilated convolutional neural network model to obtain the obstacle foreground feature map.

[0036] In the above-mentioned vision fusion perception method based on millimeter-wave radar, in S130, autocorrelation feature enhancement is respectively performed on the obstacle background feature map and the obstacle foreground feature map to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map. Specifically, in order to further enhance the expression ability of the background feature and the foreground feature, this application introduces a feature space structure consistency self-attention cross-channel enhancement module to process the obstacle background feature map and the obstacle foreground feature map respectively. Among them, the feature space structure consistency self-attention cross-channel enhancement module can utilize the channel interaction of the feature map to capture its channel context correlation information, and at the same time retain the spatial structure of the feature map to balance the robustness and richness of feature expression. Specifically, first, layer normalization processing is performed on the input feature map to eliminate the scale difference between layers and ensure the stability of feature processing. Then, through multi-layer convolution operations, the channel context correlation and spatial structure information of the feature map are captured, and its channel context correlation information is used as a query, and the spatial structure information is used as a key and a value, and feature cross-channel global correlation interaction is performed based on the self-attention mechanism. While retaining the spatial structure of the feature map, the channel correlation information between feature structures is fused in units of the feature space structure to enhance the expression ability of the feature map, thereby obtaining an enhanced obstacle background feature map and an enhanced obstacle foreground feature map. That is, performing autocorrelation feature enhancement on the obstacle background feature map and the obstacle foreground feature map respectively to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map includes: respectively inputting the obstacle background feature map and the obstacle foreground feature map into the feature space structure consistency self-attention cross-channel enhancement module to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map.

[0037] Optionally, in an embodiment of this application, Figure 4 In the vision fusion perception method based on millimeter-wave radar according to the embodiment of this application, the schematic flowchart of respectively inputting the obstacle background feature map and the obstacle foreground feature map into the feature space structure consistency self-attention cross-channel enhancement module to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map is as Figure 4As shown, inputting the obstacle background feature map and the obstacle foreground feature map into the feature space structure consistency self-attention cross-channel enhancement module respectively to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map includes: S131, performing layer normalization on the obstacle background feature map to obtain a normalized obstacle background feature map; S132, performing point convolution processing on the normalized obstacle background feature map to obtain an obstacle background channel context correlation representation feature map; S133, performing convolutional encoding on the obstacle background channel context correlation representation feature map to obtain an obstacle background spatial context correlation representation feature map; S134, performing channel-space global interaction attention fusion on the obstacle background channel context correlation representation feature map and the obstacle background spatial context correlation representation feature map to obtain the enhanced obstacle background feature map.

[0038] Optionally, in an embodiment of the present application, Figure 5 In the vision fusion perception method based on millimeter-wave radar according to the embodiment of the present application, the schematic flowchart of performing channel-space global interaction attention fusion on the obstacle background channel context correlation representation feature map and the obstacle background spatial context correlation representation feature map to obtain the enhanced obstacle background feature map is as follows. As Figure 5 shown, performing channel-space global interaction attention fusion on the obstacle background channel context correlation representation feature map and the obstacle background spatial context correlation representation feature map to obtain the enhanced obstacle background feature map includes: S1341, copying the obstacle background spatial context correlation representation feature map to obtain a backup obstacle background spatial context correlation representation feature map; S1342, performing feature shape reshaping on the obstacle background channel context correlation representation feature map, the obstacle background spatial context correlation representation feature map, and the backup obstacle background spatial context correlation representation feature map to obtain an obstacle background channel context correlation representation feature matrix, an obstacle background spatial context correlation representation feature matrix, and a backup obstacle background spatial context correlation representation feature matrix; S1343, calculating the cross-channel cross-covariance matrix between the obstacle background channel context correlation representation feature matrix and the obstacle background spatial context correlation representation feature matrix; S1344, using the Softmax function to activate the cross-channel cross-covariance matrix to obtain an obstacle background feature global interaction attention matrix; S1345, calculating the product between the backup obstacle background spatial context correlation representation feature matrix and the obstacle background feature global interaction attention matrix to obtain an attention-enhanced obstacle background feature representation matrix; S1346, performing feature shape reshaping on the attention-enhanced obstacle background feature representation matrix to obtain the enhanced obstacle background feature map.

[0039] More specifically, in an embodiment of the present application, the formula for separately performing self - correlation feature enhancement on the obstacle background feature map and the obstacle foreground feature map to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map is as follows: The following self - correlation attention enhancement formula is used to process the obstacle background feature map to obtain the enhanced obstacle background feature map, where the self - correlation attention enhancement formula is:

[0040] F ln = Layer Normalization(F i )

[0041] F Q = Conv 1×1 (F ln )

[0042] F K = Conv 3×3 (Conv 1×1 (F ln ))

[0043] F V = Copy(F K )

[0044] M Q = Reshape(F Q )

[0045] M K = Reshape(F K )

[0046] M V = Reshape(F V )

[0047]

[0048] Among them, F i represents the obstacle background feature map, Layer Normalization(·) represents the layer normalization operation, F ln represents the normalized obstacle background feature map, Conv 1×1 represents point convolution, F Q represents the obstacle background channel context - associated representation feature map, Conv 3×3 represents convolution processing based on a 3×3 convolution kernel, F K represents the obstacle background spatial context - associated representation feature map, Copy(·) represents the copy operation, F V represents the backup obstacle background spatial context - associated representation feature map, reshape(·) represents feature shape reshaping, MQ , M K and M V respectively represent the obstacle background channel context - associated representation feature matrix, the obstacle background spatial context - associated representation feature matrix, and the backup obstacle background spatial context - associated representation feature matrix. M a represents the cross - channel cross - covariance matrix, θ is a scaling factor, and softmax represents the normalized exponential function. represents matrix multiplication operation, F a represents the enhanced obstacle background feature map. Here, specifically, the self - correlation feature enhancement of the obstacle foreground feature map to obtain the enhanced obstacle foreground feature map also refers to the above formula.

[0049] In the above - mentioned millimeter - wave radar - based visual fusion perception method, in S140, the enhanced obstacle background feature map and the enhanced obstacle foreground feature map are input into the foreground - background joint perception module to obtain the foreground - background joint perception feature map. Specifically, the enhanced obstacle background feature map provides global information about the traffic environment, road layout, etc., while the enhanced obstacle foreground feature map focuses more on describing the specific detail features of the obstacles. To achieve the deep fusion of background features and foreground features in the feature space to enhance the expression of obstacle features in the case of occlusion, this application further introduces a foreground - background joint perception module to jointly encode the enhanced obstacle background feature map and the enhanced obstacle foreground feature map. Specifically, the foreground - background joint perception module adaptively adjusts the foreground features based on the dependence relationship between the background features and the foreground features to enhance the context relevance between the foreground features and the background features, and generates the foreground - background joint perception feature map. That is, the background information is used to supplement and correct the obstacle foreground features, assisting in identifying the edges and details of the obstacles, so as to more accurately understand the shape and relative position of the obstacles.

[0050] Optionally, in an embodiment of the present application, Figure 6In the vision fusion perception method based on millimeter-wave radar according to the embodiments of the present application, the schematic flowchart of inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into the foreground-background joint perception module to obtain the foreground-background joint perception feature map is as follows. Inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into the foreground-background joint perception module to obtain the foreground-background joint perception feature map includes: S141, reshaping the feature shapes of the enhanced obstacle background feature map and the enhanced obstacle foreground feature map to obtain an enhanced obstacle background feature vector and an enhanced obstacle foreground feature vector; S142, performing correlation encoding on the enhanced obstacle background feature vector and the enhanced obstacle foreground feature vector to obtain a foreground-background dependence relationship matrix; S143, performing linear interpolation on the foreground-background dependence relationship matrix to obtain a dimension-adjusted foreground-background dependence relationship matrix; S144, multiplying the dimension-adjusted foreground-background dependence relationship matrix by the enhanced obstacle foreground feature vector to obtain a foreground-background joint perception feature vector; S145, reshaping the feature shape of the foreground-background joint perception feature vector to obtain the foreground-background joint perception feature map.

[0051] In the above vision fusion perception method based on millimeter-wave radar, in S150, based on the foreground-background joint perception feature map, the type label of the obstacle is determined. Specifically, the foreground-background joint perception feature map is obtained by deeply analyzing and fusing the foreground features and background features of the obstacle. It synthesizes the local details of the obstacle and the global information of the surrounding environment, providing rich context data for obstacle recognition. Optionally, in an embodiment of the present application, determining the type label of the obstacle based on the foreground-background joint perception feature map includes: inputting the foreground-background joint perception feature map into an obstacle recognizer based on a classifier to obtain a recognition result, and the recognition result is used to represent the type label of the obstacle. Specifically, through training, the obstacle recognizer based on the classifier can effectively learn the obstacle feature representation in the foreground-background joint perception feature map and map it to a predefined obstacle category space based on the learned feature representation, thereby achieving accurate classification of the obstacle.

[0052] In a preferred embodiment, obtaining an identification result by passing the foreground-background joint perception feature map through a classifier-based obstacle recognizer includes: calculating a feature mean of the foreground-background joint perception feature map, and dividing the feature mean by a difference between a maximum feature value and a minimum feature value of the foreground-background joint perception feature map to obtain a foreground-background joint perception distribution representation value; dividing one minus the foreground-background joint perception distribution representation value by the foreground-background joint perception distribution representation value to obtain a foreground-background joint perception distribution modulation value; activating the foreground-background joint perception feature map through a probability function to obtain a probability-based foreground-background joint perception feature map; subtracting the probability-based foreground-background joint perception feature map from the foreground-background joint perception distribution modulation value, taking an absolute value, and calculating the negative of the logarithm to the base 2 to obtain a probability-based foreground-background joint perception distribution modulation information feature map; dividing the foreground-background joint perception distribution representation value by a difference between one and each feature value of the probability-based foreground-background joint perception feature map, summing all feature values of the probability-based foreground-background joint perception feature map, and dividing by a scale of the foreground-background joint perception feature map to obtain a probability-based foreground-background joint perception distribution modulation bias value; adding the probability-based foreground-background joint perception distribution modulation information feature map and a product of the probability-based foreground-background joint perception distribution modulation bias value and a weight as a hyperparameter through dot product to obtain an optimized foreground-background joint perception feature map; and inputting the optimized foreground-background joint perception feature map through the classifier-based obstacle recognizer to obtain the identification result.

[0053] Wherein, the optimized foreground-background joint perception feature map is expressed as:

[0054]

[0055] And wherein:

[0056]

[0057] Wherein, represents the feature mean of the foreground-background joint perception feature map, f max and f min respectively represent the maximum feature value and the minimum feature value in the foreground-background joint perception feature map, p represents the foreground-background joint perception distribution representation value, F represents the probability-based foreground-background joint perception feature map obtained by activating the foreground-background joint perception feature map through a probability function, f iRepresents the i-th eigenvalue of the probabilistic foreground-background joint perception feature map, ε is the weight as a hyperparameter, S is the scale of the foreground-background joint perception feature map, that is, the width of the feature matrix of the foreground-background joint perception feature map multiplied by the height and then multiplied by the number of channels of the foreground-background joint perception feature map, and F′ represents the optimized foreground-background joint perception feature map.

[0058] Specifically, considering that the enhanced obstacle background feature map and the enhanced obstacle foreground feature map respectively express the image semantic features of the background part and the foreground part in the obstacle image after the cross-channel enhancement of the feature space structure consistency self-attention, thus, when the enhanced obstacle background feature map and the enhanced obstacle foreground feature map are input into the foreground-background joint perception module, the obtained foreground-background joint perception feature map will also have insufficient foreground-background joint perception coverage due to the attention interleaving weight differences in the image spatial dimension and the channel dimension, resulting in an outlier class inference mapping deviation, which affects the accuracy of the recognition result obtained by the foreground-background joint perception feature map through the classifier-based obstacle recognizer.

[0059] Therefore, in the above-mentioned preferred example, through the Bernoulli probability modulation distribution of the foreground-background joint perception feature map with respect to the eigenvalue distribution, the eigenvalue-based probability information distribution planning of the foreground-background joint perception feature map is carried out, and the probability inverse mapping of the overall probability feature of the foreground-background joint perception feature map is used as the extended coverage of the set mapping space of the foreground-background joint perception feature map, so as to independently understand the intuitive probability information distribution and the abstract probability space mapping of the foreground-background joint perception feature map for the interaction path therebetween, in order to improve the accuracy of the recognition result obtained by the optimized foreground-background joint perception representation vector through the classifier-based obstacle recognizer by avoiding the counterfactual inference mapping from the outlier feature distribution of the foreground-background joint perception feature map to the class regression probability.

[0060] In summary, the millimeter-wave radar-based visual fusion perception method according to the embodiments of the present application is clarified. In the step of effectively recognizing obstacles based on a machine vision sensor, an obstacle image collected by a camera is obtained, and a computer vision technology based on deep learning is used to divide the foreground and background of the obstacle image collected by the camera, extract the foreground feature and the background feature of the obstacle image respectively, and through the joint perception analysis of its foreground feature and background feature, fully understand the possible occlusion relationships in the traffic scene, so as to more accurately identify the obstacle type. In this way, the recognition ability of obstacles in a complex environment can be effectively enhanced, thereby improving driving safety.

[0061] Figure 7 It is a schematic block diagram of the millimeter-wave radar-based visual fusion perception system according to the embodiments of the present application. AsFigure 7 As shown in Figure 7 , the millimeter-wave radar-based vision fusion perception system 200 according to an embodiment of the present application includes: an obstacle image acquisition module 210, configured to acquire an obstacle image collected by a camera; a foreground-background feature extraction module 220, configured to extract the background feature and the foreground feature of the obstacle image respectively to obtain an obstacle background feature map and an obstacle foreground feature map; an autocorrelation feature enhancement module 230, configured to perform autocorrelation feature enhancement on the obstacle background feature map and the obstacle foreground feature map respectively to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map; a foreground-background joint perception module 240, configured to input the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; and an obstacle type determination module 250, configured to determine the type label of the obstacle based on the foreground-background joint perception feature map.

[0062] Optionally, in an embodiment of the present application, the feature extraction module includes: a foreground-background division unit, configured to input the obstacle image into a foreground-background division module to obtain an obstacle background partial image and an obstacle foreground partial image; and an image feature extraction unit, configured to perform image feature extraction on the obstacle background partial image and the obstacle foreground partial image respectively to obtain the obstacle background feature map and the obstacle foreground feature map.

[0063] The specific operations of the above steps in the millimeter-wave radar-based vision fusion perception system have been introduced in detail in the description of the millimeter-wave radar-based vision fusion perception method above, and therefore, the repeated description thereof will be omitted. Figures 1 to 6 of the millimeter-wave radar-based vision fusion perception method, and thus, the repeated description thereof will be omitted.

[0064] An embodiment of the present invention further provides a chip system, which includes at least one processor. When program instructions are executed in the at least one processor, the method provided by the embodiment of the present application is implemented.

[0065] An embodiment of the present invention further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer executes the method of the above method embodiment.

[0066] An embodiment of the present invention further provides a computer program product including instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.

[0067] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0068] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0069] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0070] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0071] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit.

Claims

1. A visual fusion perception method based on millimeter wave radar, comprising: The millimeter wave radar and visual sensor are calibrated separately, and then the two sensors are calibrated jointly and the external parameters are calibrated; Effective determination of targets based on millimeter-wave radar; effective identification of obstacles based on machine vision sensors; construction of a fusion model of millimeter-wave radar and machine vision; different alarm methods and information prompt methods are adopted according to different types of obstacles. The characteristics are that effective identification of obstacles based on machine vision sensors include: Obtaining obstacle images captured by a camera; Extracting background features and foreground features of the obstacle image respectively to obtain an obstacle background feature map and an obstacle foreground feature map; The obstacle background feature map and the obstacle foreground feature map are respectively enhanced by autocorrelation features to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map, which includes: The obstacle background feature map and the obstacle foreground feature map are respectively input into the feature space structure consistency self-attention cross-channel enhancement module to obtain the enhanced obstacle background feature map and the enhanced obstacle foreground feature map, specifically: Performing layer normalization on the obstacle background feature map to obtain a normalized obstacle background feature map; Performing point convolution processing on the normalized obstacle background feature map to obtain an obstacle background channel context association representation feature map; Performing convolution encoding on the obstacle background channel context association representation feature map to obtain an obstacle background space context association representation feature map; The obstacle background channel context association representation feature map and the obstacle background space context association representation feature map are subjected to channel-space global interactive attention fusion to obtain the enhanced obstacle background feature map, specifically: Copying the obstacle background spatial context association representation feature map to obtain a backup obstacle background spatial context association representation feature map; Reshaping the obstacle background channel context association representation feature map, the obstacle background space context association representation feature map and the backup obstacle background space context association representation feature map to obtain an obstacle background channel context association representation feature matrix, an obstacle background space context association representation feature matrix and a backup obstacle background space context association representation feature matrix; Calculating a cross-channel cross covariance matrix between the obstacle background channel context association representation feature matrix and the obstacle background spatial context association representation feature matrix; Use the Softmax function to activate the cross-channel cross covariance matrix to obtain the obstacle background feature global interaction attention matrix; Calculating the product of the backup obstacle background spatial context association representation feature matrix and the obstacle background feature global interactive attention matrix to obtain an attention-enhanced obstacle background feature representation matrix; Reshaping the attention-enhanced obstacle background feature representation matrix to obtain the enhanced obstacle background feature map; Inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; Based on the foreground-background joint perception feature map, a type label of the obstacle is determined.

2. The visual fusion perception method based on millimeter wave radar according to claim 1 is characterized in that: Extracting background features and foreground features of the obstacle image respectively to obtain an obstacle background feature map and an obstacle foreground feature map, including: Inputting the obstacle image into a foreground-background segmentation module to obtain an obstacle background portion image and an obstacle foreground portion image; Image features are extracted from the obstacle background portion image and the obstacle foreground portion image respectively to obtain the obstacle background feature map and the obstacle foreground feature map.

3. The visual fusion perception method based on millimeter wave radar according to claim 2 is characterized in that: The image features of the obstacle background portion image and the obstacle foreground portion image are respectively extracted to obtain the obstacle background feature map and the obstacle foreground feature map, including: Inputting the obstacle background partial image into a background feature extractor based on a first hole convolutional neural network model to obtain the obstacle background feature map; The obstacle foreground partial image is input into a foreground feature extractor based on a second hole convolutional neural network model to obtain the obstacle foreground feature map.

4. The visual fusion perception method based on millimeter wave radar according to claim 3 is characterized in that: Inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map, including: Reshaping the enhanced obstacle background feature map and the enhanced obstacle foreground feature map to obtain an enhanced obstacle background feature vector and an enhanced obstacle foreground feature vector; Performing associative coding on the enhanced obstacle background feature vector and the enhanced obstacle foreground feature vector to obtain a foreground-background dependency matrix; Performing linear interpolation on the foreground-background dependency matrix to obtain a dimensionally adjusted foreground-background dependency matrix; Performing matrix multiplication of the dimensionally adjusted foreground-background dependency matrix and the enhanced obstacle foreground feature vector to obtain a foreground-background joint perception feature vector; The foreground-background joint perception feature vector is reshaped to obtain the foreground-background joint perception feature map.

5. The visual fusion perception method based on millimeter wave radar according to claim 4 is characterized in that: Determining the type label of the obstacle based on the foreground-background joint perception feature map includes: The foreground-background joint perception feature map is input into a classifier-based obstacle identifier to obtain a recognition result, and the recognition result is used to represent a type label of the obstacle.

6. A visual fusion perception system based on millimeter wave radar, used to execute the visual fusion perception method based on millimeter wave radar according to any one of claims 1 to 5, characterized in that: include: The obstacle image acquisition module is used to acquire the obstacle image collected by the camera; A foreground and background feature extraction module, used to extract the background features and foreground features of the obstacle image to obtain an obstacle background feature map and an obstacle foreground feature map; An autocorrelation feature enhancement module, used to perform autocorrelation feature enhancement on the obstacle background feature map and the obstacle foreground feature map respectively to obtain an enhanced obstacle background feature map and an enhanced obstacle foreground feature map; A foreground-background joint perception module, used for inputting the enhanced obstacle background feature map and the enhanced obstacle foreground feature map into a foreground-background joint perception module to obtain a foreground-background joint perception feature map; The obstacle type determination module is used to determine the type label of the obstacle based on the foreground and background joint perception feature map.

7. The visual fusion perception system based on millimeter wave radar according to claim 6 is characterized in that: The feature extraction module comprises: A foreground-background segmentation unit, used for inputting the obstacle image into a foreground-background segmentation module to obtain an obstacle background portion image and an obstacle foreground portion image; The image feature extraction unit is used to extract image features from the obstacle background portion image and the obstacle foreground portion image respectively to obtain the obstacle background feature map and the obstacle foreground feature map.

Citation Information

Patent Citations

  • Blind area monitoring method based on millimeter wave and visual fusion perception

    CN111060904A

  • Optical remote sensing image salient target detection method of double-flow decoding cross-task interaction network

    CN113505634A