Method, apparatus and storage medium for mobile target recognition based on radar-vision fusion

Through the lightning vision fusion method, millimeter-wave radar reconstruction and visual feature fusion are used to solve the problem of low light and long-distance object detection accuracy in intelligent transportation systems, and the detection accuracy and robustness are improved.

CN116129257BActive Publication Date: 2025-07-08GUANGXI COMPREHENSIVE TRANSPORTATION BIG DATA RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211591325.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-07-08
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

The existing intelligent transportation system has low target detection accuracy under low light and long distance conditions, the camera is not effective in low light, and the millimeter-wave radar cannot obtain object texture information, resulting in a reduced detection accuracy.

Method used

The lightning fusion method is used to characterize the target spatial position through millimeter wave radar, reconstruct the image and detect it, and combine the radar and visual features to fusion the detection accuracy.

Benefits of technology

Improve object detection accuracy under low light and long distance conditions, enhancing system robustness and detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129257B_ABST
    Figure CN116129257B_ABST
Patent Text Reader

Abstract

In the method, device and storage medium for mobile target recognition based on radar-vision fusion provided by the embodiments of the present application, first, starting from the data level, the spatial position of potential targets is characterized according to millimeter-wave radar, and the characterization result is used for dividing the long-distance target area in the original image collected by the camera. Further, the divided area image is reconstructed and detected, so as to improve the visual detection accuracy of long-distance targets. Then, modeling is carried out by fusing the radar and visual detection feature layers. In view of the characteristics that millimeter-wave radar detection is less affected by illumination and image detection has more texture information, the detection accuracy of the system in low-light environments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of machine vision. More specifically, embodiments of the present application relate to a method, device, and storage medium for mobile target recognition based on the fusion of lidar and vision. Background Art

[0002] This section aims to provide background or context for the embodiments of the present application recited in the claims. The description herein is not admitted to be prior art merely by virtue of being included in this section.

[0003] Driven by the demand for intelligent traffic management, higher requirements are also put forward for target detection technologies for vehicles, ships and other means of transportation.

[0004] In related technologies, cameras, millimeter-wave radars, etc. are usually used to detect means of transportation. Among them, the images captured by the camera can provide detailed texture information of objects and the environment, but the detection effect is not good in low-light environments; while the millimeter-wave radar is not affected by light, has strong anti-interference ability and long detection range, but cannot obtain the texture information of objects, resulting in a significant reduction in detection accuracy.

[0005] Therefore, target detection at long distances and in low light is a huge challenge in intelligent transportation systems. Summary of the Invention

[0006] Embodiments of the present application provide a method, device, and storage medium for mobile target recognition based on the fusion of lidar and vision, which are used to solve the technical problem of low accuracy of target detection in current intelligent transportation systems.

[0007] In the first aspect of the embodiments of the present application, a method for mobile target recognition based on the fusion of lidar and vision is provided, including: reconstructing an original image according to the detection result of a millimeter-wave radar for a target to be detected to obtain a to-be-detected image, where the to-be-detected image includes the original image and at least one target detection area, and the target detection area includes at least one target to be detected; inputting the to-be-detected image into a lidar-vision detection network, and detecting the target to be detected in the target detection area through the lidar-vision detection network to obtain a feature image output by the lidar-vision detection network, where the feature image includes a target detection result for recognizing the target to be detected; and performing reduction processing on the feature image according to the target detection area corresponding to the feature image and the original image to obtain a target image, where the area where the target to be detected is located in the target image is the same as the area where the target to be detected is located in the original image.

[0008] In some alternative embodiments, the original image is reconstructed based on the detection result of the millimeter-wave radar for the target to be detected to obtain the image to be detected, including: determining the second size information of the target detection region in the image to be detected according to the first size information of the original image, the number of target detection regions in the original image, and the preset distance between the target detection regions; determining the fourth size information of the original image in the image to be detected according to the third size information of the original image and the second size information of the target detection region; determining the first coordinate information of the target detection region in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second size information of the target detection region and the fourth size information of the original image in the image to be detected; and obtaining the image to be detected according to the first coordinate information and the second coordinate information.

[0009] In some alternative embodiments, determining the second size information of the target detection region in the image to be detected according to the first size information of the original image, the number of target detection regions in the original image, and the preset distance between the target detection regions includes: determining the second size information of the target detection region in the image to be detected according to the following formula:

[0010]

[0011] where cut.size2 is the second size information of the target detection region, img.size1 is the first size information of the original image, n is the number of target detection regions in the original image, and v is the preset distance between the target detection regions.

[0012] In some alternative embodiments, determining the fourth size information of the original image in the image to be detected according to the third size information of the original image and the second size information of the target detection region includes: determining the fourth size information of the original image in the image to be detected according to the following formula:

[0013] orisin.size4 = img.size3 - cut.size2

[0014] where orisin.size4 is the fourth size information of the original image in the image to be detected, and img.size3 is the third size information of the original image.

[0015] In some alternative embodiments, determining the first coordinate information of the target detection region in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second size information of the target detection region and the fourth size information of the original image in the image to be detected includes: obtaining the first coordinate information of the target detection region in the image to be detected according to the following formula:

[0016]

[0017] The second coordinate information of the original image in the image to be detected is obtained according to the following formula:

[0018]

[0019] where i is the serial number of the target detection area, is the first coordinate information of the target detection area in the image to be detected; is the second coordinate information of the original image in the image to be detected.

[0020] In some alternative embodiments, the radar-vision detection network includes a millimeter-wave radar feature extraction unit, a first vision extraction unit, a second vision extraction unit, and an adaptive feature fusion unit;

[0021] The image to be detected is input into the radar-vision detection network, and the radar-vision detection network is used to detect the target to be detected in the target detection area, and a feature image output by the radar-vision detection network is obtained, including: the millimeter-wave radar feature extraction unit is used to perform feature extraction on the image to be detected to obtain a millimeter-wave radar feature map and a first weight coefficient corresponding to the millimeter-wave radar feature map; the first vision extraction unit is used to extract low-level vision features in the image to be detected to obtain a low-level vision feature map and a second weight coefficient corresponding to the low-level vision feature map; the second vision extraction unit is used to extract high-level vision features in the image to be detected to obtain a high-level vision feature map and a third weight coefficient corresponding to the high-level vision feature map; the adaptive feature fusion unit is used to perform adaptive feature fusion processing on the millimeter-wave radar feature map, the first weight coefficient, the low-level vision feature map, the second weight coefficient, the high-level vision feature map, and the third weight coefficient to obtain a feature image.

[0022] In some alternative embodiments, the detection result includes the width and height of the target detection area in the original image, and the coordinate information of the target detection area in the original image; the moving target recognition method further includes: determining the width and height of the target detection area in the original image according to the following formula:

[0023]

[0024] Determining the coordinate information of the target detection area in the original image according to the following formula:

[0025]

[0026] Where, w is the width of the target detection area in the image to be detected, h is the height of the target detection area in the image to be detected, λ is the actual width of the target to be detected, θ is the actual height of the target to be detected, f is the focal length of the camera used to collect the original image, and d is the distance of the target to be detected detected by the millimeter-wave radar. are the coordinate information of the target detection area in the original image.

[0027] In some optional embodiments, according to the target detection area corresponding to the feature image and the original image, the feature image is restored to obtain a target image, including: obtaining the target image according to the following formula:

[0028]

[0029] Where, are the coordinates of the detection box restored from the image to be detected to the original image, are the coordinates of the detection box in the target detection area, and the detection box includes the target detection result.

[0030] In the second aspect of the embodiments of the present application, a target detection device based on radar-vision fusion is provided, including: a reconstruction unit for reconstructing the original image according to the detection result of the millimeter-wave radar on the target to be detected to obtain an image to be detected, the image to be detected includes the original image and at least one target detection area, and the target detection area includes at least one target to be detected; a detection unit for inputting the image to be detected into a radar-vision detection network, and detecting the target to be detected in the target detection area through the radar-vision detection network to obtain a feature image output by the radar-vision detection network, and the feature image includes a target detection result for identifying the target to be detected; a restoration unit for restoring the feature image according to the target detection area corresponding to the feature image and the original image to obtain a target image, and the area where the target to be detected is located in the target image is the same as the area where the target to be detected is located in the original image.

[0031] In the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, and computer-executable instructions are stored in the computer-readable storage medium. When the processor executes the computer-executable instructions, the radar-vision fusion-based moving target recognition method in the first aspect is implemented.

[0032] In the fourth aspect of the embodiments of the present application, a computing device is provided, including: at least one processor and a memory; the memory stores computer-executable instructions; at least one processor executes the computer-executable instructions stored in the memory, so that at least one processor executes the radar-vision fusion-based moving target recognition method in the first aspect.

[0033] In the fifth aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes a computer program; when the computer program is executed, it implements the method for identifying moving targets based on radar-vision fusion as described in the first aspect.

[0034] In the method, device, and storage medium for identifying moving targets based on radar-vision fusion provided by the embodiments of the present application, first, starting from the data level, the spatial position of potential targets is characterized according to the millimeter-wave radar, and the characterization result is used for dividing the long-distance target area in the original image collected by the camera. Further, the divided area image is reconstructed and detected, thereby improving the visual detection accuracy of long-distance targets. Then, modeling is carried out by fusing the radar and visual detection feature layers. Considering the characteristics that the millimeter-wave radar detection is less affected by light and the image detection has more texture information, the detection accuracy of the system in low-light environments is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] By referring to the drawings and reading the detailed description below, the above and other objects, features, and advantages of the exemplary embodiments of the present application will become readily understandable. In the drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, where:

[0036] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0037] Figure 2 It is a schematic flowchart of the method for identifying moving targets provided by an embodiment of the present application;

[0038] Figure 3 It is a schematic diagram of the principle of the method for identifying moving targets provided by an embodiment of the present application Figure 1 ;

[0039] Figure 4 It is a schematic diagram of the principle of the method for identifying moving targets provided by an embodiment of the present application Figure 2 ;

[0040] Figure 5 It is a schematic diagram of the structure and principle of the radar-vision detection network provided by an embodiment of the present application Figure 1 ;

[0041] Figure 6 It is a schematic diagram of the structure and principle of the radar-vision detection network provided by an embodiment of the present application Figure 2 ;

[0042] Figure 7 It is a schematic diagram of the structure of the storage medium provided by an embodiment of the present application;

[0043] Figure 8 It is a schematic diagram of the structure of the target detection device provided by an embodiment of the present application;

[0044] Figure 9 Schematic structural diagram of the computing device provided by the embodiment of the present application;

[0045] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0046] The principles and spirit of the present application will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are provided only to enable those skilled in the art to better understand and then implement the present application, rather than limiting the scope of the present application in any way. On the contrary, these implementation manners are provided to make the present application more thorough and complete, and to be able to fully convey the scope of the present application to those skilled in the art.

[0047] Those skilled in the art know that the implementation manners of the present application can be realized as a system, a device, an equipment, a method, or a computer program product. Therefore, the present application can be specifically realized in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0048] In addition, the number of any element in the accompanying drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning. The principles and spirit of the present application will be elaborated below with reference to several representative implementation manners of the present application.

[0049] Driven by the demand for intelligent transportation management, the target detection technology for vehicles, ships and other transportation tools has developed rapidly. However, in harsh environments such as low light, the line of sight will be poor, resulting in many limitations on the target detection performance based on pure visual images. At the same time, the actual engineering deployment requires that the detection distance of the equipment cover a long distance to reduce the number of deployed detection devices and achieve the effect of cost reduction and efficiency improvement. However, when the target distance is too far, since the number of pixels it occupies in the image is too small, the detection accuracy is greatly reduced. In summary, target detection at long distances and under low light is a huge challenge for target detection in intelligent transportation systems.

[0050] In existing intelligent transportation systems, cameras, millimeter-wave radars, etc. are commonly used in the target detection process. Among them, the images collected by cameras can provide detailed texture information of objects and the environment, but the detection effect is not good in low-light environments. Although millimeter-wave radars are not affected by light, have strong anti-interference ability and long detection range, they cannot obtain the texture information of objects. Therefore, the complementary camera-radar fusion method is beneficial to improve the accuracy and robustness of the detection system. In recent years, the research on camera-radar fusion methods has gradually increased. From the perspective of the fusion method, it is mainly divided into decision-level fusion, data-level fusion and feature-level fusion. Among them, the idea of camera-radar decision-level fusion is that the sensor first detects and obtains preliminary data, and then sends the data of each sensor to the decision terminal. By analyzing multiple groups of data, the information of the target is finally output.

[0051] For the camera-radar decision-level fusion method at the decision level, it associates the target distance measured by the radar with the same target in the image, providing additional distance information for image detection, but does not consider improving the problem of low accuracy of the camera detection under low-light and long-distance conditions according to the radar information. The idea of camera-radar data-level fusion is to use the millimeter-wave radar to generate a list of target boxes. The target boxes in the list are hypotheses about the existence of targets, and then use the visual detection system to verify the hypotheses.

[0052] The idea of camera-radar feature-level fusion is to send the data of both the millimeter-wave radar and the visual sensor into the network, and combine the features of both to output the detection result. Decision-level and data-level fusion can often only detect targets that can be recognized by both the millimeter-wave radar and the camera, while the feature-level fusion method can extract features from the data of the millimeter-wave radar and the camera, enabling the detection model to learn the association between the millimeter-wave radar and the visual data. Even when the detection effect of a certain sensor is not good, the data of another sensor can be used for supplementation, which is the most effective method for simultaneously using the information of the millimeter-wave radar and the camera.

[0053] In view of this, the embodiments of the present application provide a mobile target recognition method, device and storage medium based on camera-radar fusion. First, starting from the data level, the potential target spatial position is characterized according to the millimeter-wave radar, and the characterization result is used for the division of the long-distance target area in the image collected by the camera; further, the image of the divided area is reconstructed and detected, so as to improve the visual detection accuracy of the long-distance target; then, modeling is carried out from the camera-radar detection feature layer fusion. In view of the characteristics that the millimeter-wave radar detection is less affected by light and the image detection has more texture information, the detection accuracy of the system under low-light environments is further improved.

[0054] First refer to Figure 1 , Figure 1 which is a schematic diagram of the application scenario provided by the embodiment of the present application. The devices involved in this application scenario include the target to be detected, the millimeter-wave radar, the camera unit and the calculation module.

[0055] In practical applications, the camera module can be connected to the computing module through a network interface, so as to send the images it captures to the computing module through the network interface. The millimeter-wave radar can be connected to the GPIO of the input / output interface of the computing module through a CAN transceiver, so as to send the millimeter-wave radar signals it captures to the computing module.

[0056] Correspondingly, the computing module is used to implement target detection according to the moving target recognition method provided in the embodiments of the present application, and output an image containing the detection result.

[0057] In some specific embodiments, an Nvidia AGX Xavier embedded device can be used as the computing module, which is connected to the CAN transceiver through the GPIO expansion port of the computing module, and the CAN transceiver is then connected to the millimeter-wave radar for communication. The RJ45 type gigabit Ethernet ports built in the camera unit and the computing module are connected through a network cable to complete the communication between the two.

[0058] It should be understood that the type of the computing module is not particularly limited in the embodiments of the present application. The computing module can be, for example: a personal digital assistant (PDA) device, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a vehicle-mounted device, a wearable device (such as a smart watch, a smart bracelet), a smart vehicle-mounted device (such as a smart display device), etc., which are not particularly limited in the embodiments of the present application.

[0059] The embodiments of the present application can be applied to various target detection scenarios. For example, in the intelligent transportation scenario, vehicles, ships and other transportation tools are detected. The "targets" mentioned in the embodiments of the present application are vehicles, ships and other transportation tools.

[0060] In some optional embodiments, a power supply module may also be included in this application scenario, which is used to supply power to one or more of the computing module, the camera module and the millimeter-wave radar.

[0061] The following Figure 1 application scenario is referred to Figures 2 - 6 to describe the moving target recognition method according to the exemplary embodiments of the present application. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard. On the contrary, the embodiments of the present application can be applied to any applicable scenario.

[0062] Refer to Figure 2 , Figure 2Flow diagram of the moving target recognition method provided by the embodiments of the present application Figure 1 As shown Figure 2 in the figure, the moving target recognition method includes the following steps:

[0063] S201. Reconstruct the original image according to the detection result of the millimeter-wave radar for the target to be detected, and obtain the image to be detected.

[0064] The image to be detected includes the original image and at least one target detection area, and the target detection area includes at least one target to be detected.

[0065] Exemplarily, Figure 3 Principle diagram of the moving target recognition method provided by the embodiments of the present application Figure 1 As shown Figure 3 in the figure, P1 is the original image collected by the camera unit, and C1, C2,..., Cn are the targets to be detected collected by the millimeter-wave radar, which are mapped to the original image and divided into target candidate areas.

[0066] In this step, after splicing the enlarged target detection area and the reduced original image P1, the image P2 to be detected can be formed.

[0067] In some embodiments, the detection result of the millimeter-wave radar for the target to be detected includes the width and height of the target detection area in the original image, and the coordinate information of the target detection area in the original image. In the embodiments of the present application, the width and height of the target detection area in the original image can be determined according to the following formula:

[0068]

[0069] where w is the width of the target detection area in the image to be detected, h is the height of the target detection area in the image to be detected, λ is the actual width of the target to be detected, is the actual height of the target to be detected, f is the focal length of the camera for collecting the original image, and d is the distance of the target to be detected detected by the millimeter-wave radar. Alternatively, the coordinate information of the target detection area in the original image can be determined according to the following formula:

[0070]

[0071] where is the coordinate information of the target detection area in the original image, x r , y r are the position coordinates of the millimeter-wave radar for collecting the target to be detected mapped to the original image.

[0072] Specifically, are the vertex coordinates of the upper left corner of the target area, are the vertex coordinates of the lower left corner of the target area, are the vertex coordinates of the upper right corner of the target area, are the vertex coordinates of the lower right corner of the target area.

[0073] Furthermore, considering the situation where the traffic flow of vehicles (i.e., the targets to be detected) is large in the real road conditions, there will be an overlapping phenomenon between the obtained target detection areas. In view of this, in the embodiments of the present application, the method of non-maximum suppression (NMS) can be used to remove redundant target detection areas.

[0074] In some embodiments of the present application, the image to be detected can be obtained through the following steps:

[0075] (1) Determine the second size information of the target detection area in the image to be detected according to the first size information of the original image, the number of target detection areas in the original image, and the preset distance between the target detection areas.

[0076] Specifically, the second size information of the target detection area in the image to be detected can be determined according to the following formula:

[0077]

[0078] Among them, cut.size2 is the second size information of the target detection area, img.size1 is the first size information of the original image, n is the number of target detection areas in the original image, and v is the preset distance between the target detection areas.

[0079] In some optional implementation manners, the size of the original image P1 is the same as the size of the image to be detected P2, or the size of the original image P1 and the size of the image to be detected P2 can also be different. When the size of the original image P1 is the same as the size of the image to be detected P2, the detection efficiency of the subsequent radar-vision detection network can be improved.

[0080] It should be noted that there are various splicing methods between the target detection area and the original image. For example, Figure 3 In the splicing method shown in, the target detection areas are arranged horizontally on the upper side of the target detection image, and the original image is located on the lower side of the target detection image; or, in some optional implementation manners, the target detection areas can also be arranged horizontally on the lower side of the target detection image, and the original image is located on the upper side of the target detection image; or, the target detection areas can also be arranged vertically on the right side (or left side) of the target detection image, and the original image is located on the left side (or right side) of the target detection image. Take Figure 3Taking the splicing method shown as an example, the first dimension information img.size1 is the width of the original image P1 (or the image P2 to be detected), and the second dimension information cut.size2 is the width of the target detection area.

[0081] It should be noted that since the radar-vision detection network is likely to recognize different objects in two adjacent candidate areas as the same one and separate the adjacent candidate areas by a certain distance. In the embodiments of the present application, the distance between each target detection area is set to v, which can prevent the phenomenon that the target detection box in the image to be detected finally output spans two candidate areas, thereby improving the accuracy of the detection result. It should be understood that the specific value of v is not limited in the embodiments of the present application.

[0082] As an alternative solution, the distances between each target detection area can also be different. For example, the distance between target detection areas C1 and C2 is V1, and the distance between target detection area C2 and target detection area C3 is V2...

[0083] (2) Determine the fourth dimension information of the original image in the image to be detected according to the third dimension information of the original image and the second dimension information of the target detection area.

[0084] Specifically, the fourth dimension information of the original image in the image to be detected can be determined according to the following formula:

[0085] orisin.size4 = img.size3 - cut.size2

[0086] Among them, orisin.size4 is the fourth dimension information of the original image in the image to be detected, and img.size3 is the third dimension information of the original image.

[0087] It should be noted that in the embodiments of the present application, taking the width and height of the target detection area as the same as an example, the second dimension information of the target detection area here is the height of the target detection area.

[0088] Still taking the above as an example, the fourth dimension information orisin.size4 is the height of the original image P1 in the image to be detected P2, and the third dimension information img.size3 is the original height of the original image P1.

[0089] (3) Determine the first coordinate information of the target detection area in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second dimension information of the target detection area and the fourth dimension information of the original image in the image to be detected;

[0090] Specifically, on the one hand, the first coordinate information of the target detection area in the image to be detected can be obtained according to the following formula:

[0091]

[0092] Among them, i is the serial number of the target detection area, which is the first coordinate information of the target detection area in the image to be detected.

[0093] Specifically, is the coordinate of the upper left corner of the target detection area in the image to be detected, is the coordinate of the lower left corner of the target detection area in the image to be detected, is the coordinate of the upper right corner of the target detection area in the image to be detected, is the coordinate of the lower right corner of the target detection area in the image to be detected.

[0094] On the other hand, the second coordinate information of the original image in the image to be detected can be obtained according to the following formula:

[0095]

[0096] Among them, is the second coordinate information of the original image in the image to be detected.

[0097] Specifically, is the coordinate of the upper left corner of the original image in the image to be detected, is the coordinate of the lower left corner of the original image in the image to be detected, is the coordinate of the upper right corner of the original image in the image to be detected, is the coordinate of the lower right corner of the original image in the image to be detected.

[0098] (4) Obtain the image to be detected according to the first coordinate information and the second coordinate information.

[0099] Specifically, after obtaining the above coordinate information, splice them in sequence from the upper left corner of the picture according to the serial number of each target detection area, the first coordinate information of each target detection area, and the second coordinate information of the original image, so as to obtain the image to be detected.

[0100] S202. Input the image to be detected into the radar-vision detection network, and detect the target to be detected in the target detection area through the radar-vision detection network to obtain the feature image output by the radar-vision detection network.

[0101] Among them, the feature image includes the target detection result for identifying the target to be detected.

[0102] As for the structure and implementation principle of the radar-vision detection network, they are shown in the subsequent embodiments.

[0103] S203. According to the target detection region corresponding to the feature image and the original image, perform a reduction process on the feature image to obtain a target image.

[0104] Among them, the region where the target to be detected is located in the target image is the same as the region where the target to be detected is located in the original image.

[0105] It should be noted that since the above target detection region is generated based on the image to be detected after coordinate transformation and mapping of the original image, it is necessary to restore the feature image to obtain the target image.

[0106] Figure 4 Schematic diagram of the principle of the moving target recognition method provided by the embodiment of the present application Figure 2 As Figure 4 shown, P3 is the feature image, and P4 is the target image restored based on the feature image.

[0107] In the embodiment of the present application, during the image restoration process, first determine the target detection region where the detection frame is located, and then restore the detection frame according to the coordinate information of the target candidate region in the original image. Among them, the detection frame includes the target detection result, and the target detection result is also a means of transportation such as a vehicle or a ship.

[0108] Specifically, the target image can be obtained through the following formula:

[0109]

[0110] Among them, is the coordinate of the detection frame restored from the image to be detected to the original image, is the coordinate of the detection frame in the target detection region, and the detection frame includes the target detection result.

[0111] In this formula, first determine the ratio of the length and width of the detection frame to the target detection region, multiply it by the length and width of the target detection region in the original image, and then add to obtain the coordinate of the detection frame on the original image.

[0112]

[0113] In this formula, is the coordinate of the detection frame in the restored target image.

[0114] Next, the structure and principle of the radar-vision detection network will be described in detail in combination with specific embodiments:

[0115] Figure 5 Schematic diagram of the structure of the radar-vision detection network provided by the embodiment of the present application Figure 1 As Figure 5As shown in the figure, the radar-vision detection network provided by the embodiment of the present application includes: a millimeter-wave radar feature extraction unit, a first vision extraction unit, a second vision extraction unit, and an adaptive feature fusion unit.

[0116] Among them, the first vision extraction unit is used to extract low-level vision feature maps, and the second vision extraction unit is used to extract high-level vision feature maps. By fusing the two, the radar-vision detection network can obtain more semantic information.

[0117] When implementing the above step S202, it specifically includes the following steps:

[0118] (1) Through the millimeter-wave radar feature extraction unit, perform feature extraction on the image to be detected, and obtain a millimeter-wave radar feature map and a first weight coefficient corresponding to the millimeter-wave radar feature map;

[0119] (2) Through the first vision extraction unit, extract low-level vision features in the image to be detected, and obtain a low-level vision feature map and a second weight coefficient corresponding to the low-level vision feature map;

[0120] (3) Through the second vision extraction unit, extract high-level vision features in the image to be detected, and obtain a high-level vision feature map and a third weight coefficient corresponding to the high-level vision feature map;

[0121] In the embodiment of the present application, the millimeter-wave radar feature extraction unit, the first vision extraction unit, and the second vision extraction unit are adjusted to have the same number of channels by using a convolutional layer with a size of 1×1. Next, the sizes are adjusted through upsampling or downsampling operations, and the obtained feature maps are the millimeter-wave radar feature map, the low-level vision feature map, and the high-level vision feature map respectively.

[0122] Furthermore, after passing the millimeter-wave radar feature map, the low-level vision feature map, and the high-level vision feature map through a 1×1 convolution, the corresponding first weight parameter α, second weight coefficient β, and third weight coefficient γ can be obtained respectively.

[0123] (4) Through the adaptive feature fusion unit, perform adaptive feature fusion processing on the millimeter-wave radar feature map, the first weight coefficient, the low-level vision feature map, the second weight coefficient, the high-level vision feature map, and the third weight coefficient, and obtain a feature image.

[0124] Furthermore, by multiplying the millimeter-wave radar feature map, the low-level vision feature map, and the high-level vision feature map by the first weight parameter α, the second weight coefficient β, and the third weight coefficient γ respectively through the adaptive feature fusion unit, a weighted fusion feature image can be obtained. Specifically, the feature image can be obtained through the following formula:

[0125] y = α·x 1 + β·x 2 + γ·x3

[0126] Among them, y is the feature image, and x 1 is the millimeter-wave radar feature map, and x 2 is the low-level visual feature map, and x 3 is the high-level visual feature map, and α + β + γ = 1 (α, β, γ ∈ [0, 1]).

[0127] In the embodiments of the present application, since the detection of millimeter-wave radar is less affected by light, and the image detection has more texture information, the feature advantages of the two are weighted and fused, which is beneficial to improving the detection accuracy of the system in low-light environments. However, the differences in the structural mechanisms, perspectives, etc. between millimeter-wave radar and cameras will result in different sizes and spatial positions. At the same time, for target detection, there are differences in the feature contribution degrees of different layers. If the millimeter-wave radar feature map and the visual feature map are simply subjected to spatial transformation and adjusted to the same size and then superimposed, the characteristics of different feature maps cannot be fully utilized. In order to make full use of the features of radar and vision.

[0128] In addition, the detection accuracy can be further improved by increasing the width and depth of the network. The output after weighted fusion is passed through different branch networks, and the feature maps are processed through convolutions of different sizes on the branches to extract information of different receptive fields in the map. Finally, the output results of the three branches are combined to obtain a stronger image representation ability and improve the accuracy of the detection results.

[0129] Based on the above ideas of the visual enhancement algorithm and the feature weighted fusion method, the design of the radar-vision detection network is carried out, which includes two parts: the fusion calibration of millimeter-wave radar and camera and the design of the network structure.

[0130] Among them, in the radar-vision detection network, both visual enhancement and feature weighted calculation need to obtain the spatial coordinate relationship of the data collected by the millimeter-wave radar and the camera. Therefore, first, the millimeter-wave radar and the camera need to be fusion-calibrated to realize the mapping of the target points in the millimeter-wave radar coordinate system to the image coordinate system. Next, the millimeter-wave radar data is processed to generate a millimeter-wave image. Among them, the distance D, speed V, and RCS detected by the millimeter-wave radar are respectively converted into pixel values of different channels (R, G, B), and the conversion relationship is as follows:[[]]

[0131]

[0132] Among them, V max represents the maximum speed, that is, the speed limit value of the current road; V min represents the minimum speed; RCS max represents the average value of the maximum RCS output by the millimeter-wave radar in multiple measurements; RCS min represents the minimum RCS; D maxRepresents the maximum distance that the millimeter-wave radar can detect. D min Represents the minimum detected distance.

[0133] Next, this application designs a radar-vision detection network based on the YOLOv4-tiny basic framework. Figure 6 Schematic structure of the radar-vision detection network provided by the embodiments of this application Figure 2 . As Figure 6 shown, the radar-vision detection network mainly includes three parts: a millimeter-wave radar feature extraction branch (Radar_Net), a visual image feature extraction branch (CSPDarknet53-tiny), and a radar-vision weighted fusion framework (A-Net).

[0134] In the embodiments of this application, the millimeter-wave feature extraction branch Radar-Net is composed of basic convolutional units CBLM and CBL, while the visual extraction branch CSPDarknet53-tiny is mainly composed of CBL units and Resblock units. Among them, CBLM consists of a convolutional layer, batch normalization processing, an activation function, and a max pooling layer, CBL is composed of a convolutional layer, batch normalization processing, and an activation function combination, and Resblock nests 4 CBLs in a residual manner and then processes them through a max pooling layer.

[0135] To fuse the millimeter-wave radar and the visual feature map to improve the detection accuracy, first, the feature map Layer1 with a size of 26×26 in Radar_Net is upsampled and then concatenated and fused with the feature map Layer2 with a size of 52×52 in CSPDarknet53-Tiny and input into the subsequent visual extraction network. Then, the feature map Layer3 with a size of 26×26, the feature map Layer4 with a size of 13×13, and Layer1 in CSPDarknet53-Tiny are input into the A-Net framework, and the output result is sent to the YOLO Head to complete the detection.

[0136] Experimental results and analysis: To test the application effect of the algorithm proposed in this paper, the algorithm proposed in this paper is tested and compared with pure vision algorithms YOLOv4, YOLOv4-tiny, and a typical radar-vision fusion algorithm RVNet. The tests are carried out on the Nuscenes public dataset and the dataset of urban roads collected by the project team under low-light conditions at night.

[0137] The detection situations of the test model in two scenarios of low brightness at night and normal brightness during the day on the Nuscenes test set are shown in Tables 1 and 2:

[0138] Table 1 Comparison of the performance of different networks on Nuscenes night data

[0139] Network AP (%) FPS YOLOv4 - tiny 47.01 33 YOLOv4 57.86 10 RVNet 60.70 14 The radar - vision detection network of this application 67.26 12

[0140] Table 2 Comparison of Data Performance of Different Networks in Nuscenes Daylight

[0141]

[0142]

[0143] As can be seen from Table 2, under the test of normal brightness scenarios, the AP value of the radar-vision fusion network in the embodiments of the present application is only 5% higher than that of the pure vision network. However, in low brightness scenarios, Table 1 shows that the AP value of the method in this paper is 10% higher than that of YOLOv4 and 7% higher than that of RVNet. The results show that the radar-vision detection network provided by the embodiments of the present application has improved detection performance for targets both at night with low brightness and during the day.

[0144] To further illustrate the detection effect of the method proposed in the embodiments of the present application in actual scenarios, evaluation and verification were carried out on a self-made dataset. The experimental results are shown in Table 3 below:

[0145] Table 3 Comparison of Performance of Different Networks on the Self-made Dataset

[0146] Network AP (%) FPS YOLOv4 - tiny 41.29 33 YOLOv4 50.52 10 RVNet 65.29 14 The radar - vision detection network of this application 70.60 12

[0147] As can be seen from Table 3, the AP value of the radar-vision fusion network provided by the embodiments of the present application has improved compared with the performance of other networks.

[0148] Table 4 shows the effects of each algorithm when detecting targets at different distances. Through on-site measurement, the positions 40 meters, 80 meters, and 120 meters away from the detection device were marked. The experimental vehicle was driven near the specified positions to collect the vehicle target test sets at different distances. The YOLOv4 algorithm and the RVNet algorithm with high detection accuracy were used to conduct comparative tests with the radar-vision detection network provided by the embodiments of the present application. Since the YOLOv4-tiny algorithm has no advantage in detecting long-distance targets, no comparison was made.

[0149] Table 4 Average Precision of Vehicle Detection at Different Distances

[0150]

[0151] As shown in Table 4, when the distance to the target is 40 meters, each algorithm can have a good detection effect. As the distance increases, the detection effects of the YOLOv4 and RVNet algorithms gradually decline. When the distance increases to more than 120 meters, only the radar-vision detection network provided by the embodiments of the present application can detect the target. It uses the method of image reconstruction using millimeter-wave distance information and has more advantages than the other two algorithms when detecting vehicle targets at long distances.

[0152] Table 5 shows the performance test results of the radar-vision weighted fusion framework using different combinations of convolutional layers:

[0153] Table 5 Detection performance test under different convolutional layer configurations

[0154]

[0155] As shown in Table 5, among the three combinations of convolutional layers, 1x1, 5x5, and 7x7 have better effects, and the detection results of the four-convolutional-layer combination are second. Therefore, in the embodiments of the present application, a radar-vision plus detection network is composed of 1x1, 5x5, and 7x7 convolutional layers.

[0156] Table 6 tests the detection effects of weighted fusion of millimeter-wave radar feature maps and visual feature maps of different sizes. Feature maps with sizes of 104, 52, and 26 are selected from the YOLOv4-Tiny backbone feature extraction network and fused with the millimeter-wave radar feature maps respectively.

[0157] Table 6 Detection effects of weighted fusion of visual feature maps of different sizes

[0158]

[0159] As shown in Table 6, the fusion at the position where the size of the visual image feature map is 52 is better.

[0160] To test the effects of radar-vision backbone network fusion, radar-vision weighted fusion framework, and image reconstruction on the detection effect, ablation experiments were conducted, and the model was trained and tested on the same training set and test set. The test results are shown in Table 7 below:

[0161]

[0162]

[0163] As shown in Table 7, it can be seen that all three improvement methods optimize the detection effect, and the detection effect of using all three methods simultaneously is the best.

[0164] After introducing the methods of the exemplary embodiments of the present application, next, refer to Figure 7 to describe the storage medium of the exemplary embodiments of the present application.

[0165] Refer to Figure 7 As shown, in the storage medium 700, there is stored a program product for implementing the above method according to the embodiments of the present application. It can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present application is not limited to this.

[0166] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0167] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium.

[0168] The program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN).

[0169] Reference Figure 8 , Figure 8 The structural schematic of the target detection device provided for the embodiments of the present application Figure 1 。As Figure 8 shown, the target detection device 800 includes:

[0170] A reconstruction unit 801, configured to reconstruct the original image according to the detection result of the millimeter-wave radar on the target to be detected, and obtain a to-be-detected image, where the to-be-detected image includes the original image and at least one target detection area, and the target detection area includes at least one target to be detected;

[0171] A detection unit 802, configured to input the to-be-detected image into a radar-vision detection network, and detect the target to be detected in the target detection area through the radar-vision detection network, and obtain a feature image output by the radar-vision detection network, where the feature image includes a target detection result for identifying the target to be detected;

[0172] The restoration unit 803 is configured to perform a restoration process on the feature image according to the target detection region corresponding to the feature image and the original image, so as to obtain a target image, where the region where the target to be detected is located in the target image is the same as the region where the target to be detected is located in the original image.

[0173] In a possible implementation manner, the reconstruction unit 801 is specifically configured to: determine the second size information of the target detection region in the image to be detected according to the first size information of the original image, the number of target detection regions in the original image, and the preset distance between the target detection regions; determine the fourth size information of the original image in the image to be detected according to the third size information of the original image and the second size information of the target detection region; determine the first coordinate information of the target detection region in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second size information of the target detection region and the fourth size information of the original image in the image to be detected; and obtain the image to be detected according to the first coordinate information and the second coordinate information.

[0174] In a possible implementation manner, the reconstruction unit 801 is specifically configured to: determine the second size information of the target detection region in the image to be detected according to the following formula:

[0175]

[0176] where cut.size2 is the second size information of the target detection region, img.size1 is the first size information of the original image, n is the number of target detection regions in the original image, and v is the preset distance between the target detection regions.

[0177] In a possible implementation manner, the reconstruction unit 801 is specifically configured to: determine the fourth size information of the original image in the image to be detected according to the following formula:

[0178] orisin.size4 = img.size3 - cut.size2

[0179] where orisin.size4 is the fourth size information of the original image in the image to be detected, and img.size3 is the third size information of the original image.

[0180] In a possible implementation manner, the reconstruction unit 801 is specifically configured to: obtain the first coordinate information of the target detection region in the image to be detected according to the following formula:

[0181]

[0182] Obtain the second coordinate information of the original image in the image to be detected according to the following formula:

[0183]

[0184] wherein, i is the serial number of the target detection area, is the first coordinate information of the target detection area in the image to be detected; is the second coordinate information of the original image in the image to be detected.

[0185] In a possible implementation manner, the radar-vision detection network includes a millimeter-wave radar feature extraction unit, a first vision extraction unit, a second vision extraction unit, and an adaptive feature fusion unit;

[0186] The detection unit 802 is specifically configured to: through the millimeter-wave radar feature extraction unit, perform feature extraction on the image to be detected to obtain a millimeter-wave radar feature map and a first weight coefficient corresponding to the millimeter-wave radar feature map;

[0187] through the first vision extraction unit, extract low-level vision features in the image to be detected to obtain a low-level vision feature map and a second weight coefficient corresponding to the low-level vision feature map; through the second vision extraction unit, extract high-level vision features in the image to be detected to obtain a high-level vision feature map and a third weight coefficient corresponding to the high-level vision feature map; through the adaptive feature fusion unit, perform adaptive feature fusion processing on the millimeter-wave radar feature map, the first weight coefficient, the low-level vision feature map, the second weight coefficient, the high-level vision feature map, and the third weight coefficient to obtain a feature image.

[0188] In a possible implementation manner, the detection result includes the width and height of the target detection area in the original image, and the coordinate information of the target detection area in the original image; the target detection device 800 further includes: a determination unit 804, configured to determine the width and height of the target detection area in the original image according to the following formula:

[0189] Determine the coordinate information of the target detection area in the original image according to the following formula:

[0190]

[0191] wherein, w is the width of the target detection area in the image to be detected, h is the height of the target detection area in the image to be detected, λ is the actual width of the target to be detected, θ is the actual height of the target to be detected, f is the focal length of the camera for collecting the original image, d is the distance of the target to be detected detected by the millimeter-wave radar, is the coordinate information of the target detection area in the original image.

[0192] In a possible implementation manner, the restoration unit 803 is specifically configured to: obtain a target image according to the following formula:

[0193]

[0194] Among them, are the coordinates of the detection box restored from the image to be detected to the original image, are the coordinates of the detection box in the target detection area, and the detection result of the target is included in the detection box.

[0195] It should be understood that the target detection device 1000 provided in the embodiments of the present application is used to implement the moving target recognition method in any of the above method embodiments on the side of the collection device, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0196] After introducing the methods, media and devices of the exemplary embodiments of the present application, next, refer to Figure 9 to describe the computing device of the exemplary embodiments of the present application. It should be understood that Figure 9 the displayed computing device 900 is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0197] Figure 9 is a schematic structural diagram of the computing device provided in the embodiments of the present application. As Figure 9 shown, the computing device 900 is presented in the form of a general-purpose computing device. The components of the computing device 900 may include but are not limited to: the above-mentioned at least one processing unit 901, the above-mentioned at least one storage unit 902, and a bus 903 connecting different system components (including the processing unit 901 and the storage unit 902).

[0198] The bus 903 includes a data bus, a control bus, and an address bus. The storage unit 902 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 913 and / or a cache memory 922, and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 932.

[0199] The storage unit 902 may also include a program / utilities 952 having a set (at least one) of program units 942, and such program units 942 include but are not limited to: an operating system, one or more application programs, other program units, and program data, and the implementation of a network environment may be included in each or some combination of these examples.

[0200] The computing device 900 may also communicate with one or more external devices 904 (such as a keyboard, a pointing device, etc.). This communication may be carried out through the input / output (I / O) interface 905. Moreover, the computing device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 906. As Figure 9 shown, the network adapter 906 communicates with other units of the computing device 900 through the bus 903. It should be understood that, although not shown in the figure, other hardware and / or software units may be used in conjunction with the computing device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0201] It should be noted that although several units / sub-units of the timing update device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units / sub-units may be embodied in one unit / sub-unit. Conversely, the features and functions of one of the above-described units / sub-units may be further divided and embodied by multiple units / sub-units.

[0202] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0203] Although the spirit and principles of the present application have been described with reference to several specific embodiments, it should be understood that the present application is not limited to the disclosed specific embodiments, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of description. The present application aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A mobile target recognition method based on radar-vision fusion, characterized in that Including: Reconstructing the original image according to the detection result of the millimeter-wave radar for the target to be detected to obtain the image to be detected, where the image to be detected includes the original image and at least one target detection area, and the target detection area includes at least one target to be detected; Inputting the image to be detected into the radar-vision detection network, and detecting the target to be detected in the target detection area through the radar-vision detection network to obtain the feature image output by the radar-vision detection network, where the feature image includes the target detection result for identifying the target to be detected; Restoring the feature image according to the target detection area corresponding to the feature image and the original image to obtain the target image, where the area where the target to be detected is located in the target image is the same as the area where the target to be detected is located in the original image; Among them, the reconstructing the original image according to the detection result of the millimeter-wave radar for the target to be detected to obtain the image to be detected includes: Determining the second size information of the target detection area in the image to be detected according to the first size information of the original image, the number of target detection areas in the original image, and the preset distance between the target detection areas; Determining the fourth size information of the original image in the image to be detected according to the third size information of the original image and the second size information of the target detection area; Determining the first coordinate information of the target detection area in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second size information of the target detection area and the fourth size information of the original image in the image to be detected; Obtaining the image to be detected according to the first coordinate information and the second coordinate information; The restoring the feature image according to the target detection area corresponding to the feature image and the original image to obtain the target image includes: obtaining the target image according to the following formula: Among them, are the coordinates of the detection box restored from the image to be detected to the original image, are the coordinates of the detection box in the target detection area, and the detection box includes the target detection result, are the coordinate information of the target detection area in the original image, v is the preset distance between the target detection areas, and cut.size2 is the second size information of the target detection area.

2. The mobile target recognition method according to claim 1, wherein The determining the second size information of the target detection area in the image to be detected according to the first size information of the original image, the number of target detection areas in the original image, and the preset distance between the target detection areas includes: Determining the second size information of the target detection area in the image to be detected according to the following formula: Where cut.size2 is the second size information of the target detection area, img.size1 is the first size information of the original image, n is the number of target detection areas in the original image, and v is the preset distance between the target detection areas.

3. The mobile target recognition method according to claim 2, wherein The determining the fourth size information of the original image in the image to be detected according to the third size information of the original image and the second size information of the target detection area includes: Determining the fourth size information of the original image in the image to be detected according to the following formula: orisin.size4 = img.size3 - cut.size2 Among them, orisin.size4 is the fourth size information of the original image in the image to be detected, and img.size3 is the third size information of the original image.

4. The mobile target recognition method according to claim 3, wherein The determining the first coordinate information of the target detection region in the image to be detected and the second coordinate information of the original image in the image to be detected according to the second size information of the target detection region and the fourth size information of the original image in the image to be detected includes: Obtaining the first coordinate information of the target detection region in the image to be detected according to the following formula: Obtaining the second coordinate information of the original image in the image to be detected according to the following formula: where i is the serial number of the target detection region, is the first coordinate information of the target detection region in the image to be detected; is the second coordinate information of the original image in the image to be detected.

5. The mobile target recognition method according to any one of claims 1 to 4, characterized in that, The radar-vision detection network includes a millimeter-wave radar feature extraction unit, a first vision extraction unit, a second vision extraction unit, and an adaptive feature fusion unit; The inputting the image to be detected into the radar-vision detection network, and detecting the target to be detected in the target detection region through the radar-vision detection network to obtain the feature image output by the radar-vision detection network includes: Through the millimeter-wave radar feature extraction unit, performing feature extraction on the image to be detected to obtain a millimeter-wave radar feature map and a first weight coefficient corresponding to the millimeter-wave radar feature map; through the first vision extraction unit, extracting low-level vision features in the image to be detected to obtain a low-level vision feature map and a second weight coefficient corresponding to the low-level vision feature map; Through the second vision extraction unit, extracting high-level vision features in the image to be detected to obtain a high-level vision feature map and a third weight coefficient corresponding to the high-level vision feature map; Through the adaptive feature fusion unit, performing adaptive feature fusion processing on the millimeter-wave radar feature map, the first weight coefficient, the low-level vision feature map, the second weight coefficient, the high-level vision feature map, and the third weight coefficient to obtain the feature image.

6. A target detection device based on radar-vision fusion, which is used to execute the moving target recognition method as described in any one of claims 1 to 5, characterized in that, Including: A reconstruction unit for reconstructing the original image according to the detection result of the target to be detected by the millimeter-wave radar to obtain an image to be detected, where the image to be detected includes the original image and at least one target detection region, and the target detection region includes at least one target to be detected; A detection unit for inputting the image to be detected into the radar-vision detection network, and detecting the target to be detected in the target detection region through the radar-vision detection network to obtain the feature image output by the radar-vision detection network, where the feature image includes a target detection result for identifying the target to be detected; A reduction unit for reducing the feature image according to the target detection region corresponding to the feature image and the original image to obtain a target image, where the region where the target to be detected is located in the target image is the same as the region where the target to be detected is located in the original image.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the moving target recognition method according to any one of claims 1 to 5 is implemented.