Target identification method and device, controller, vehicle and storage medium

By combining the recognition results of infrared and visible cameras, the yolov7 model and feature extraction module are improved, and the problem of low recognition accuracy of infrared thermal imaging in complex driving scenarios is solved, achieving more efficient target recognition.

CN120472406APending Publication Date: 2025-08-12BYD CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510311724.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing target recognition method based on infrared thermal imaging has low recognition accuracy in complex driving scenarios, resulting in poor target recognition efficiency.

Method used

Combining the recognition results of infrared cameras and visible light cameras, the target recognition results are determined through the improved yolov7 model and feature extraction module, integrating infrared and visible light image features, and a deep learning classification model is used to determine the target recognition results.

Benefits of technology

It improves the accuracy and efficiency of target recognition in complex driving scenarios, and improves the overall effect of target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472406A_ABST
    Figure CN120472406A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a target identification method and device, a controller, a vehicle and a storage medium. The method comprises the following steps: performing target identification on images acquired by an infrared camera and a visible light camera corresponding to a target vehicle to obtain a first identification result of the infrared camera and a second identification result of the visible light camera; and determining a target recognition result based on the first recognition result and the second recognition result. Therefore, the target recognition result of the collected image is determined according to the first recognition result of the infrared camera and the second recognition result of the visible light camera, the accuracy of target recognition can be improved by combining the infrared recognition result and the visible light recognition result, accurate target recognition in a complex driving scene is achieved, and the accuracy of target recognition is improved. And the target identification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target recognition technology, and in particular to a target recognition method, device, controller, vehicle and storage medium. Background Art

[0002] The vehicle's infrared thermal imaging function is suitable for low light, glare and other insufficient lighting conditions at night, entering and exiting tunnels, as well as in severe weather such as fog, haze, rain, and snow. It provides lane line markings, road target markings, safety warnings, perception fusion and other functions in special scenarios.

[0003] However, in the process of research and practice of existing technologies, it was found that the existing target recognition method based on infrared thermal imaging has low accuracy in identifying targets (vehicles, pedestrians, etc.), and is difficult to cope with the more complex target recognition needs in real driving scenarios, resulting in poor target recognition efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a target recognition method, device, controller, vehicle and storage medium, which can combine infrared recognition results and visible light recognition results to improve the accuracy of target recognition, realize accurate target recognition in complex driving scenarios, and thus improve target recognition efficiency.

[0005] In order to achieve the above-mentioned object, according to a first aspect of the present application, a target recognition method is provided, the method comprising:

[0006] Performing target recognition on images captured by an infrared camera and a visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera;

[0007] A target recognition result is determined based on the first recognition result and the second recognition result.

[0008] According to a second aspect of the present application, a target recognition device is provided, the device comprising:

[0009] An identification module is used to perform target recognition on images captured by the infrared camera and the visible light camera corresponding to the target vehicle, and obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera;

[0010] A determination module is used to determine a target recognition result based on the first recognition result and the second recognition result.

[0011] According to a third aspect of the present application, a controller is provided, comprising a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to implement the target recognition method provided in an embodiment of the present application.

[0012] According to a fourth aspect of the present application, a vehicle is provided, comprising the controller provided in the third aspect of the present application.

[0013] According to a fifth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for loading by a processor to execute the steps in any target recognition method provided in the embodiments of the present application.

[0014] According to the sixth aspect of the present application, a computer program product is provided, which includes a computer program, and the computer program is stored in a computer-readable storage medium; when the processor of the controller reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the controller executes the steps in the target recognition method provided in the embodiment of the present application.

[0015] In the target recognition method, device, controller, vehicle, and storage medium of the embodiments of the present application, target recognition is performed on images captured by the infrared camera and visible light camera corresponding to the target vehicle to obtain a first recognition result from the infrared camera and a second recognition result from the visible light camera; and a target recognition result is determined based on the first recognition result and the second recognition result. Thus, by determining the target recognition result of the captured image based on the first recognition result from the infrared camera and the second recognition result from the visible light camera, the accuracy of target recognition can be improved by combining the infrared recognition result and the visible light recognition result, thereby achieving accurate target recognition in complex driving scenarios and thereby improving target recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a schematic diagram of an implementation scenario of a target recognition method provided in an embodiment of the present application;

[0018] Figure 2 This is a flow chart of a target recognition method provided in an embodiment of the present application;

[0019] Figure 3a This is a schematic diagram of a feature extraction module of a target recognition method provided by an embodiment of the present application;

[0020] Figure 3bThis is a schematic diagram of a feature aggregation module of a target recognition method provided in an embodiment of the present application;

[0021] Figure 3c This is a schematic diagram of a second recognition model structure of a target recognition method provided by an embodiment of the present application;

[0022] Figure 4a This is a schematic diagram of the overall process of a target recognition method provided by an embodiment of the present application;

[0023] Figure 4b This is a schematic diagram of a specific process of a target recognition method provided in an embodiment of the present application;

[0024] Figure 5 is a schematic structural diagram of a target recognition device provided in an embodiment of the present application;

[0025] Figure 6 It is a schematic diagram of the structure of the controller provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0027] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0028] Embodiments of the present application provide a target recognition method, device, controller, vehicle, and storage medium. The target recognition device can be integrated into a controller, which can be applied to a server or a terminal. Alternatively, the controller can be an electronic device, which can be a server or a terminal.

[0029] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as basic cloud computing services such as big data and artificial intelligence platforms. Terminals may include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Terminals and servers can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.

[0030] The controller may be integrated into a vehicle, which may be a fuel vehicle, a plug-in hybrid vehicle, a new energy vehicle, etc. This application does not impose any specific limitation on this.

[0031] See also Figure 1 , taking the target recognition device integrated into the controller as an example, Figure 1 A schematic diagram of an implementation scenario of the target recognition method provided in an embodiment of the present application, wherein the controller can be integrated into a vehicle or communicatively connected to the vehicle, and can perform target recognition on images captured by an infrared camera and a visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera; based on the first recognition result and the second recognition result, a target recognition result is determined.

[0032] It should be noted that Figure 1 The schematic diagram of the implementation environment scenario of the target recognition method shown is only an example. The implementation environment scenario of the target recognition method described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application and does not constitute a limitation on the technical solution provided by the embodiment of the present application. It is known to those skilled in the art that with the evolution of target recognition and the emergence of new business scenarios, the technical solution provided in this application is also applicable to similar technical problems.

[0033] The solutions provided in the embodiments of the present application are specifically described by the following embodiments. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0034] This embodiment will be described from the perspective of a target recognition device, which may be integrated into a controller.

[0035] See also Figure 2 , Figure 2 : is a flow chart of a target recognition method provided in an embodiment of the present application. The target recognition method includes:

[0036] Step S101 , performing target recognition on images captured by the infrared camera and the visible light camera corresponding to the target vehicle, and obtaining a first recognition result of the infrared camera and a second recognition result of the visible light camera.

[0037] The target vehicle may be the vehicle currently undergoing target identification, and the infrared camera and visible light camera may be installed on the target vehicle or may establish a network link with the target vehicle. The infrared camera may be a camera that utilizes infrared imaging, for example, an infrared thermal imaging camera. The infrared thermal imaging camera may be a camera based on infrared thermal imaging technology, which captures and measures the distribution of heat by sensing infrared radiation energy emitted by a target object, thereby generating a thermal image. The visible light camera may be a camera that utilizes visible light for imaging. The first recognition result may be the result of target identification performed on an image captured by the infrared camera, and the second recognition result may be the result of target identification performed on an image captured by the visible light camera.

[0038] Optionally, there are multiple ways to capture images through the infrared camera and visible light camera corresponding to the target vehicle. For example, when the target vehicle is in driving mode, the on-board far-infrared thermal imaging camera can be used to capture real-time images of the area in front of the vehicle. The captured images can be three-channel YUV (a color encoding method) images, the resolution of the images captured by the infrared camera can be 640×480, and the frequency of obtaining the images captured by the infrared camera can be 30 frames per second, etc.

[0039] Optionally, there are multiple ways to capture images through the visible light camera corresponding to the target vehicle. For example, when the target vehicle is in driving mode, the on-board forward-looking visible light camera can be used to capture real-time images of the area in front of the vehicle. The captured image can be a three-channel YUV image, and the resolution of the image captured by the visible light camera can be 3860×2160, etc.

[0040] In a specific embodiment, the images captured by the infrared camera and the images captured by the visible light camera need to be ensured to be taken at the same timestamp. The infrared camera and the visible light camera can be jointly calibrated to facilitate image capture of the same area and accurate target recognition.

[0041] Among them, there are many ways to perform target recognition on the images captured by the infrared camera and the visible light camera corresponding to the target vehicle, and obtain the first recognition result of the infrared camera and the second recognition result of the visible light camera. For example, a first recognition model can be used to perform target recognition on the images captured by the infrared camera corresponding to the target vehicle, and obtain the first recognition result corresponding to the infrared camera; a second recognition model can be used to perform target recognition on the images captured by the visible light camera corresponding to the target vehicle, and obtain the second recognition result corresponding to the visible light camera.

[0042] Among them, the first recognition model can be a model for performing target recognition on images captured by an infrared camera, used to implement an infrared recognition algorithm provided in an embodiment of the present application, and the second recognition model can be a model for performing target recognition on images captured by a visible light camera, used to implement a visible light recognition algorithm provided in an embodiment of the present application.

[0043] Optionally, the first recognition model and the second recognition model in the embodiment of the present application may be models improved based on the target detection model (yolov7 model).

[0044] Among them, the first recognition model is used to perform target recognition on the image captured by the infrared camera corresponding to the target vehicle. There are many ways to obtain the first recognition result corresponding to the infrared camera. For example, two adjacent frames of images captured by the infrared camera corresponding to the target vehicle can be spliced to obtain a spliced infrared image; the first recognition model is used to perform target recognition on the spliced infrared image to obtain the first recognition result corresponding to the infrared camera.

[0045] The spliced infrared image may be an image obtained by splicing two adjacent frames of images captured by an infrared camera.

[0046] Among them, there can be multiple ways to perform target recognition on the stitched infrared image through the first recognition model to obtain the first recognition result corresponding to the infrared camera. For example, the first recognition model includes a feature extraction module and a feature aggregation module. The feature extraction module extracts the initial infrared image features from the stitched infrared image, and the feature aggregation module performs multi-scale feature fusion on the initial infrared image features to obtain fused infrared image features, so as to recognize the fused infrared image features and obtain the first recognition result corresponding to the infrared camera.

[0047] Optionally, the feature extraction module may include multiple convolution heads, the multiple convolution heads including at least one of a first convolution head based on ordinary convolution, a second convolution head based on dilated convolution, a third convolution head based on grouped convolution, and a fourth convolution head based on depthwise separable convolution.

[0048] Among them, there can be multiple ways to extract the initial infrared image features from the stitched infrared image through the feature extraction module. For example, multiple convolution heads may include a first convolution head, a second convolution head, a third convolution head and a fourth convolution head. The first feature can be extracted from the stitched infrared image through the first convolution head, and the second feature can be extracted from the stitched infrared image through the second convolution head; the first feature and the second feature are extracted through the third convolution head and the fourth convolution head to obtain the initial infrared image features.

[0049] Among them, there are many ways to obtain the initial infrared image features by extracting the first feature and the second feature through the third convolution head and the fourth convolution head. For example, the first feature and the second feature can be fused to obtain the third feature; the fourth feature can be extracted from the third feature through the third convolution head; the fourth feature and the first feature can be fused to obtain the fifth feature; the fifth feature can be extracted through the fourth convolution head to obtain the initial infrared image features.

[0050] Among them, there are many ways to obtain fused infrared image features by performing multi-scale feature fusion on the initial infrared image features through the feature aggregation module. For example, the initial infrared image features can be subjected to multi-scale convolution processing through the feature aggregation module to obtain at least three different multi-scale convolution features; and the multi-scale convolution features can be subjected to feature fusion to obtain fused infrared image features.

[0051] Optionally, the second recognition model may include an extraction module, a global average pooling layer and a fifth convolution head, the extraction module includes convolution of multiple different convolution kernels, and the fifth convolution head includes grouped convolution.

[0052] Among them, the second recognition model is used to perform target recognition on the image captured by the visible light camera corresponding to the target vehicle. There are many ways to obtain the second recognition result corresponding to the visible light camera. For example, the extraction module can be used to perform feature extraction on the image captured by the visible light camera corresponding to the target vehicle to obtain initial visible light image features, and the initial visible light image features can be pooled through the global average pooling layer to obtain pooled image features. The visible light image features are extracted from the pooled image features through the fifth convolution head, so as to obtain the second recognition result corresponding to the visible light camera based on the visible light image features.

[0053] Correspondingly, an embodiment of the present application provides a target recognition method, including: splicing two adjacent frames of images captured by an infrared camera corresponding to a target vehicle to obtain a spliced infrared image; performing target recognition on the spliced infrared image through a first recognition model to obtain a first recognition result corresponding to the infrared camera, and determining the target recognition result of the image based on the first recognition result and the second recognition result corresponding to the visible light camera.

[0054] The target recognition result may be a result of target recognition performed on the captured image based on the first recognition result and the second recognition result.

[0055] Among them, there can be multiple ways to perform target recognition on the stitched infrared image through the first recognition model and obtain the first recognition result corresponding to the infrared camera. For example, the first recognition model can include a feature extraction module and a feature aggregation module. The feature extraction module can be used to extract the initial infrared image features from the stitched infrared image; the feature aggregation module can be used to perform multi-scale feature fusion on the initial infrared image features to obtain fused infrared image features, so as to recognize the fused infrared image features and obtain the first recognition result corresponding to the infrared camera.

[0056] The initial infrared image feature may be an image feature extracted from the spliced infrared image by a feature extraction module, and the fused infrared image feature may be an image feature obtained by fusing the initial infrared image feature.

[0057] Optionally, the feature extraction module may include multiple convolution heads, and the multiple convolution heads may include at least one of a first convolution head based on ordinary convolution, a second convolution head based on void convolution, a third convolution head based on grouped convolution, and a fourth convolution head based on depth-separable convolution.

[0058] The multiple convolution heads may include at least one convolution, a batch normalization layer, and an activation function layer, and may also include a global average pooling layer. The batch normalization layer (BN) can be used before the activation function to make the output distribution of the previous layer have a mean of 0 and a variance of 1, that is, to normalize the input of the next layer.

[0059] In one embodiment, the first convolution head may include a normal convolution with a convolution kernel of 3×3 and a stride of 2, a batch normalization layer, and an activation function layer (e.g., Sigmoid Linear Unit, Silu for short); the second convolution head may include a dialed convolution with a convolution kernel of 5×5 and a dilation rate of 2, a batch normalization layer, a linear rectified activation function layer (Rectified Linear Unit, Relu for short), and a global average pooling layer. The third convolution head may include a group convolution with a convolution kernel of 3×3×3, a batch normalization layer, an activation function layer (e.g., Swish activation function), and a global average pooling layer. The fourth convolution head may include a depthwise separable convolution with a convolution kernel of 3×3 and a stride of 2, a batch normalization layer, a Sigmoid activation function layer, and a global average pooling layer.

[0060] Optionally, there may be multiple ways to extract initial infrared image features from the stitched infrared image through the feature extraction module. For example, multiple convolution heads may include a first convolution head, a second convolution head, a third convolution head, and a fourth convolution head. The first feature may be extracted from the stitched infrared image through the first convolution head, and the second feature may be extracted from the stitched infrared image through the second convolution head; the first feature and the second feature may be extracted through the third convolution head and the fourth convolution head to obtain the initial infrared image feature.

[0061] The first feature may be a feature extracted by the first convolution head in the stitched infrared image, and the second feature may be a feature extracted by the second convolution head in the stitched infrared image. The feature may be information in the form of a feature map.

[0062] Among them, there are many ways to obtain the initial infrared image features by extracting the first feature and the second feature through the third convolution head and the fourth convolution head. For example, the first feature and the second feature can be fused to obtain the third feature; the fourth feature can be extracted from the third feature through the third convolution head; the fourth feature and the first feature can be fused to obtain the fifth feature; the fifth feature can be extracted through the fourth convolution head to obtain the initial infrared image features.

[0063] Among them, the third feature can be a feature obtained by fusing the first feature and the second feature, the fourth feature can be a feature extracted by the third convolution head from the third feature, and the fifth feature can be a feature obtained by fusing the fourth feature and the first feature.

[0064] Among them, there are many ways to obtain fused infrared image features by performing multi-scale feature fusion on the initial infrared image features through the feature aggregation module. For example, the initial infrared image features can be subjected to multi-scale convolution processing through the feature aggregation module to obtain at least three different multi-scale convolution features; and the multi-scale convolution features can be subjected to feature fusion to obtain fused infrared image features.

[0065] The multi-scale convolution feature may be a feature obtained by performing multi-scale convolution processing on the initial infrared image feature.

[0066] In one embodiment, the infrared recognition algorithm used in the first recognition model in the embodiment of the present application can be improved based on the yolov7 model. The input size of the existing yolov7 model is 640×640×3. After passing through the four basic modules (CBS modules) in the backbone feature extraction network (Backbone), the output feature map is 160×160×128 in size. Existing target detection models are often used for target recognition of visible light images. Taking into account the difference between infrared imaging and visible light imaging, the embodiment of the present application improves the feature extraction method of the model, and fuses the infrared images of the front and rear two frames captured by the infrared camera, a total of 6 channels, and inputs them into the first recognition model. For example, please refer to Figure 3a , Figure 3aThis is a schematic diagram of a feature extraction module of a target recognition method provided in an embodiment of the present application. After the two frames of images are spliced and resized, the size of the spliced infrared image becomes 640×640×6, wherein the input of the feature extraction module is a 640×640×6 image. Convolution head A (i.e., the first convolution head) includes a common convolution with a convolution kernel of 3×3 and a stride of 2, a BN layer, and a Silu activation function layer. After convolution head A extracts features, the output feature map size is 320×320×32. Convolution head B (i.e., the second convolution head) includes a dilated convolution with a convolution kernel of 5×5 and an expansion rate of 2, a BN layer, a Relu activation function layer, and a global average pooling layer. After convolution head B extracts features, the output feature map size is 320×320×64. The convolution head C (i.e., the third convolution head) includes a grouped convolution with a kernel size of 3×3×3, a batch normalization layer, a Swish activation function layer, and a global average pooling layer. After the convolution head C extracts features, the output feature map size is 320×320×96. The convolution head D (i.e., the fourth convolution head) can include a depthwise separable convolution with a kernel size of 3×3 and a stride of 2, a batch normalization layer, a sigmoid activation function layer, and a global average pooling layer. After the convolution head D extracts features, the output feature map size can be 160×160×128. In this way, the first feature can be extracted from the input stitched infrared image through the convolution head A, and the second feature can be extracted from the input stitched infrared image through the convolution head B. Then, the first feature and the second feature are spliced to obtain the third feature. The fourth feature is extracted from the third feature through the convolution head C. The fourth feature and the first feature are spliced to obtain the fifth feature. The fifth feature is extracted through the convolution head D to obtain the initial infrared image feature.

[0067] Optionally, the feature aggregation module in the first recognition model is an ELAN module obtained by improving the Efficient Layer Aggregation Network (ELAN) module in the existing Yolov7 model in the embodiment of the present application, so as to enable more features to be integrated into the module. For example, please refer to Figure 3b , Figure 3bThis is a schematic diagram of a feature aggregation module of a target recognition method provided by an embodiment of the present application, wherein the convolution kernel size of the convolution layer included in the yellow CBS module is 1×1 and the step size is 1, and the convolution kernel size of the convolution layer included in the orange CBS module is 3×3 and the step size is 1. These two different CBS modules have different convolution kernel sizes and can be used to collect different features. The convolution with a larger convolution kernel can extract features with a wider local range, and the convolution with a smaller convolution kernel can extract features with a narrower local range, thereby fusion can obtain features with a wider range, that is, fusion infrared image features.

[0068] For details, please continue to refer to Figure 3b , the input features of the input ELAN module can be input into a yellow CBS module and four orange CBS modules in turn for convolution processing to obtain the first convolution feature, and the input features can be input into the yellow CBS module, orange CBS module, yellow CBS module, orange CBS module, and yellow CBS module in turn for convolution processing to obtain the second convolution feature, and the input features can be input into three yellow CBS modules for convolution processing respectively, and the output of the CBS module can be spliced and fused. Then, the fused features can be input into two yellow CBS modules in turn. The module performs convolution processing to obtain the first sub-convolution feature, and the input feature is sequentially input into the yellow CBS module, the orange CBS module, and the orange CBS module for convolution processing to obtain the second sub-convolution feature. The input feature is sequentially input into the yellow CBS module, the orange CBS module, and the yellow CBS module for convolution processing to obtain the third sub-convolution feature. The first sub-convolution feature, the second sub-convolution feature, and the third sub-convolution feature can then be spliced and fused. The fused features are sequentially input into the two yellow CBS modules for convolution processing to obtain the third convolution feature. Then, the first convolution feature, the second convolution feature, and the third convolution feature are spliced (cat), and the spliced result is convolved through the orange CBS module to obtain the fused infrared image feature. Therefore, based on the fused infrared image feature, the first recognition result of the image captured by the infrared camera can be accurately identified.

[0069] Correspondingly, an embodiment of the present application also provides a target recognition method, including: using a second recognition model to perform target recognition on the image captured by the visible light camera corresponding to the target vehicle, obtaining a second recognition result corresponding to the visible light camera, and determining the target recognition result of the image based on the second recognition result and the first recognition result corresponding to the infrared camera.

[0070] Among them, the second recognition model may include an extraction module, a global average pooling layer and a fifth convolution head, the extraction module includes convolution of multiple different convolution kernels, and the fifth convolution head includes grouped convolution.

[0071] Among them, the second recognition model is used to perform target recognition on the image captured by the visible light camera corresponding to the target vehicle. There are many ways to obtain the second recognition result corresponding to the visible light camera. For example, the image captured by the visible light camera corresponding to the target vehicle can be extracted through the extraction module to obtain initial visible light image features; the initial visible light image features can be pooled through the global average pooling layer to obtain pooled image features; the visible light image features can be extracted from the pooled image features through the fifth convolution head to obtain the second recognition result corresponding to the visible light camera based on the visible light image features.

[0072] In one embodiment, the visible light recognition algorithm used in the second recognition model in the embodiment of the present application can be improved based on the yolov7 model. Since the resolution of the image captured by the original visible light camera is 3860×2160 pixels, in order to more effectively use the high-resolution image, more feature extraction can be performed on the front end of the model, and the backbone feature extraction network in the yolov7 model can be improved. For example, please refer to Figure 3c , Figure 3c This is a schematic diagram of the second recognition model structure of a target recognition method provided by an embodiment of the present application. The improved visible light camera-based extraction module of the present embodiment is shown in the figure. First, a region of interest (ROI) can be captured based on the shared field of view of the infrared camera and the visible light camera. Then, the ROI area can be resized to 1280×1280×6 as input.

[0073] Among them, the extraction module includes convolutions of multiple different convolution kernels. For example, the extraction module may include four convolutions, which can be represented as CBS1, CBS2, CBS3 and CBS4 respectively. The main difference between CBS1-CBS4 is the different convolution kernels. "B" and "S" can represent a BN layer and a Silu activation function layer respectively. For example, CBS1 may include a convolution with a convolution kernel of 3×3 and a stride of 2, CBS2 may include a dilated convolution with a convolution kernel of 3×3, a dilation rate of 2, and a stride of 1, CBS3 may include a convolution with a convolution kernel of 5×5 and a stride of 2, and CBS4 may include a dilated convolution with a convolution kernel of 5×5, a dilation rate of 2, and a stride of 1. Optionally, the number of channels output by each CBS module can be 32.

[0074] The global average pooling layer can include two global average pooling layers, which can be expressed as cat1 and cat2. Among them, the cat1 module is used to convert the output of the CBS1-CBS4 module into 640×640 through a global average pooling layer, and then splice in the channel direction. The size of the feature map after splicing is 640×640×128. The cat2 module is used to convert the output of the CBS1-CBS4 module into 320×320 through a global average pooling layer, and then splice in the channel direction. The size of the feature map after splicing is 320×320×128.

[0075] The fifth convolutional head may include a grouped convolution with a 3×3 kernel and a stride of 1, a dilated convolution with a 3×3 kernel, a dilation rate of 2, and a stride of 1, as well as a batch normalization layer, a Reinforced Lu (ReLU) layer, and a global average pooling layer. After the fifth convolutional head, the output feature map has a size of 160×160×128, i.e., the visible light image features are obtained. Based on the visible light image features, the second recognition model is used to obtain the second recognition result corresponding to the visible light camera.

[0076] Step S102: Determine a target recognition result based on the first recognition result and the second recognition result.

[0077] The target recognition result may be a result of target recognition performed on the captured image based on the first recognition result and the second recognition result.

[0078] Among them, based on the first recognition result and the second recognition result, there can be multiple ways to determine the target recognition result. For example, the confidence corresponding to the first recognition result and the second recognition result can be determined; based on the confidence, the first recognition result and the second recognition result are fused to obtain the target recognition result.

[0079] The confidence level may indicate the credibility of the first recognition result and the second recognition result, and may be used to determine which camera's recognition result should be trusted.

[0080] Among them, there are many ways to determine the confidence levels corresponding to the first recognition result and the second recognition result. For example, the confidence levels corresponding to the first recognition result and the second recognition result can be predicted through a classification model based on images captured by an infrared camera and a visible light camera.

[0081] The classification model may be a deep learning model for predicting the confidence levels corresponding to the first recognition result and the second recognition result.

[0082] For example, see Figure 4a , Figure 4aThis is a schematic diagram of the overall process of a target recognition method provided in an embodiment of the present application. The image captured by the infrared camera can be used for target recognition through the infrared recognition algorithm corresponding to the first recognition model to obtain a first recognition result. The image captured by the visible light camera can be used for target recognition through the visible light recognition algorithm corresponding to the second recognition model to obtain a second recognition result. The first recognition result and the second recognition result can be fused through the target fusion algorithm corresponding to the classification model to obtain a target recognition result, thereby realizing accurate target recognition of the captured image.

[0083] Among them, there can be multiple ways to predict the confidence corresponding to the first recognition result and the second recognition result through the classification model. For example, the classification model can include a convolution head module, a feature splicing and extraction module and a backbone network. Based on the images captured by the infrared camera and the visible light camera, the convolution head module can be used to extract the first image features corresponding to the infrared camera and the second image features corresponding to the visible light camera in the image; the feature splicing and extraction module can be used to fuse the first image features and the second image features to obtain the target image features; and the backbone network can be used to predict the confidence corresponding to the first recognition result and the second recognition result based on the target image features.

[0084] The convolution head module may include a convolution head 1 for extracting features from images captured by an infrared camera, and a convolution head 2 for extracting features from images captured by a visible light camera. The feature splicing and extraction module may be a module for fusing first image features with second image features, and the backbone network may be a network for outputting confidence levels corresponding to first and second recognition results.

[0085] Optionally, the convolution head 1 may include a convolution with a kernel size of 3×3 and a stride of 1, a convolution with a kernel size of 5×5 and a stride of 1, a dilated convolution with a kernel size of 3×3, a dilation rate of 2, and a stride of 1, a batch normalization layer, a Relu activation function layer, and a global average pooling layer. The convolution head 2 may include a convolution with a kernel size of 3×3 and a stride of 2, a convolution with a kernel size of 5×5 and a stride of 1, a dilated convolution with a kernel size of 3×3, a dilation rate of 2, and a stride of 1, a batch normalization layer, a Relu activation function layer, and a global average pooling layer.

[0086] The feature splicing and extraction module can include a convolution with a convolution kernel of 3×3 and a stride of 2, a convolution with a convolution kernel of 5×5 and a stride of 1, a BN layer, a Relu activation function layer, and a global average pooling layer.

[0087] Optionally, the backbone network may be a model structure obtained by removing one convolutional layer (conv1) from a deep convolutional neural network (ResNet50) based on a residual network architecture.

[0088] Optionally, in order to ensure the accuracy of the fusion of the first recognition result and the second recognition result, the two cameras need to be calibrated in advance, that is, the position in the image captured by the infrared camera can be found in the corresponding position in the image captured by the visible light camera. The image captured by the infrared camera and the image captured by the visible light camera need to be fused with similar targets first. Specifically, if there are two pedestrians walking side by side in the image, the detection frames of the targets corresponding to the two pedestrians given by the target detection algorithm can be merged into one detection frame when the distance is close and there are many repetitions. Among them, the ranging method can adopt a monocular ranging method. For example, when the distance difference between two similar targets from the vehicle is less than 0.5 meters and the intersection-over-union (IOU) of the detection frames is greater than 0.5, the detection frames of the two targets can be merged into one.

[0089] In one embodiment, please refer to Figure 4b , Figure 4b This is a specific flow chart of a target recognition method provided by an embodiment of the present application. The image captured by the infrared camera (i.e., the infrared camera picture) needs to be resized to a picture of 480×480 pixels, and the image captured by the visible light camera (i.e., the visible light camera picture) needs to be resized to a picture of 1280×1280 pixels first. The convolution head 1 can include a convolution with a convolution kernel of 3×3 and a step size of 1, a convolution with a convolution kernel of 5×5 and a step size of 1, a convolution with a convolution kernel of 3×3, an expansion rate of 2, and a dilated convolution with a step size of 1, a BN layer, a Relu activation function layer, and a global average pooling layer. After the infrared camera picture is subjected to feature extraction by the convolution head 1, the output feature map size can be 224×224×3, i.e., the first image feature. Convolution head 2 includes a convolution with a 3×3 kernel and a stride of 2, a convolution with a 5×5 kernel and a stride of 1, a dilated convolution with a 3×3 kernel, a dilation rate of 2, and a stride of 1, a batch normalization layer, a ReLU activation function layer, and a global average pooling layer. After feature extraction from the visible light camera image through convolution head 1, the output feature map is 224×224×3, representing the second image features. The first and second image features are then concatenated channel-wise using the feature concatenation and extraction module. After feature extraction by the feature concatenation and extraction module, the output feature map is 112×112×64, representing the target image features. This allows the backbone network to predict the confidence levels of the first and second recognition results based on the target image features.

[0090] In one embodiment, since the recognition capabilities of infrared cameras and visible light cameras are different in different scenarios, the target recognition result can be accurately determined by determining the confidence level of each camera in the current scenario. To this end, the embodiment of the present application adopts a classification model based on deep learning. The classification model needs to be trained on existing data before use. The existing data can include images collected using a forward-looking visible light camera and an infrared camera. The corresponding labels can be the confidence levels of each camera in the current scenario. The confidence levels can be manually determined based on which camera should be more trusted in the current scenario. The sum of the confidence levels can be 1, so that a classification model can be trained.

[0091] After determining the confidence levels corresponding to the first and second recognition results, the first and second recognition results can be fused based on the confidence levels to obtain a target recognition result. There are multiple ways to obtain a target recognition result by fusing the first and second recognition results based on the confidence levels. For example, if the confidence level of the first recognition result is greater than the confidence level of the second recognition result, the first recognition result can be determined as the target recognition result; if the confidence level of the first recognition result is not greater than the confidence level of the second recognition result, the second recognition result can be determined as the target recognition result.

[0092] Among them, when the confidence level of the first recognition result is greater than the confidence level of the second recognition result, it can be shown that the credibility of the infrared recognition result is greater than that of the visible light recognition result. The first recognition result can be determined as the target recognition result. The confidence level of the first recognition result is equal to the confidence level of the second recognition result. Since the visible light recognition result is more accurate than the infrared recognition result, the second recognition result can be determined as the target recognition result. When the confidence level of the first recognition result is less than the confidence level of the second recognition result, the second recognition result can be determined as the target recognition result. In this way, based on the confidence level, the target recognition result with higher accuracy in the current scenario can be determined between the first recognition result and the second recognition result, thereby achieving accurate target recognition in complex driving scenarios and improving target recognition efficiency.

[0093] A vehicle's infrared thermal imaging function is suitable for use in low-light conditions such as nighttime low light and glare, in tunnels, and in inclement weather such as fog, haze, rain, and snow. It provides lane marking, road target identification, safety warnings, and sensor fusion for these special scenarios. Existing infrared thermal imaging-based target recognition methods rely primarily on deep learning, with far-infrared cameras primarily used for infrared target detection and near-infrared cameras for infrared target display. Current algorithms for infrared target detection have low accuracy, which limits the practical application of infrared detection results. If deep learning algorithms can be rationally improved to make them more suitable for infrared detection tasks and integrated with other sensors, higher-performance infrared target detection capabilities will be achieved, contributing to the implementation of all-weather intelligent driving.

[0094] To this end, the existing infrared thermal imaging-based recognition method is not accurate enough in identifying infrared targets (vehicles, pedestrians, etc.), and it is difficult to cope with the more complex road recognition problems in real driving scenarios. The embodiment of the present application provides a target recognition method based on a combination of an infrared thermal imaging camera and a visible light camera, and at the same time combines multiple advanced methods such as deep learning detection algorithms and fusion algorithms. A far-infrared front camera is used for target recognition, and a monocular front-view camera is used for target recognition. The recognition results are then fused, and the final target recognition result is displayed on the human-machine interface (Human Machine Interface, referred to as HMI) on the one hand, and can be input into the collision warning and automatic emergency braking (AEB) functions on the other hand, to achieve effective recognition of road targets in front of the vehicle.

[0095] Specifically, the embodiment of the present application first improves the infrared recognition model. The current deep learning target detection network is mainly based on visible light cameras, and the recognition results of infrared camera images are not accurate enough. Based on the characteristics of infrared thermal imaging cameras themselves (small pixels and image imaging effects similar to grayscale images), the embodiment of the present application improves the CBS module of the original Yolov7 model, adds multiple different types of convolution heads, and more effectively extracts features from multiple angles, so that different infrared features can be effectively abstracted. At the same time, the original ELAN module is improved so that it can more effectively use the features transmitted by the convolution head. In this way, based on the improved first recognition model, a target detection solution more suitable for infrared data is proposed, thereby improving the accuracy of infrared recognition. At the same time, based on the characteristic that the original resolution of optical camera images can be larger, the embodiment of the present application improves the feature extraction module based on the original Yolov7 model, abstracts more features and integrates them into the deep learning model, thereby achieving more accurate target recognition based on the improved second recognition model for visible light recognition, thereby improving the accuracy of visible light target recognition. In addition, the embodiment of the present application post-fuses the visible light recognition results with the infrared recognition results. At the same time, considering that the recognition capabilities of infrared cameras and visible light cameras are different under different light and temperature conditions, different confidence levels are designed to fuse the visible light recognition results with the infrared recognition results. The improved fusion algorithm can take into account scene factors such as light, and realize the effective fusion of infrared cameras and visible light cameras, so that the accuracy of the target recognition results determined after fusion can surpass the recognition results of a single camera, thereby realizing accurate target recognition in complex driving scenarios.

[0096] As can be seen from the above, the embodiment of the present application performs target recognition on the images captured by the infrared camera and the visible light camera corresponding to the target vehicle, obtaining a first recognition result from the infrared camera and a second recognition result from the visible light camera; and determining the target recognition result based on the first recognition result and the second recognition result. In this way, by determining the target recognition result of the captured image based on the first recognition result of the infrared camera and the second recognition result of the visible light camera, the accuracy of target recognition can be improved by combining the infrared recognition result and the visible light recognition result, achieving accurate target recognition in complex driving scenarios, and thereby improving target recognition efficiency.

[0097] To facilitate better implementation of the target recognition method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above target recognition method. The meanings of the terms are the same as those in the above target recognition method, and the specific implementation details can be referred to the description in the method embodiment.

[0098] For example, Figure 5FIG. 2 is a schematic diagram of the structure of a target recognition device provided in an embodiment of the present application. The target recognition device may include a recognition module 201 and a determination module 202, as follows:

[0099] Identification module 201, configured to perform target recognition on images captured by the infrared camera and the visible light camera corresponding to the target vehicle, and obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera;

[0100] The determination module 202 is configured to determine a target recognition result based on the first recognition result and the second recognition result.

[0101] In one embodiment, the determination module 202 is configured to:

[0102] Determining confidence levels corresponding to the first recognition result and the second recognition result;

[0103] Based on the confidence level, the first recognition result and the second recognition result are fused to obtain the target recognition result.

[0104] In one embodiment, the above-mentioned determination of the confidence levels corresponding to the first recognition result and the second recognition result is used to:

[0105] Based on the images captured by the infrared camera and the visible light camera, the confidence level corresponding to the first recognition result and the second recognition result is predicted by the classification model.

[0106] In one embodiment, the classification model includes a convolution head module, a feature splicing and extraction module, and a backbone network. The above-mentioned images captured by the infrared camera and the visible light camera are used to predict the confidence level corresponding to the first recognition result and the second recognition result through the classification model, and are used to:

[0107] Extracting a first image feature corresponding to the infrared camera and a second image feature corresponding to the visible light camera from the image through a convolution head module;

[0108] The first image feature and the second image feature are fused through the feature splicing and extraction module to obtain the target image feature;

[0109] Through the backbone network, based on the target image features, the confidence levels corresponding to the first and second recognition results are predicted.

[0110] In one embodiment, the first recognition result and the second recognition result are fused based on the confidence level to obtain the target recognition result, including:

[0111] If the confidence level of the first recognition result is greater than the confidence level of the second recognition result, the first recognition result is determined as the target recognition result;

[0112] If the confidence level of the first recognition result is not greater than the confidence level of the second recognition result, the second recognition result is determined as the target recognition result.

[0113] In one embodiment, the identification module 201 is configured to:

[0114] Using the first recognition model, target recognition is performed on the image captured by the infrared camera corresponding to the target vehicle to obtain a first recognition result corresponding to the infrared camera;

[0115] The second recognition model is used to perform target recognition on the image captured by the visible light camera corresponding to the target vehicle to obtain a second recognition result corresponding to the visible light camera.

[0116] In one embodiment, the first recognition model is used to perform target recognition on an image captured by an infrared camera corresponding to the target vehicle, and a first recognition result corresponding to the infrared camera is obtained, including:

[0117] Splicing two adjacent frames of images captured by the infrared camera corresponding to the target vehicle to obtain a spliced infrared image;

[0118] The target is recognized by the spliced infrared image using the first recognition model to obtain a first recognition result corresponding to the infrared camera.

[0119] As can be seen from the above, in the embodiment of the present application, the recognition module 201 performs target recognition on the images captured by the infrared camera and the visible light camera corresponding to the target vehicle, thereby obtaining a first recognition result from the infrared camera and a second recognition result from the visible light camera; the determination module 202 determines the target recognition result based on the first recognition result and the second recognition result. In this way, by determining the target recognition result of the captured image based on the first recognition result of the infrared camera and the second recognition result of the visible light camera, the accuracy of target recognition can be improved by combining the infrared recognition result and the visible light recognition result, thereby achieving accurate target recognition in complex driving scenarios and thereby improving target recognition efficiency.

[0120] Accordingly, the embodiment of the present application also provides a controller, such as Figure 6 As shown, Figure 6 Schematic diagram of the structure of the controller provided in an embodiment of the present application. The controller 300 includes a processor 301 having one or more processing cores, a memory 302 having one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 is electrically connected to the memory 302. It will be understood by those skilled in the art that the controller structure shown in the figure does not constitute a limitation of the controller, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0121] The processor 301 is the control center of the controller 300. It connects the various parts of the entire controller 300 using various interfaces and lines. It executes various functions of the controller 300 and processes data by running or loading software programs and / or units stored in the memory 302 and calling data stored in the memory 302. The processor 301 can be a processor CPU, a graphics processor GPU, a network processor (NP), etc., and can implement or execute the various methods, steps, and logic blocks disclosed in the embodiments of this application.

[0122] In the embodiment of the present application, the processor 301 in the controller 300 loads instructions corresponding to one or more application processes into the memory 302 according to the following steps, and the processor 301 runs the application stored in the memory 302 to implement various functions, such as:

[0123] Target recognition is performed on images captured by the infrared camera and the visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera; and a target recognition result is determined based on the first recognition result and the second recognition result.

[0124] Furthermore, various functions implemented by running the application stored in the memory 302 can also be described in the aforementioned embodiments and will not be repeated here.

[0125] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0126] Optional, such as Figure 6 As shown, the controller 300 further includes: a touch screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. Among them, the processor 301 is electrically connected to the touch screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307 respectively. Those skilled in the art will understand that Figure 6 The controller structure shown in the figure does not constitute a limitation to the controller, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0127] The touch display screen 303 can be used for displaying a graphical user interface and receiving the operation instructions generated by the user acting on the graphical user interface. The touch display screen 303 may include a display panel and a touch panel. Among them, the display panel can be used for displaying the information input by the user or the information provided to the user and various graphical user interfaces of the controller, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light emitting diode (OLED), or the like. The touch panel can be used for collecting the touch operation of the user thereon or near it (such as the user uses any suitable object or accessory such as a finger, a stylus on the touch panel or near the touch panel), and generates corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 301, and can receive the command sent by the processor 301 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 301 to determine the type of touch event, and then the processor 301 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 303 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 303 can also be used as part of the input unit 306 to realize the input function.

[0128] The RF circuit 304 may be used to transmit and receive RF signals, so as to establish wireless communication with a network device or other controllers through wireless communication, and to transmit and receive signals with the network device or other controllers.

[0129] The audio circuit 305 can be used to provide an audio interface between the user and the controller via a speaker and microphone. The audio circuit 305 can convert received audio data into electrical signals and transmit them to the speaker, which then converts them into sound signals for output. The microphone, on the other hand, converts collected sound signals into electrical signals, which are then received by the audio circuit 305 and converted into audio data. The audio data is then output to the processor 301 for processing, and then sent to another controller via the RF circuit 304. Alternatively, the audio data can be output to the memory 302 for further processing. The audio circuit 305 may also include an earphone jack to allow communication between an external headset and the controller.

[0130] The input unit 306 may be configured to receive input target video and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0131] Power supply 307 is used to supply power to various components of controller 300. Optionally, power supply 307 can be logically connected to processor 301 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 307 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0132] although Figure 6 Not shown in the figure, the controller 300 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be repeated here.

[0133] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant descriptions of other embodiments. It should be noted that the controller provided in the embodiment of the present application and the target recognition method in the above embodiment are based on the same concept. The specific implementation process is detailed in the above method embodiment and will not be repeated here.

[0134] As can be seen from the above, the controller provided in the embodiment of the present application can perform target recognition on the images captured by the infrared camera and the visible light camera corresponding to the target vehicle, thereby obtaining a first recognition result of the infrared camera and a second recognition result of the visible light camera; and determine the target recognition result based on the first recognition result and the second recognition result. In this way, by determining the target recognition result of the captured image based on the first recognition result of the infrared camera and the second recognition result of the visible light camera, the accuracy of target recognition can be improved by combining the infrared recognition result and the visible light recognition result, thereby achieving accurate target recognition in complex driving scenarios and thereby improving target recognition efficiency.

[0135] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0136] To this end, an embodiment of the present application provides a computer-readable storage medium including a computer program. When the computer program is executed on a controller, the computer program is used to cause the controller to perform any of the target recognition methods provided in the embodiments of the present application. For example, the computer program can perform the following steps of the target recognition method:

[0137] Target recognition is performed on images captured by the infrared camera and the visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera; and a target recognition result is determined based on the first recognition result and the second recognition result.

[0138] Furthermore, for the detailed steps of the above method steps, please refer to the description in the above embodiments, which will not be repeated here.

[0139] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0140] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0141] Since the computer program stored in the computer-readable storage medium can execute any target recognition method provided in the embodiments of the present application, the beneficial effects that can be achieved by any target recognition method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0142] According to one aspect of the present application, a computer program product is also provided, including a computer program, which is stored in a computer-readable storage medium; when a processor of a controller reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the controller executes the methods provided in various optional implementations of the above embodiments.

[0143] In the above-mentioned target recognition device, computer-readable storage medium, controller, and computer program product embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a particular embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes and beneficial effects of the target recognition device, computer-readable storage medium, computer program product, controller, and their corresponding units described above can be referred to the description of the target recognition method in the above embodiments, and the specific details will not be repeated here.

[0144] The above is a detailed introduction to a target recognition method, device, controller, vehicle, computer-readable storage medium and computer program product provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A target recognition method, characterized in that: include: Performing target recognition on images captured by an infrared camera and a visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera; A target recognition result is determined based on the first recognition result and the second recognition result.

2. The target recognition method according to claim 1, characterized in that: The determining a target recognition result based on the first recognition result and the second recognition result includes: Determining confidence levels corresponding to the first recognition result and the second recognition result; Based on the confidence level, the first recognition result and the second recognition result are fused to obtain a target recognition result.

3. The target recognition method according to claim 2, characterized in that: The determining of the confidence levels corresponding to the first recognition result and the second recognition result includes: Based on the images captured by the infrared camera and the visible light camera, confidence levels corresponding to the first recognition result and the second recognition result are predicted using a classification model.

4. The target recognition method according to claim 3, characterized in that: The classification model includes a convolution head module, a feature splicing and extraction module, and a backbone network. The confidence corresponding to the first recognition result and the second recognition result is predicted by the classification model based on the image captured by the infrared camera and the visible light camera, including: Extracting, from the image, a first image feature corresponding to the infrared camera and a second image feature corresponding to the visible light camera through the convolution head module; fusing the first image feature and the second image feature through the feature splicing and extraction module to obtain a target image feature; The backbone network is used to predict the confidence levels corresponding to the first recognition result and the second recognition result based on the target image features.

5. The target recognition method according to claim 2, characterized in that: The fusing the first recognition result and the second recognition result based on the confidence level to obtain a target recognition result includes: If the confidence level of the first recognition result is greater than the confidence level of the second recognition result, determining the first recognition result as the target recognition result; If the confidence level of the first recognition result is not greater than the confidence level of the second recognition result, the second recognition result is determined as the target recognition result.

6. The target recognition method according to any one of claims 1 to 5, characterized in that: The performing target recognition on images captured by the infrared camera and the visible light camera corresponding to the target vehicle to obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera includes: Using a first recognition model, performing target recognition on an image captured by an infrared camera corresponding to a target vehicle, and obtaining a first recognition result corresponding to the infrared camera; A second recognition model is used to perform target recognition on an image captured by a visible light camera corresponding to the target vehicle to obtain a second recognition result corresponding to the visible light camera.

7. The target recognition method according to claim 6, characterized in that: The first recognition model is used to perform target recognition on an image captured by an infrared camera corresponding to the target vehicle to obtain a first recognition result corresponding to the infrared camera, including: Splicing two adjacent frames of images captured by the infrared camera corresponding to the target vehicle to obtain a spliced infrared image; Target recognition is performed on the spliced infrared image using a first recognition model to obtain a first recognition result corresponding to the infrared camera.

8. A target recognition method, characterized in that: include: Splicing two adjacent frames of images captured by the infrared camera corresponding to the target vehicle to obtain a spliced infrared image; The target recognition is performed on the spliced infrared image through a first recognition model to obtain a first recognition result corresponding to the infrared camera, so as to determine the target recognition result of the image according to the first recognition result and a second recognition result corresponding to the visible light camera.

9. The target recognition method according to claim 8, characterized in that: The first recognition model includes a feature extraction module and a feature aggregation module. The target recognition is performed on the spliced infrared image by the first recognition model to obtain a first recognition result corresponding to the infrared camera, including: Extracting initial infrared image features from the spliced infrared image by the feature extraction module; The feature aggregation module performs multi-scale feature fusion on the initial infrared image features to obtain fused infrared image features, so as to identify the fused infrared image features and obtain a first recognition result corresponding to the infrared camera.

10. The target recognition method according to claim 9, characterized in that: The feature extraction module includes multiple convolution heads, and the multiple convolution heads include at least one of a first convolution head based on normal convolution, a second convolution head based on void convolution, a third convolution head based on grouped convolution, and a fourth convolution head based on depth-separable convolution.

11. The target recognition method according to claim 10, characterized in that: The multiple convolution heads include the first convolution head, the second convolution head, the third convolution head, and the fourth convolution head, and extracting the initial infrared image features from the spliced infrared image by the feature extraction module includes: extracting a first feature from the stitched infrared image using the first convolution head, and extracting a second feature from the stitched infrared image using the second convolution head; The first feature and the second feature are extracted by the third convolution head and the fourth convolution head to obtain initial infrared image features.

12. The target recognition method according to claim 11, characterized in that: The extracting the first feature and the second feature by the third convolution head and the fourth convolution head to obtain an initial infrared image feature includes: fusing the first feature and the second feature to obtain a third feature; Extracting a fourth feature from the third feature by the third convolution head; Fusing the fourth feature with the first feature to obtain a fifth feature; The fourth convolution head is used to extract the fifth feature to obtain an initial infrared image feature.

13. The target recognition method according to claim 9, characterized in that: The performing multi-scale feature fusion on the initial infrared image features by the feature aggregation module to obtain fused infrared image features includes: Performing multi-scale convolution processing on the initial infrared image features through the feature aggregation module to obtain at least three different multi-scale convolution features; The multi-scale convolution features are subjected to feature fusion to obtain fused infrared image features.

14. A target recognition method, characterized in that: include: A second recognition model is used to perform target recognition on the image captured by the visible light camera corresponding to the target vehicle to obtain a second recognition result corresponding to the visible light camera, so as to determine the target recognition result of the image based on the second recognition result and the first recognition result corresponding to the infrared camera.

15. The target recognition method according to claim 14, characterized in that: The second recognition model includes an extraction module, a global average pooling layer and a fifth convolution head. The extraction module includes convolution of multiple different convolution kernels, and the fifth convolution head includes grouped convolution.

16. The target recognition method according to claim 15, characterized in that: The second recognition model is used to perform target recognition on an image captured by a visible light camera corresponding to the target vehicle to obtain a second recognition result corresponding to the visible light camera, including: The extraction module extracts features from the image captured by the visible light camera corresponding to the target vehicle to obtain initial visible light image features; Performing pooling processing on the initial visible light image features through the global average pooling layer to obtain pooled image features; The fifth convolution head is used to extract visible light image features from the pooled image features, so as to obtain a second recognition result corresponding to the visible light camera based on the visible light image features.

17. A target recognition device, characterized in that: include: An identification module is used to perform target recognition on images captured by the infrared camera and the visible light camera corresponding to the target vehicle, and obtain a first recognition result of the infrared camera and a second recognition result of the visible light camera; A determination module is used to determine a target recognition result based on the first recognition result and the second recognition result.

18. A controller, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is enabled to perform the steps of any one of the methods of claims 1 to 16.

19. A vehicle, characterized in that: The vehicle includes the controller of claim 18.

20. A computer-readable storage medium, characterized in that The method comprises a computer program, and when the computer program is run on a controller, the computer program is used to make the controller execute the steps of the method according to any one of claims 1 to 16.

21. A computer program product, characterized in that The method comprises a computer program or instructions, which implements the steps of any one of claims 1 to 16 when executed by a processor.

Citation Information

Cited By

  • Infrared target identification method and system based on deep learning, medium and product

    CN121053371A