Target detection and collision early warning method, image processing equipment and vehicle
By registering and channel splicing images of different spectra, multi-channel data is formed, which solves the problem that dual-light image fusion cannot effectively utilize information in the prior art, and improves the accuracy and efficiency of object detection.
Patent Information
- Application Number
- CN202510314552.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The existing dual-light image fusion method cannot effectively utilize dual-light image information, resulting in limited improvement in object detection accuracy.
After obtaining images of different spectra for registration, channel splicing is performed to obtain multi-channel data, and object detection is performed based on the multi-channel data.
This method avoids the loss of image information during the fusion process, retains all information of the dual-light image, and improves the accuracy and efficiency of object detection.
Smart Images

Figure CN120164181A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of image processing technology and intelligent driving technology, and particularly to a method for target detection and collision warning, an image processing device, and a vehicle. Background Art
[0002] With the popularization of autonomous driving technology, in order to ensure driving safety during the autonomous driving process, the requirements for detecting targets such as people and vehicles are getting higher and higher.
[0003] To improve the accuracy of target detection, by equipping the vehicle with dual-light cameras (such as a visible light camera and an infrared camera), and then according to a certain method, fusing the dual-light images to simultaneously utilize the respective characteristics of the dual-light images, and then performing target detection and judgment based on the fused image.
[0004] To a certain extent, this method improves the accuracy of target detection, but due to the fact that the existing dual-light image fusion method cannot effectively utilize the dual-light image information, there is still a problem that the improvement of the accuracy of target detection is limited. Summary of the Invention
[0005] To solve the existing technical problems, the present application provides a method for target detection and collision warning, an image processing device, and a vehicle that can improve the accuracy of target detection.
[0006] In a first aspect, a method for target detection is provided, and the method includes:
[0007] Obtain a first image and a second image respectively captured using different spectra for the same scene, and register the first image and the second image;
[0008] Perform channel splicing on the registered first image and second image to obtain multi-channel data;
[0009] Perform target detection based on the multi-channel data to obtain a target detection result.
[0010] In a second aspect, a method for collision warning is provided, and the method includes:
[0011] Adopt the method for target detection in the above embodiment to obtain the category of the target in the current scene and the position of the target;
[0012] Determine a warning target with a collision risk from the targets according to the category of the target, the position of the target, and the vehicle driving parameters;
[0013] Output a warning prompt for the warning target.
[0014] In a third aspect, there is provided an image processing device, including a processor and a memory connected to the processor. A computer program executable by the processor is stored on the memory. When the computer program is executed by the processor, the steps of the above-mentioned target detection method or the steps of the above-mentioned collision warning method are implemented.
[0015] In a fourth aspect, there is provided a vehicle, including a dual-light imaging device, the above-mentioned image processing device, and an in-vehicle central control system. The dual-light imaging device includes a first imaging device for collecting a first image and a second imaging device for collecting a second image; the in-vehicle central control system is configured to output corresponding assisted driving information according to the target detection result and / or warning prompt of the image processing device.
[0016] For the above-mentioned target detection method, a first image and a second image obtained by photographing the same scene using different spectra are acquired, and the first image and the second image are registered. The registered first image and second image are subjected to channel splicing to obtain multi-channel data, and then target detection is performed based on the multi-channel data to obtain a target detection result. Since the multi-channel data is obtained by splicing the registered first image and second image in channels, it avoids the loss of image information during the fusion process and retains all the information of the dual-light image. Therefore, performing target detection based on the multi-channel data can improve the accuracy of target detection. At the same time, this method does not require performing detection or feature extraction on the dual-light images separately before splicing, and the required amount of calculation is small, which can improve the efficiency of target detection.
[0017] The collision warning method, image processing device, and vehicle provided in the above embodiments belong to the same concept as the corresponding target detection method embodiments, and thus have the same technical effects as the corresponding target detection method embodiments, which will not be elaborated here. Description of the Drawings
[0018] Figure 1 It is a schematic structural diagram of an image processing device in an embodiment.
[0019] Figure 2 It is a schematic structural diagram of a vehicle in an embodiment.
[0020] Figure 3 It is a flowchart of a target detection method in an embodiment.
[0021] Figure 4 It is a schematic diagram of channel splicing in an embodiment.
[0022] Figure 5 It is a schematic framework diagram of a target detection model in an embodiment.
[0023] Figure 6 It is a schematic structural diagram of a feature extraction module in a training stage in an embodiment.
[0024] Figure 7 It is a schematic structural diagram of the feature extraction module in the inference stage in an embodiment.
[0025] Figure 8 It is a flowchart of training an object detection model in an embodiment.
[0026] Figure 9 It is a flowchart of a collision warning method in an embodiment.
[0027] Figure 10 It is a flowchart of object detection and collision warning methods in an embodiment.
[0028] Figure 11 It is a flowchart of object detection and collision warning methods in an embodiment. Detailed implementation manners
[0029] The technical solution of the present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0030] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0031] In the following description, the expression "some embodiments" is involved, which describes a subset of all possible embodiments. It should be noted that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0032] In the following description, the terms "first, second, third" involved are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first, second, third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0033] With the popularization of autonomous driving technology, in order to ensure driving safety during the autonomous driving process, the requirements for detecting targets such as people and vehicles are getting higher and higher. Traditional target detection methods rely on equipping vehicles with visible light cameras and performing target detection based on the captured images of the visible light cameras. However, for ordinary visible light cameras, the road images captured by them are easily affected by many external factors, such as changes in lighting conditions (poor lighting at night), extreme weather conditions (such as haze and smoke), and glare when meeting other vehicles. In such environments, the accuracy of target detection is affected, which may lead to potential dangers during the driving process.
[0034] To improve the accuracy of target detection, an infrared camera is equipped on the vehicle. Taking advantage of the fact that the infrared camera can capture clear images at night or in low-light environments, it makes up for the above deficiencies of visible light images. First, according to a certain weighting method, the visible light image and the infrared image are weighted and fused to simultaneously utilize the characteristics of the two types of images, and then target detection and judgment are performed based on the fused image. This method can perform target detection by using the respective features of the two types of images, and to a certain extent, it improves the accuracy of target detection. However, during the fusion process of the two types of images, certain information of each type of image is lost, which in turn affects the accuracy of target detection. In the prior art, there is also a feature fusion method based on deep learning, that is, after the visible light image and the infrared image are respectively passed through a feature extraction network to extract features in their respective modalities, they are then passed through a specially designed feature fusion module, where the fusion of the two types of features is performed. This method can combine the information from two different sensor sources to obtain a more comprehensive and robust feature representation, thereby improving the recognition performance of subsequent detection tasks. However, this method requires manual design of the feature extraction and fusion models, which is difficult, and the feature extraction of the two types of images is performed separately, resulting in a large computational amount and affecting the efficiency of target detection.
[0035] In view of the above problems, the present application provides an image processing device, as Figure 1 shown. The image processing device includes a processor 101 and a memory 102 connected to the processor. A computer program executable by the processor 101 is stored on the memory 102. When the computer program is executed by the processor 101, it implements the steps of a target detection method and / or the steps of a collision warning method.
[0036] This image processing device can be deployed on the vehicle side, as Figure 2 shown. A vehicle includes a dual-light imaging device, the above-mentioned image processing device, and an in-vehicle central control system. The dual-light imaging device includes a first imaging device for collecting a first image and a second imaging device for collecting a second image. The in-vehicle central control system is configured to output corresponding assisted driving information according to the target detection result and / or warning prompt of the image processing device.
[0037] Among them, the dual - light imaging device can be set in different directions around the vehicle body to detect targets in different directions of the vehicle and / or give early warnings to collision - warning targets.
[0038] Among them, the first imaging device and the second imaging device in the dual - light imaging device can select a suitable combination method according to specific application requirements and target scenarios. In one embodiment, the first imaging device can be an infrared camera for collecting infrared images, and the second imaging device can be a visible - light camera for collecting visible - light images. This combination can provide users with rich environmental information.
[0039] In addition, according to different application requirements and target scenarios, the dual - light imaging device can also select a variety of different types of imaging devices for combination. In one embodiment, the dual - light imaging device can be a combination of an infrared camera and a thermal imaging camera. This combination can be used to detect and analyze the heat distribution of objects and can be applied to fields such as security monitoring and fire warning. In one embodiment, the dual - light imaging device can be a combination of a visible - light camera and an ultraviolet camera. The ultraviolet camera can capture images of ultraviolet radiation. This combination can be used in fields such as environmental monitoring and evaluation of ultraviolet protection effects. In one embodiment, the dual - light imaging device can be a combination of an infrared camera and a LiDAR (Light Detection and Ranging) camera. The LiDAR camera generates a three - dimensional image by emitting laser light and receiving the reflected signal. The combination with the infrared camera can provide more detailed environmental information and is suitable for scenarios such as autonomous driving and robot navigation. In one embodiment, the dual - light imaging device can be a combination of a visible - light camera and a multispectral camera. The multispectral camera can capture images in multiple spectral ranges. The combination with the visible - light camera can provide richer color and texture information and is suitable for fields such as agricultural monitoring and geological exploration. In one embodiment, the dual - light imaging device can be a combination of an infrared camera and a fluorescence camera. The fluorescence camera can capture the fluorescence emitted by an object under a specific light source. This combination can be used in fields such as biomedical research and materials science.
[0040] A target detection method, as Figure 3 shown, includes the following steps:
[0041] Step 302, obtain a first image and a second image respectively taken using different spectra for the same scene, and register the first image and the second image.
[0042] Among them, the types of the first image and the second image are related to the types of the imaging devices.
[0043] Taking the first imaging device as an infrared camera and the second imaging device as a visible light camera as an example, the original infrared image and visible light image obtained by the infrared camera and the visible light camera shooting the same scene are acquired.
[0044] Since the installation positions, viewing angles, and imaging parameters of the first imaging device and the second imaging device are different, the first image and the second image are usually inconsistent in space. To improve the quality of image fusion, the first image and the second image can be first registered to obtain the transformation relationship between the first image and the second image, and the first image and the second image are aligned to the same coordinate system according to the transformation relationship to make their spatial positions and directions consistent.
[0045] In one embodiment, the registration matrix is obtained through the internal and external parameter values of the dual-light imaging device collected before equipment installation, the dual-light images are registered to obtain their homography matrix, and the original first image is transformed according to the homography matrix to obtain the registered first image, ensuring that the same target is in the same position in the first image and the second image.
[0046] Step 304: The registered first image and the second image are channel-spliced to obtain multi-channel data.
[0047] Specifically, the registered first image and the second image are spliced on the channels to obtain multi-channel data. In this fusion method, the number of channels of the obtained multi-channel data is related to the number of channels of the first image and the second image. For example, if the first image includes three different channels and the second image is a single-channel image, then after splicing the channels of the first image and the second image, four-channel data is obtained. This fusion method directly splices the channels of the first image and the second image, and can obtain multi-channel data that retains all the information of the dual-light images, avoiding the loss of image information during the fusion process. Moreover, this fusion method does not require detecting or feature extraction on the dual-light images respectively before fusion, and the required amount of calculation is small.
[0048] Step 306: Object detection is performed based on the multi-channel data to obtain an object detection result.
[0049] Using the multi-channel data obtained by splicing on the channels for object detection, since the multi-channel data retains all the information of the dual-light images and no image information is lost during the fusion process. Object detection based on the multi-channel data with no lost image information can improve the accuracy of object detection.
[0050] The above target detection method obtains a first image and a second image that are respectively captured using different spectra for the same scene, registers the first image and the second image, splices the registered first image and second image in channels to obtain multi-channel data, and then performs target detection based on the multi-channel data to obtain a target detection result. Since the multi-channel data is obtained by splicing the registered first image and second image in channels, it avoids the loss of image information during the fusion process and retains all the information of the dual-light images. Therefore, performing target detection based on the multi-channel data can improve the accuracy of target detection. At the same time, this method does not require detecting or feature extraction on the dual-light images separately before splicing, and the required computational amount is small, which can improve the efficiency of target detection.
[0051] In technical fields such as the military, aerospace, industrial inspection, security monitoring, and intelligent driving, by utilizing the color and texture information of visible light images and the target detection ability of infrared images at night or in bad weather, and fusing the two, comprehensive and accurate high-quality images can be provided, enabling targets to be detected more robustly in complex backgrounds.
[0052] In this application scenario, in one embodiment, the first image is a visible light image, the second image is an infrared image, and target detection is performed based on the multi-channel data to obtain a target detection result, including: splicing the three channels of the registered visible light image with the single channel of the infrared image to obtain four-channel data.
[0053] Specifically, a visible light image has three color channels, which respectively represent three colors: red (R), green (G), and blue (B). These channels together constitute the color and brightness information of the image. An infrared image usually has only one channel, representing the thermal radiation information of an object. Infrared images can clearly show the target contour at night or in harsh environments, but lack color and detail information. As Figure 4 shown, by splicing the three channels of the registered visible light image with the single channel of the infrared image, and adding the single channel of the infrared image as a new channel after the three RGB channels of the visible light image, four-channel data can be obtained. For example, a fused image with four channels, where the three RGB channels are from the visible light image and the fourth channel is from the infrared image. This fusion method can retain all the information of the visible light image and the infrared image, avoiding information loss or quality degradation during the fusion process.
[0054] In the prior art, after fusing dual-light images, a detection model is used for target detection. However, the network of the detection model is relatively complex, and multiple models need to be deployed in edge devices. In scenarios with tight computing power, the consumption of computing resources is too large.
[0055] To address this problem, object detection is performed based on multi-channel data to obtain object detection results, including: inputting the multi-channel data into a trained object detection model, which includes a feature extraction module, a feature fusion module, and a detection head connected in sequence. Among them, the feature extraction module is deployed from a multi-branch structure in the training stage to a single-branch structure after merging the operators of multiple branches; the object detection model outputs the category of the object in the multi-channel data and the prediction box indicating the position of the object.
[0056] In this embodiment, although the method of inputting the fused dual-light image into the detection model for detection is still adopted, it is different from the existing object detection models. The object detection model of this embodiment is as Figure 5 shown, including a feature extraction module 502, a feature fusion module 504, and a detection head 506 connected in sequence.
[0057] Among them, the feature extraction module 502 is used to extract useful feature information from the input multi-channel data, which is crucial for subsequent object recognition and positioning. The features can include edges, textures, colors, shapes, etc. The feature fusion module 504 is responsible for integrating the feature information from different layers or different feature maps. The detection head 506 is used to predict the category and position of the object based on the extracted features.
[0058] Among them, as Figure 6 shown, the feature extraction module 502 adopts a multi-branch structure in the training stage, which can enhance the representation ability of the model and extract richer features. As Figure 7 shown, the trained feature extraction module is deployed to merge the operators of multiple branches into a single-branch structure. In this way, it is more beneficial for edge devices with limited performance, which can make full use of their computing resources, accelerate the inference speed when the model is actually used, and ensure real-time performance.
[0059] In one embodiment, as Figure 8 shown, the steps of predicting the object detection model include:
[0060] Step 802, obtain a training data set, which includes multiple multi-channel data, as well as the object category and the ground truth box representing the position of the object for each multi-channel data. The multi-channel data is obtained by splicing the channels of the first image and the second image registered in the same scene.
[0061] In one embodiment, taking the autonomous driving application scenario as an example, the operation steps for obtaining the training data set are as follows:
[0062] (1), Prepare a dual-light imaging device, and pre-perform a registration operation according to the relative position of its camera installation to obtain the homography matrix between the dual-light imaging devices;
[0063] (2) Install a dual - light data acquisition device at a fixed position with good visibility above the vehicle. Drive the vehicle to collect road driving data (including targets such as people and vehicles) in dual - light scenarios at different driving scenarios such as urban areas and suburbs during two time points: day and night.
[0064] (3) According to the pre - obtained homography matrix, convert visible light (or infrared) to the same field of view as another modality data to obtain training data.
[0065] (4) Use annotation software to annotate the people, vehicles, etc. that need to be detected in the above - mentioned training pictures.
[0066] (5) Stitch the corresponding dual - light data into a four - channel fused image.
[0067] (6) Divide the annotated fused images into a training set, a validation set, and a test set. In one example, 100,000 pieces of data are collected for annotation, among which 80,000 images are used as the training set to train samples, and 20,000 images are used as validation and test samples.
[0068] Step 804: Input the multi - channel data of the training data set into the target detection model to be trained, and obtain the predicted category of the target and the prediction box indicating the position of the target.
[0069] In one embodiment, the target detection model adopts a single - stage network, which can accurately and quickly detect the input. The main framework is as Figure 5 shown. It is mainly composed of three parts: a feature extraction module (Backbone) 502, a feature fusion module (neck) 504, and a detection head (head) 506.
[0070] Among them, the feature extraction module 502 adopts a network with a RepVGG structure to extract useful feature information from the input fused image. These information are crucial for subsequent target recognition and positioning. During the training stage, the feature extraction part of the model adopts a multi - branch structure, which can enhance the model's representation ability and extract richer features. In the inference stage of the model, after converting the parameters of the multi - branch structure of the model into a single - branch 3 * 3 convolution, it is more beneficial for edge devices with limited performance, which can make full use of their computing resources and speed up the inference speed during the actual use of the model to ensure real - time performance.
[0071] The feature fusion module 504 is responsible for integrating feature information from different layers or different feature maps. The detection head 506 is used to predict the category and position of the target according to the extracted features.
[0072] Step 806: Calculate the loss value based on the difference between the predicted category of the target and the annotated category of the target, the intersection over union (IoU) between the predicted bounding box and the ground truth box of the target's position, and the bounding box regression loss.
[0073] Specifically, the loss value consists of two parts, namely the classification loss and the regression loss.
[0074] Among them, the classification loss is calculated based on the difference between the predicted category of the target and the annotated category of the target, and is usually the cross-entropy loss between the predicted category and the annotated category.
[0075] The regression loss includes two parts. One part is the IoU loss, which is the intersection over union between the predicted bounding box and the ground truth box of the target's position. The other part is the bounding box regression loss (Distribution Focal Loss, DFL Loss). Its formula is:
[0076] DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 ))
[0077] Among them, S is the probability value output by the model. As can be seen from the above formula, y is the label value, y i and y i+1 are two adjacent points for which the probability values need to be predicted. By adding this loss, the model can quickly converge towards the coordinates of the label position, and more accurate target box position prediction information can be obtained when occlusion commonly occurs in actual application scenarios.
[0078] In one embodiment, the loss value is the weighted sum of the classification loss and the regression loss, where the weighting weights can be adjusted according to the specific task.
[0079] Step 808: Adjust the parameters of the object detection model based on the loss value until the iteration end condition is reached, and then complete the training to obtain the trained object detection model.
[0080] Specifically, after obtaining the loss value, perform backpropagation based on the loss value. In the backpropagation stage, the loss value will be propagated backward along the model hierarchy to calculate the gradients of the parameters of each layer. The gradient reflects the sensitivity of the loss value to the parameters of each layer, that is, the degree of influence of parameter changes on the loss value. Using the calculated gradients and combining with optimization algorithms (such as Stochastic Gradient Descent SGD, Adam, etc.), update the model parameters to minimize the loss value as much as possible, thereby improving the prediction accuracy of the model.
[0081] The steps of model prediction (forward propagation), loss calculation, backpropagation, and parameter update constitute a training iteration. During the training process, these steps are continuously repeated, and each iteration uses new or re-ordered training samples. As the iterations progress, the model parameters are gradually adjusted, the loss value usually gradually decreases, and the prediction performance of the model is gradually improved.
[0082] When the preset number of iterations is reached, or the loss value on the validation set stops decreasing significantly, or metrics such as accuracy reach stability, it can be considered that the training is complete, and the trained object detection model is obtained.
[0083] In the training process of this object detection model, in the training stage, the feature extraction part of the model adopts a multi-branch structure, which can enhance the model's representation ability and extract richer features. When performing target box coordinate regression, in addition to the traditional IOU Loss, DFL Loss (Distribution Focal Loss) is also used, which can obtain more accurate target box position prediction information when occlusion commonly occurs in actual application scenarios.
[0084] As Figure 5 shown, the detection head 506 is decoupled into a classification task branch and a regression task branch. The classification task branch is used to determine the category of the detected target, and the regression task branch is used to predict the prediction box corresponding to the position of the target.
[0085] This object detection model adopts a decoupled head design, which decouples the category branch and the regression task branch into two parts, can avoid conflicts between these two tasks, improve the model accuracy, and improve the model's convergence speed.
[0086] For the above object detection results, they can be further processed according to the requirements in actual applications. In one embodiment, the object detection method further includes: displaying a fused image corresponding to multi-channel data, and marking the object detection results on the fused image.
[0087] In this embodiment, by displaying the fused image and the object detection results, users can understand the target situation of the current scene in real time through the display interface, including the category of the target and the position, distance, etc. of the target. Specifically, if the object detection results include the category of the target and the prediction box indicating the position of the target, then the category of the target and the prediction box of the target can be marked on the fused image.
[0088] For the above object detection results, there can be different application requirements in different application fields. Taking the application in the field of intelligent driving as an example, the object detection results can be used to further detect targets with collision risks, or the results of object detection can be combined with road information, traffic information to plan an autonomous driving path.
[0089] Taking the application of object detection results to collision detection as an example, as Figure 9 shown, a collision warning method includes the following steps:
[0090] Step 902, obtain the category of the object and the position of the object in the current scene. Among them, the category and position of the object are obtained based on the above object detection method.
[0091] Using the above object detection method, obtain the first image and the second image taken with different spectra for the same scene respectively, register the first image and the second image, splice the registered first image and the second image in channels to obtain multi-channel data, and then perform object detection based on the multi-channel data to obtain object detection results. Since the multi-channel data is obtained by splicing the registered first image and the second image in channels, it avoids the loss of image information during the fusion process and retains all the information of the dual-light images. Therefore, performing object detection based on the multi-channel data can improve the accuracy of object detection. At the same time, this method does not need to perform detection or feature extraction on the dual-light images separately before splicing, and the required computational amount is small, which can improve the efficiency of object detection.
[0092] In this way, using this object detection method can provide a reliable data source for subsequent collision warnings.
[0093] Step 904, determine the warning objects with collision risks from the objects according to the category of the object, the position of the object, and the vehicle driving parameters.
[0094] Specifically, in the autonomous driving scenario, the category of the object may include vehicles, pedestrians, bicycles, animals, road obstacles, etc. The moving objects and static objects can be distinguished by the category of the object.
[0095] The vehicle driving parameters may include vehicle speed, acceleration, steering wheel angle, brake state, etc.
[0096] Among them, for static objects, the relative position relationship and distance between the object and the vehicle can be determined according to the position of the object, and then the collision risk can be evaluated in combination with the vehicle driving parameters.
[0097] For dynamic objects, by tracking the position of the object, the predicted motion trajectory of the object can be obtained. Combining the predicted motion trajectory of the object and the vehicle driving speed, the collision risk is evaluated.
[0098] The objects with an evaluation result of having a collision risk with the vehicle are determined as warning objects.
[0099] Step 906, output a warning prompt for the warning objects.
[0100] The way to output a warning prompt for a warning target of collision risk can be to send a reminder message of text and / or picture indicating the current warning target with collision risk through the display interface of the vehicle-mounted central control system by the assisted driving system; it can also refer to sending a reminder message indicating the warning target with collision risk in the front through the indicator light and / or speaker in the assisted driving system. In some embodiments, it can also include various reminder methods that can remind the driver to be aware of the warning target with recognized collision risk in a timely manner, such as sending an alarm prompt message to the associated terminal device, enabling a reminder message for maximum speed limit control for the vehicle itself, etc.
[0101] Based on the above-mentioned collision warning method, on the basis that the target detection method can provide a reliable data source for subsequent collision warnings, combined with the target detection results and the driving parameters of the vehicle, it is determined whether the target has a collision risk, and the warning target with a collision risk is determined. Based on the reliable target detection results, the reliability of detecting the warning target with a collision risk can be improved, and a warning prompt for the warning target is output, so as to actively give a risk warning to the driver.
[0102] In a specific example, the target detection method and collision warning method of the present application are applied to the target detection and collision warning of a vehicle, such as Figure 10 and Figure 11 shown, including the following steps:
[0103] Step 1: Dual-light data acquisition.
[0104] In this step, the original first image and second image obtained by the first imaging device and the second imaging device respectively shooting the same scene are acquired.
[0105] Specifically, an infrared night vision device and a visible light camera installed on the vehicle are used to collect real-time dual-light images for subsequent recognition processing. Compared with the commonly used single visible light camera, this solution has better acquisition capabilities in scenarios such as night and glare, and can greatly expand the application scenarios.
[0106] Step 2: Target intelligent detection.
[0107] In this step, for the dual-light images, according to the preset registration parameters, the dual-light images are registered, and the registered dual-light images are spliced on the channels, and the spliced multi-channel data is used as the input of the improved target detection algorithm module to obtain accurate target position and category information, etc.
[0108] This step includes sub-steps of registration conversion, channel fusion, and target detection.
[0109] Among them, in the registration conversion sub-step, according to the transformation relationship between the original first image and the second image, the original first image is transformed to obtain the registered first image.
[0110] For the sub-step of channel fusion, the registered first image and second image are stitched in channels to obtain multi-channel data.
[0111] For the sub-step of object detection, the multi-channel data is input into a trained object detection model. The object detection model includes a feature extraction module, a feature fusion module, and a detection head connected in sequence. Among them, the feature extraction module is deployed from a multi-branch structure in the training stage to a single-branch structure after merging the operators of multiple branches; the category of the object in the multi-channel data and the prediction box indicating the position of the object are output through the object detection model.
[0112] Step 3: Display.
[0113] In this step, the fused image corresponding to the multi-channel data can be displayed, and the object detection result can be marked on the fused image. For example, the display device of the vehicle-mounted central control system, according to the above object detection results (information such as the position and size of the vehicle and pedestrian objects), combined with the position information of the vehicle itself, displays the relative position information of the current vehicle and surrounding vehicle and pedestrian objects in real time.
[0114] Step 4: Collision warning.
[0115] Specifically, according to the category of the object, the position of the object, and the vehicle driving parameters, the warning object with a collision risk is determined from the objects; a warning prompt for the warning object is output.
[0116] Among them, for the object detection result of step 2, the object detection result can be directly displayed, or a collision warning can be performed. It is also possible to display the object detection result and the collision warning result after the collision warning.
[0117] The object detection method of the present application is more applicable in low light and glare at night compared with the current single visible light device. Moreover, the proposed stitching fusion between the dual-light data channels can avoid the loss of dual-light data information compared with the current more commonly used pre-fusion and feature fusion of dual-light data, providing a reliable data source for subsequent accurate detection algorithms. Using a feature extraction network based on the RepVGG network, a decoupled head, and a deep learning object detection model optimized by adding DFL loss, a multi-branch structure can be adopted during model training to obtain more comprehensive and diverse features, providing richer features for the subsequent detection part. When actually deploying the model at the edge, the operators of multiple branches are merged to accelerate the inference speed and reduce resource occupancy; and the added DFL loss can obtain more accurate object position information in occlusion situations in vehicle-mounted and other scenarios.
[0118] The collision warning method of the present application, based on the fact that the target detection method can provide a reliable data source for subsequent collision warning, combines the target detection result and the driving parameters of the vehicle to determine whether the target has a collision risk and identify the warning target with a collision risk. Based on the reliable target detection result, the reliability of detecting the warning target with a collision risk can be improved, and a warning prompt for the warning target can be output, so as to actively give a risk warning to the driver.
[0119] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned target detection method or each process of the collision warning method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0120] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements each process of the target detection method or the collision warning method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0121] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0122] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0123] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A target detection method, characterized in that: The method comprises: Acquire a first image and a second image respectively captured using different spectra for the same scene, and register the first image with the second image; Perform channel stitching on the registered first image and the second image to obtain multi-channel data; Target detection is performed based on the multi-channel data to obtain a target detection result.
2. The target detection method according to claim 1, characterized in that: The first image is a visible light image, the second image is an infrared image, and the first image and the second image after registration are channel-joined to obtain multi-channel data, including: The three channels of the registered visible light image and the single channel of the infrared image are spliced to obtain four-channel data.
3. The target detection method according to claim 1 or 2, characterized in that: The target detection result includes the category of the target and a prediction box indicating the location of the target; Performing target detection based on the multi-channel data to obtain a target detection result includes: Inputting the multi-channel data into a trained target detection model, wherein the target detection model includes a feature extraction module, a feature fusion module and a detection head connected in sequence, wherein the feature extraction module is deployed from a multi-branch structure in a training phase to a single-branch structure after merging operators of multiple branches; The target detection model outputs the category of the target in the multi-channel data and a prediction box indicating the location of the target.
4. The target detection method according to claim 3, characterized in that: The method further comprises: Acquire a training data set, the training data set comprising a plurality of multi-channel data, and an object category annotated for each multi-channel data and a real box representing a position of the object, the multi-channel data being obtained by splicing channels of a first image and a second image registered in the same scene; Inputting the multi-channel data of the training data set into the target detection model to be trained to obtain the predicted category of the target and the predicted box indicating the position of the target; Calculate the loss value according to the difference between the predicted category of the target and the marked category of the target, the intersection-over-union ratio between the predicted box and the real box at the position of the target, and the box regression loss; The parameters of the target detection model are adjusted based on the loss value until an iteration end condition is reached, and the training is completed to obtain a trained target detection model.
5. The target detection method according to claim 3, characterized in that: The detection head is decoupled into a classification task branch and a regression task branch, the classification task branch is used to determine the category of the detected target, and the regression task branch is used to predict the prediction box corresponding to the position of the target.
6. The target detection method according to claim 4, characterized in that: The bounding box regression loss is: DFL(S i ,S i+1 )=-((y i+1 -y)log(S i )+(y-y i )log(S i+1 )) Among them, S is the probability value output by the model, y is the label value, and y i and i+1 are two adjacent points whose probability values need to be predicted.
7. The target detection method according to claim 1 or 2, characterized in that: The method further comprises: The fused image corresponding to the multi-channel data is displayed, and the target detection result is marked on the fused image.
8. A collision warning method, characterized in that: The method comprises: Using the target detection method described in any one of claims 1 to 8, obtaining the category of the target in the current scene and the position of the target; Determining a warning target with a collision risk from among the targets according to the type of the target, the location of the target and the vehicle driving parameters; Output a warning prompt for the warning target.
9. An image processing device, characterized in that: It comprises a processor and a memory connected to the processor, wherein the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the steps of the target detection method as described in any one of claims 1 to 7, or the steps of the collision warning method as described in claim 8 are implemented.
10. A vehicle, characterized in that: It includes a dual-light imaging device, an image processing device as described in claim 9, and a vehicle-mounted central control system, wherein the dual-light imaging device includes a first imaging device for acquiring a first image, and a second imaging device for acquiring a second image; the vehicle-mounted central control system is used to output corresponding auxiliary driving information according to the target detection results and / or early warning prompts of the image processing device.