Target detection method, device, equipment, medium and product
By performing object detection on image data and classifying reflection scenes, using pre-training prediction network to correct key points, the problem of large ground point recognition error in reflection scenes is solved, and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510547087.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-12
AI Technical Summary
In the reflection scene, the target detection algorithm based on 2D images cannot accurately calculate the grounding point, resulting in large distance measurement errors, affecting the vehicle's planning and control logic decisions. Especially when the reflection exists, the target's grounding point identification error is significant.
By performing object detection on image data, reflecting scene classification is performed while using pre-training prediction network to regress the initial detection results, correcting key points to eliminate the impact of reflection, and determining the final object detection result.
It realizes accurate identification of docking ground points in reflection scenes, improves the accuracy and efficiency of target detection, and solves the problem of misrecognition of reflections.
Smart Images

Figure CN120472423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a target detection method, device, equipment, medium and product. Background Art
[0002] Current vehicles are equipped with numerous cameras and a wide range of intelligent driving functions, making them increasingly intelligent. The LAPA (Long-Range Autonomous Parking Assistant) beyond-visual-range memory parking function under the surround-view 4-way fisheye camera enables a more intelligent and convenient parking experience.
[0003] LAPA requires real-time object perception to provide information. Limited by performance and computing power, 2D image detection algorithms are currently the mainstream approach. Object detection in 2D images determines the object's bounding box and ground point. Then, through camera calibration parameters, the corresponding world distance and other information are calculated and output to the planning and control module, enabling real-time obstacle avoidance.
[0004] However, calculating the position of the ground point of an object in an image in the world coordinate system is highly dependent on its accuracy. Significant deviations from the ground point can lead to significant ranging errors, impacting subsequent planning and control logic decisions. This is especially true when the object is reflected, where the detected ground point is located. This can cause a significant discrepancy between the measured distance and the true value, leading to the failure of memory parking obstacle avoidance. Summary of the Invention
[0005] The present invention provides a target detection method, device, equipment, medium and product to realize the judgment of reflection scene and the correction of docking location when reflection exists.
[0006] According to a first aspect of the present invention, there is provided a target detection method, comprising:
[0007] Get image data;
[0008] Performing target detection on the image data to obtain a target detection result, wherein the target detection result includes an initial detection result and a reflection scene classification result;
[0009] When the reflection scene classification result is that a reflection exists, ground point regression is performed on the initial detection result to determine a corrected final target detection result.
[0010] According to a second aspect of the present invention, there is provided an object detection device, comprising:
[0011] A data acquisition module, used for acquiring image data;
[0012] A first determination module is configured to perform target detection on the image data to obtain a target detection result, wherein the target detection result includes an initial detection result and a reflection scene classification result;
[0013] The second determination module is used to perform ground point regression on the initial detection result when the reflection scene classification result is that there is a reflection, so as to determine a corrected final target detection result.
[0014] According to a third aspect of the present invention, there is provided an electronic device, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the target detection method described in any embodiment of the present invention.
[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the target detection method described in any embodiment of the present invention when executed.
[0019] According to a fifth aspect of the present invention, an embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the target detection method of any embodiment of the present invention.
[0020] The technical solution of the embodiment of the present invention is to obtain image data; perform target detection on the image data to obtain a target detection result, which includes an initial detection result and a reflection scene classification result; when the reflection scene classification result indicates the presence of a reflection, perform ground point regression on the initial detection result to determine a corrected final target detection result. By performing target detection on the image data and simultaneously performing reflection scene classification, in a scene with a reflection, performing ground point regression on the obtained detection result to determine the true ground point after excluding the influence of the reflection, and obtaining the final target monitoring result, the reflection data can be processed quickly and efficiently, thereby achieving accurate identification of the ground line and solving the problem of misidentification of target reflections.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a flowchart of a target detection method provided in accordance with the first embodiment of the present invention;
[0024] Figure 2 This is a flow chart of a target detection method provided according to the second embodiment of the present invention;
[0025] Figure 3 This is an example diagram of a reflection in a target detection method provided in Embodiment 2 of the present invention;
[0026] Figure 4 This is an example flow chart of a target detection method provided according to the second embodiment of the present invention;
[0027] Figure 5 2 is a schematic structural diagram of a target detection device provided according to a third embodiment of the present invention;
[0028] Figure 6 It is a schematic structural diagram of an electronic device implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] Example 1
[0032] Figure 1 A flow chart of a target detection method is provided for the first embodiment of the present invention. This embodiment is applicable to the accurate detection of images with reflections. The method can be performed by a target detection device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110: Acquire image data.
[0034] In this embodiment, the image data can be understood as a 2D image captured by a camera or other equipment.
[0035] Specifically, the processor may obtain image data captured by a camera or other device.
[0036] For example, in a vehicle scenario, a camera will be turned on to capture the parking environment during automatic parking. The vehicle's controller can receive image data captured by the camera to analyze the image data to determine obstacles around the vehicle. When the water quality or lighting on the ground causes obstacles to form reflections on the ground, if the obstacles in the reflection are mistakenly used as the basis for judging the parking distance, the distance between the vehicle and the obstacle will be inaccurately judged.
[0037] S120 , performing target detection on the image data to obtain a target detection result, where the target detection result includes an initial detection result and a reflection scene classification result.
[0038] In this embodiment, the target detection result can be understood as the target recognition result and scene classification result detected after analyzing the image data, where the initial detection result includes the identified target box and key points. The reflection scene classification result can be a binary classification result, indicating whether a reflection is present or absent in the scene.
[0039] Specifically, the processor can perform target detection on the image data through the optimized target detection module, perform target detection based on the feature fusion results in the target detection module, and obtain the target box and key points detected in the image data as the initial detection results. It can also perform scene classification through the feature fusion results to determine whether there are reflections in the image data, obtain the reflection scene classification results, and use the initial detection results and the reflection scene classification results as the target detection results.
[0040] S130: When the reflection scene classification result indicates that a reflection exists, ground point regression is performed on the initial detection result to determine a corrected final target detection result.
[0041] In this embodiment, the final target detection result can be understood as the result after correcting the key points.
[0042] Specifically, when the reflection scene classification result indicates the presence of a reflection, the key points in the initial detection result are usually identified as the grounding points in the reflection, which is inaccurate. The processor can perform grounding point regression on the feature map of the target box in the initial detection result, and perform reflection verification based on the corrected key points to determine whether there is a symmetry phenomenon when the corrected key points are used as grounding points, so as to verify again whether there is a reflection. When it is verified that there is a reflection, the corrected key points and target box are used as the corrected final target detection result. When there is no reflection, the initial detection result is used as the final target detection result.
[0043] The technical solution of the embodiment of the present invention performs reflection scene classification while performing target detection on image data. In scenes with reflections, the grounding point regression is performed on the detection results obtained to determine the true grounding point after excluding the influence of reflections, and the final target monitoring result is obtained. It can process reflection data quickly and efficiently, realize accurate identification of grounding lines, and solve the problem of misidentification of target reflections.
[0044] As a first optional embodiment of the first embodiment, based on the above embodiment, it further includes:
[0045] When the reflection scene classification result is that there is no reflection, the initial detection result is directly used as the final target detection result.
[0046] Specifically, when the reflection scene classification result is that there is no reflection, there is no need to perform the subsequent ground point regression step, and the initial detection result can be used as the final target detection result.
[0047] In a first optional embodiment of the first embodiment, when the classification result indicates that there is no reflection, the initial detection result is directly output, thereby improving target detection conditions in different situations.
[0048] Example 2
[0049] Figure 2 This is a flow chart of a target detection method provided in the second embodiment of the present invention. This embodiment is a further refinement of the above embodiment. Figure 2 As shown, the method includes:
[0050] S201: Acquire image data.
[0051] S202. Extract features from the image data through the backbone network of the target detection module to obtain a feature atlas.
[0052] In this embodiment, the object detection module can be understood as a deep learning model used to identify the location and range of specific objects in an image. It typically includes a backbone network, a neck network, and a head network. The backbone network is used to extract features from the image. The feature atlas can be understood as a collection of features extracted at different levels.
[0053] Specifically, the processor can input the image data into the target detection module, perform feature extraction on the image data through the backbone network of the target detection module, and obtain features extracted at different levels as a feature atlas.
[0054] For example, three feature maps, namely Feat1, Feat2 and Feat3, can be extracted from the image data through the backbone network to form a feature map set.
[0055] S203. Perform feature fusion on the feature atlas through the neck network of the target detection module to obtain a fused feature result set.
[0056] In this embodiment, the neck network can be understood as a network that fuses features at different levels to obtain a richer and more representative feature representation. The fused feature result set can be understood as a set of fused results.
[0057] Specifically, the neck network can be used to perform feature fusion on different feature maps in the feature map set to obtain fused feature results at different levels to form a fused feature result set.
[0058] For example, using Feat1, Feat2, and Feat3 obtained in the above example, the neck network can fuse Feat3 to obtain the fusion feature Fusion3, which is then fused upward with Feat2 to obtain Fusion2. Fusion2 is then fused with Feat1 to obtain Fusion1, which is then fused with Fusion11 to obtain Fusion11. Fusion11 is fused with Fusion2 to obtain Fusion22, and Fusion22 is fused with Fusion3 to obtain Fusion33. Fusion11, Fusion22, and Fusion33 are the fusion feature results, forming the fusion feature result set.
[0059] S204: Process the fusion feature result set through the head network of the target detection module to obtain an initial detection result.
[0060] In this embodiment, the head network can be understood as the part used for target detection, which maps the features provided by the neck network to the final output space to generate the final prediction result of the network.
[0061] Specifically, each fusion feature result in the fusion feature result set can be mapped through the head network to obtain multiple prediction results, and then the multiple prediction results are processed through corresponding algorithms (such as non-maximum suppression, soft non-maximum suppression and scoring, etc.), so as to select the final target box as the initial detection result of target detection in the image data.
[0062] Exemplarily, following the above example, Fusion11 obtains the prediction result Det_Head1 through the head network, Fusion22 obtains the prediction result Det_Head2 through the head network, and Fusion33 obtains the prediction result Det_Head3 through the head network, and then processes them through the preset algorithm to obtain the initial detection result.
[0063] S205 , performing scene classification on the fusion feature result set through the scene classification submodule of the target detection module to obtain a reflection scene classification result.
[0064] In this embodiment, the scene classification submodule can be understood as a module for scene classification.
[0065] Specifically, the scene classification submodule may perform scene classification on the target features in the fusion feature result set to obtain a reflection scene classification result for distinguishing whether there is a reflection in the scene.
[0066] Furthermore, based on the above embodiment, the scene classification submodule of the target detection module can be used to perform scene classification on the fusion feature result set to obtain the reflection scene classification result. The steps are refined as follows:
[0067] Filter the feature map to be classified from the fusion feature result set; use the scene classification submodule to classify the feature map according to the feature to be classified Figure 3 The height dimension of the three-dimensional channel is used to fuse the feature map to be classified up and down to obtain the upper output feature vector and the lower output feature vector; the scene classification submodule is used to fuse the feature map to be classified left and right according to the width dimension of the three-dimensional channel to obtain the left output feature vector and the right output feature vector; the scene is classified according to the upper output feature vector, the lower output feature vector, the left output feature vector and the right output feature vector to obtain the reflection scene classification result.
[0068] In this embodiment, the feature map to be classified can be understood as the fusion feature result classified in the fusion feature result set, and the fusion feature result at the bottom layer can be selected. The bottom layer has high aggregation characteristics and smaller pixels, so the processing speed is the fastest. The three-dimensional channel can be understood as a way of representing a color image in three dimensions, generally CHW, where C (Channel) represents the channel dimension, H (Height) represents the height dimension, and W (Width) represents the width dimension. The left-hand output feature vector and the right-hand output feature vector can be understood as the output features of each part after slicing in the height dimension. The upper-hand output feature vector and the lower-hand output feature vector can be understood as the output features of each part after slicing in the width dimension.
[0069] Specifically, the target detection module can filter the feature map to be classified from the fusion feature result set, and then use the scene classification submodule to classify the feature map according to the feature map to be classified. Figure 3 The height dimension of the three-dimensional channel is used to fuse the feature map to be classified up and down. For example, the height dimension can be sliced according to H (height dimension size) / 2, and the size of a single slice is the channel dimension size * width dimension size. The upper part can be processed in order from top to bottom, and the lower part can be processed in order from bottom to top. A convolution operation can be performed on each part, and the output results after convolution are connected through a fully connected FC layer to obtain the upper output feature vector and the lower output feature vector. The scene classification submodule is used to fuse the feature map to be classified left and right according to the width dimension of the three-dimensional channel. For example, the width dimension can be sliced according to W (width dimension size) / 2, and the size of a single slice is the channel dimension size * height dimension size. The left part can be processed in order from left to right, and the right part can be processed in order from right to left. A convolution operation can be performed on each part, and the output results after convolution are connected through a fully connected FC layer to obtain the left output feature vector and the right output feature vector. Scene classification is performed based on the upper output feature vector, the lower output feature vector, the left output feature vector, and the right output feature vector to obtain the reflection scene classification result.
[0070] For example, the feature map to be classified is the above-mentioned Fusion33, which is sliced according to H / 2 in the height dimension of the three-dimensional channel. The upper half is represented by P1 and the lower half is represented by P2. The size of each feature slice is C*W. Taking P1 as an example, the formula for processing the P1 part is as follows:
[0071]
[0072] Among them, Feat_U′ i,j,kRepresents the output feature vector of the upper part, i, j, k correspond to the CHW three-dimensional channel on the feature map Fusion33, K is the corresponding convolution kernel, a total of H / 2*1*3*3, Feat_U i,j,k Represents each feature slice in the upper part of the feature map to be classified, and performs convolution operation f() on each feature slice. The final output after convolution one by one is the output feature vector of the upper part, and the size is C*W; each output feature vector of the upper part is connected through the fully connected FC layer, and the obtained output feature vector of the upper half is 1*C in size.
[0073] The formula for processing the P2 part is as follows:
[0074]
[0075] Among them, Feat_D′ i,j,k Represents the output feature vector of the lower part, Feat_D i,j,k Represents each feature slice in the lower part of the feature map to be classified. Similarly, the output feature vectors of the left part and the right part can also be determined by similar formulas, only the H dimension is replaced by the W dimension.
[0076] S206: Use the initial detection result and the reflection scene classification result as the target detection result.
[0077] S207. When the reflection scene classification result is that there is a reflection, a correction key point is determined in the target feature map to which the target frame belongs according to the pre-trained prediction network and the initial detection result.
[0078] The initial detection results include a target bounding box and initial key points. The target bounding box is used to select the detected influential target. The initial key point can be understood as the connection point between the target identified by the target bounding box and the ground. Generally speaking, in scenes with reflections, the initial key point is identified as the lowest point of the reflection.
[0079] In this embodiment, the pre-trained prediction network can be understood as the network model used to determine key points. The target feature map can be understood as the feature map where the target box is located. The corrected key points can be understood as the actual key points after eliminating the influence of reflections, that is, the connection points between the target and the ground.
[0080] Specifically, when the reflection scene classification result is that there is a reflection, the processor can use the target feature map where the target box in the initial detection result is located as input, input it into the pre-trained prediction network for key point prediction, and use the predicted key point position as the correction key point determined in the target feature map to which the target box belongs.
[0081] Furthermore, based on the above embodiment, the step of determining the correction key points in the target feature map to which the target frame belongs according to the pre-trained prediction network and the initial detection results can be refined as follows:
[0082] The target feature map is predicted based on a pre-trained prediction network to obtain a predicted heat map; wherein, the pre-trained prediction network is integrated with a pre-trained heat map generation module; the predicted key point position information is extracted from the predicted heat map, and the corrected key points are determined in the target feature map to which the target box belongs.
[0083] In this embodiment, the predicted heat map can be understood as a visualization tool that maps data values to colors and uses color gradients to represent different data ranges to predict key points in the target feature map. The heat map generation module can be understood as a module for generating heat maps. The key point location information can be understood as the position coordinates of the predicted key points in the target feature map.
[0084] Specifically, the processor can input the target feature map into a pre-trained prediction network, whose main body can be, for example, a convolutional neural network feature extraction network (CNN), and directly use the highest value in the predicted heat map as the predicted key point position information, and determine the corrected key point in the target feature map based on the position information.
[0085] S208. Determine the reflection verification result according to the target feature map and the reflection verification threshold.
[0086] In this embodiment, the reflection verification threshold is a threshold set for determining whether a reflection exists.
[0087] Specifically, the processor can determine the reflection part and the target part in the image domain based on the corrected key points in the target feature map and the pixel size of the target frame, and determine the similarity between the two, and then compare them with the reflection verification threshold to determine whether the reflection is similar to the target, thereby obtaining the reflection verification result.
[0088] Furthermore, based on the above embodiment, the steps of determining the reflection verification result according to the target feature map and the reflection verification threshold can be refined as follows:
[0089] According to the target frame pixel attributes and correction key points of the target frame, the reflection pixel information and the target pixel information are determined in the target feature map; based on the reflection pixel information and the target pixel information, the similarity between the reflection and the actual target in the target feature map is determined; if the similarity is less than the reflection verification threshold, the existence of the reflection is taken as the reflection verification result; otherwise, the non-existence of the reflection is taken as the reflection verification result.
[0090] In this embodiment, the target frame pixel attributes can be understood as the attribute information used to characterize the pixel length and width of the target frame. The reflection pixel information can be understood as the pixels of the reflection part. The target pixel information can be understood as the pixels of the target part. The reflection can be understood as the part of the target feature map that is the reflection of the real target. The actual target can be understood as the part of the target feature map that is the real target. The similarity can be understood as a numerical value used to reflect the degree of similarity between the reflection and the actual target.
[0091] Specifically, the processor can use the corrected key point as the boundary position between the real target and the reflection in the target feature map, and determine the reflection pixel information of the reflection and the target pixel information of the real target according to the target frame pixel attributes of the target frame as the selection range of the real target and the reflection. The processor can determine the similarity between the reflection and the actual target in the target feature map based on the reflection pixel information and the target pixel information. The processor can compare the similarity with the reflection verification threshold. If the similarity is less than the reflection verification threshold, the existence of the reflection is regarded as the reflection verification result; otherwise, the non-existence of the reflection is regarded as the reflection verification result.
[0092] For example, using a specific example as a demonstration, Figure 3 This is an example diagram of a reflection in a target detection method provided in the second embodiment of the present invention, such as Figure 3 As shown, in the image domain, taking the corrected key point p1 determined by the above steps as an example, according to the target frame pixel attributes of the target frame, the pixel information of a certain pixel length and width in the upward and downward directions near it is intercepted for statistics, where the pixel length h can be selected as the length of 1 / n of the target frame in the target frame pixel attributes (for example, 1 / 4), that is, the position of the upper arrow in the figure, and the upward part of p1 is recorded as the actual target Pup, the size is (the width w in the target frame pixel attributes, that is, the red horizontal line in the figure) h*w, and the downward part is the reflection Pdown, the size is also h*w, the reflection Pdown part is rotated 180 degrees, and the rotation is recorded as PdownT. The processor can determine the similarity HistSim between Pup and PdownT by the following formula:
[0093]
[0094] Here, z represents a pixel.
[0095] S209: If the reflection verification result shows that a reflection exists, the corrected key points and target frame are used as the corrected final target detection result.
[0096] Specifically, if the reflection verification result shows that a reflection exists, the initial key points in the initial detection result are inaccurate, and the processor can use the corrected key points and target box as the corrected final target detection result.
[0097] S210: Otherwise, the initial detection result is used as the final target detection result.
[0098] Specifically, if the reflection verification result is that there is no reflection, there may be a misjudgment of the reflection scene. There is no need to correct the key points, and the initial detection result is the final target monitoring result.
[0099] The technical solution of the embodiment of the present invention is to obtain the reflection scene classification result by integrating the scene classification submodule into the target detection module, performing global feature fusion through the feature map output by the target detection module, and performing scene classification based on the fused features. The scene judgment after fusion is more accurate, and at the same time, feature reuse reduces performance consumption. The target feature map is predicted by a pre-trained prediction network, and the feature map output by the target detection module is also reused for prediction. The dedicated network of its regression module is small, and the key point position information is determined by predicting the heat map, so that the coordinates obtained by regression are more accurate to eliminate the influence of reflections. The correction key points are checked for the presence of reflections through the pixel attributes of the target frame to further confirm the situation of the reflection, thereby solving the problem of misidentification of target reflections, with low performance overhead and higher accuracy.
[0100] For example, a specific example is used as an example. Figure 4 This is an example flow chart of a target detection method provided in the second embodiment of the present invention, such as Figure 4 As shown, the steps may include:
[0101] S301, acquiring image data;
[0102] S302: Detect the image data through the target detection module to determine the target frame and the initial grounding point. The scene classification submodule performs scene classification on the fusion feature result set to obtain the reflection scene classification result.
[0103] S303: Check if there is a reflection scene in the reflection scene classification result; if so, jump to step S304; otherwise, jump to step S308;
[0104] S304: Perform ground point regression based on the pre-trained prediction network and the initial detection results, and determine the correction key points in the target feature map to which the target frame belongs;
[0105] S305: Determine the reflection verification result according to the target feature map and the reflection verification threshold;
[0106] S306: Check if the reflection verification result is that the reflection exists; if so, jump to step S307; if not, jump to step S308;
[0107] S307: Using the corrected key points and target frame as the corrected final target detection result;
[0108] S308: Use the initial detection result as the final target detection result.
[0109] Example 3
[0110] Figure 5 This is a schematic diagram of the structure of a target detection device provided by the third embodiment of the present invention. Figure 5 As shown, the device includes: a data acquisition module 51 , a first determination module 52 and a second determination module 53 .
[0111] A data acquisition module 51 is used to acquire image data;
[0112] A first determination module 52 is configured to perform target detection on the image data to obtain a target detection result, wherein the target detection result includes an initial detection result and a reflection scene classification result;
[0113] The second determination module 53 is configured to perform ground point regression on the initial detection result to determine a corrected final target detection result when the reflection scene classification result indicates that a reflection exists.
[0114] The technical solution of the embodiment of the present invention performs reflection scene classification while performing target detection on image data. In scenes with reflections, the grounding point regression is performed on the detection results obtained to determine the true grounding point after excluding the influence of reflections, and the final target monitoring result is obtained. It can process reflection data quickly and efficiently, realize accurate identification of grounding lines, and solve the problem of misidentification of target reflections.
[0115] Furthermore, the first determining module 52 includes:
[0116] A first determining unit is configured to extract features from the image data using a backbone network of an object detection module to obtain a feature atlas;
[0117] A second determining unit is configured to perform feature fusion on the feature atlas through the neck network of the target detection module to obtain a fused feature result set;
[0118] A third determining unit is configured to process the fusion feature result set through the head network of the target detection module to obtain an initial detection result;
[0119] a fourth determining unit, configured to perform scene classification on the fusion feature result set using the scene classification submodule of the object detection module to obtain a reflection scene classification result;
[0120] A fifth determining unit is configured to use the initial detection result and the reflection scene classification result as the target detection result.
[0121] The fourth determining unit is specifically configured to:
[0122] Filtering the feature graph to be classified from the fusion feature result set;
[0123] By using the scene classification submodule according to the features to be classified Figure 3 The feature map to be classified is fused up and down in the height dimension of the dimensional channel to obtain an upper output feature vector and a lower output feature vector;
[0124] By using the scene classification submodule, the feature map to be classified is fused left and right according to the width dimension of the three-dimensional channel to obtain a left output feature vector and a right output feature vector;
[0125] Scene classification is performed according to the upper output feature vector, the lower output feature vector, the left output feature vector, and the right output feature vector to obtain a reflection scene classification result.
[0126] Furthermore, the initial detection result includes a target frame and initial key points. Accordingly, the second determination module 53 includes:
[0127] A sixth determining unit, configured to determine, in the target feature map to which the target frame belongs, a correction key point based on the pre-trained prediction network and the initial detection result;
[0128] a seventh determining unit, configured to determine a reflection verification result according to the target feature map and the reflection verification threshold;
[0129] an eighth determining unit, configured to use the corrected key points and the target frame as a corrected final target detection result if the reflection verification result indicates that a reflection exists;
[0130] A ninth determining unit is configured to: otherwise, use the initial detection result as the final target detection result.
[0131] The sixth determining unit is specifically configured to:
[0132] Predicting the target feature map based on a pre-trained prediction network to obtain a predicted heat map; wherein the pre-trained prediction network is integrated with a pre-trained heat map generation module;
[0133] The predicted key point position information is extracted from the predicted heat map, and the corrected key points are determined in the target feature map to which the target box belongs.
[0134] The seventh determining unit is specifically configured to:
[0135] Determining the reflection pixel information and the target pixel information in the target feature map according to the target frame pixel attributes of the target frame and the correction key points;
[0136] Determining the similarity between the reflection and the actual target in the target feature map based on the reflection pixel information and the target pixel information;
[0137] If the similarity is less than the reflection verification threshold, the existence of the reflection is taken as the reflection verification result;
[0138] Otherwise, the absence of reflection is used as the reflection verification result.
[0139] Optionally, the device further includes:
[0140] The third determination module is configured to directly use the initial detection result as the final target detection result when the reflection scene classification result is that there is no reflection.
[0141] The target detection device provided in the embodiment of the present invention can execute the target detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0142] Example 4
[0143] Figure 6 A schematic diagram of the structure of an electronic device 60 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0144] like Figure 6As shown, the electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., which is communicatively connected to the at least one processor 61. The memory stores a computer program that can be executed by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or the computer program loaded from the storage unit 68 into the random access memory (RAM) 63. Various programs and data required for the operation of the electronic device 60 can also be stored in the RAM 63. The processor 61, ROM 62, and RAM 63 are connected to each other via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.
[0145] Multiple components in the electronic device 60 are connected to the I / O interface 65, including an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0146] The processor 61 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the target detection method.
[0147] In some embodiments, the target detection method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as a storage unit 68. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded into the RAM 63 and executed by the processor 61, one or more steps of the target detection method described above can be performed. Alternatively, in other embodiments, the processor 61 can be configured to perform the target detection method in any other suitable manner (e.g., by means of firmware).
[0148] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0149] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0150] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0152] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0153] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0154] In one embodiment, the present invention further includes a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, it implements the target detection method of any embodiment of the present invention.
[0155] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0157] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A target detection method, characterized in that: include: Get image data; Performing target detection on the image data to obtain a target detection result, wherein the target detection result includes an initial detection result and a reflection scene classification result; When the reflection scene classification result is that a reflection exists, ground point regression is performed on the initial detection result to determine a corrected final target detection result.
2. The method according to claim 1, characterized in that The performing target detection on the image data to obtain a target detection result includes: Performing feature extraction on the image data through the backbone network of the target detection module to obtain a feature atlas; Performing feature fusion on the feature atlas through the neck network of the target detection module to obtain a fused feature result set; Processing the fusion feature result set through the head network of the target detection module to obtain an initial detection result; Performing scene classification on the fusion feature result set by the scene classification submodule of the target detection module to obtain a reflection scene classification result; The initial detection result and the reflection scene classification result are used as the target detection result.
3. The method according to claim 2, characterized in that The performing scene classification on the fusion feature result set by the scene classification submodule of the target detection module to obtain a reflection scene classification result includes: Filtering the feature graph to be classified from the fusion feature result set; The scene classification submodule fuses the feature map to be classified up and down according to the height dimension of the three-dimensional channel of the feature map to be classified to obtain an upper output feature vector and a lower output feature vector; By using the scene classification submodule, the feature map to be classified is fused left and right according to the width dimension of the three-dimensional channel to obtain a left output feature vector and a right output feature vector; Scene classification is performed according to the upper output feature vector, the lower output feature vector, the left output feature vector, and the right output feature vector to obtain a reflection scene classification result.
4. The method according to claim 1, wherein The initial detection result includes a target frame and initial key points. Accordingly, performing ground point regression on the initial detection result to determine a corrected final target detection result includes: Determining a correction key point in a target feature map to which the target frame belongs based on a pre-trained prediction network and the initial detection result; Determining a reflection verification result according to the target feature map and the reflection verification threshold; If the reflection verification result is that a reflection exists, the corrected key points and the target frame are used as the corrected final target detection result; Otherwise, the initial detection result is used as the final target detection result.
5. The method according to claim 4, characterized in that The determining of the correction key points in the target feature map to which the target frame belongs based on the pre-trained prediction network and the initial detection result includes: Predicting the target feature map based on a pre-trained prediction network to obtain a predicted heat map; wherein the pre-trained prediction network is integrated with a pre-trained heat map generation module; The predicted key point position information is extracted from the predicted heat map, and the corrected key points are determined in the target feature map to which the target box belongs.
6. The method according to claim 4, characterized in that Determining the reflection verification result according to the target feature map and the reflection verification threshold includes: Determining the reflection pixel information and the target pixel information in the target feature map according to the target frame pixel attributes of the target frame and the correction key points; Determining the similarity between the reflection and the actual target in the target feature map based on the reflection pixel information and the target pixel information; If the similarity is less than the reflection verification threshold, the existence of the reflection is taken as the reflection verification result; Otherwise, the absence of reflection is used as the reflection verification result.
7. The method according to claim 1, characterized in that Also includes: When the reflection scene classification result is that there is no reflection, the initial detection result is directly used as the final target detection result.
8. A target detection device, characterized in that: include: A data acquisition module, used for acquiring image data; A first determination module is configured to perform target detection on the image data to obtain a target detection result, wherein the target detection result includes an initial detection result and a reflection scene classification result; The second determination module is used to perform ground point regression on the initial detection result when the reflection scene classification result is that there is a reflection, so as to determine a corrected final target detection result.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the target detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the target detection method according to any one of claims 1 to 7 when executed.
11. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the target detection method according to any one of claims 1 to 7.