Automatic driving target detection method and electronic equipment

Through the combination of the feature pyramid module, boundary extraction module and mask refinement module, high resolution semantic information and edge enhancement features are generated, which solves the problems of blurred boundary, high leakage detection rate of small targets and poor real-time performance in autonomous driving target detection, and achieves efficient target detection.

CN120564152APending Publication Date: 2025-08-29CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510692948.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing autonomous driving target detection methods have blurred boundaries, high detection rate of small targets, poor real-time performance and sensitive background interference in complex scenarios, making it difficult to meet the real-time requirements of vehicle-mounted embedded platforms.

Method used

The feature pyramid module is used to extract multi-scale semantic features, combine the boundary extraction module and mask refinement module, and generate the target mask through edge enhancement features and semantic features. Finally, the minimum rotation rectangle is generated through the post-processing module to achieve target detection.

Benefits of technology

It improves the detection accuracy of small targets, improves boundary continuity and positioning accuracy in complex scenarios, reduces the invalid area of ​​bounding boxes, and meets the real-time requirements of the vehicle-mounted embedded platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564152A_ABST
    Figure CN120564152A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving target detection method and electronic equipment, and the method comprises the steps: obtaining a to-be-detected image, inputting the to-be-detected image into a pre-trained target detection model, extracting multi-scale semantic features through a feature pyramid module in the model, and carrying out the extraction of the multi-scale semantic features; then determining multi-scale edge enhancement features through a boundary extraction module, the to-be-detected image and the multi-scale semantic features in the model, further determining a target mask through a mask refining module, the multi-scale semantic features and the multi-scale edge enhancement features in the model, and finally performing post-processing on the target mask through a post-processing module in the model. And determining a minimum rotation rectangle corresponding to the mask contour in the target mask, and obtaining a target bounding box, thereby realizing target detection, improving the detection accuracy of a small target and the boundary continuity in a complex scene, solving the problems of boundary blur and background interference sensitivity, avoiding excessive smoothness, improving the positioning precision of a deformed target, and improving the positioning accuracy of the deformed target. And the target detection real-time performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to an autonomous driving target detection method and electronic equipment. Background Art

[0002] With the rapid development of autonomous driving technology, object detection plays a vital role in autonomous driving systems. To ensure that autonomous vehicles can accurately perceive surrounding objects in complex environments, object recognition and perception systems must possess high precision, high robustness, and the ability to adapt to various driving environments.

[0003] Existing deep learning-based object detection frameworks (such as Mask R-CNN and the YOLO series) mainly use the following technologies: 1. General mask generation: extracting simple feature maps from the backbone network through regions of interest and predicting the target mask based on them; 2. Traditional optimization methods: relying solely on image morphological post-processing for mask smoothing to suppress noise.

[0004] However, the above-mentioned related technologies have the following problems: 1. Boundary fuzziness: In complex scenarios (such as target occlusion and lighting changes), jagged artifacts or breakage are prone to appear at the mask boundary; 2. High small target missed detection rate: The mask prediction integrity of small targets (such as pedestrians and traffic signs) is insufficient, resulting in failure to identify key obstacles; 3. Real-time bottleneck: The multi-level network structure (such as Mask R-CNN) has a large number of parameters and is difficult to meet the real-time requirements of the vehicle embedded platform; 4. Sensitivity to background interference: Dynamic backgrounds in road scenes (such as swaying trees and ground reflections) can easily cause false detections, reducing system reliability. Summary of the Invention

[0005] In view of the above-mentioned defects or deficiencies in the prior art, the present application aims to provide an autonomous driving target detection method and electronic equipment to solve the problems in the related technology such as blurred boundaries, high missed detection rate of small targets, poor real-time performance, and sensitivity to background interference.

[0006] This embodiment of the present application provides a method for detecting an object in an autonomous driving process, the method comprising: Acquire an image to be detected and input the image to be detected into a pre-trained target detection model, wherein the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module, and a post-processing module; Extracting multi-scale semantic features from the image to be detected based on the feature pyramid module, and determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features; Determining a target mask based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features; A minimum rotated rectangle corresponding to the mask outline in the target mask is determined based on the post-processing module to obtain a target bounding box.

[0007] Optionally, the feature pyramid module includes a bottom-up path and a top-down path, and extracting multi-scale semantic features from the image to be detected based on the feature pyramid module includes: Downsampling the image to be detected based on the bottom-up path to obtain a multi-scale first feature; Upsampling the multi-scale first features based on the top-down path to obtain multi-scale second features; Based on the top-down path, the first features and the second features of each scale are horizontally connected, and convolution processing is performed on the multi-scale features after the horizontal connection to obtain multi-scale semantic features.

[0008] Optionally, the boundary extraction module includes a basic feature extraction unit, an edge detection unit, an edge enhancement unit, and an alignment unit. Determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features includes: Extracting multi-scale basic features from the image to be detected based on the basic feature extraction unit; Extracting initial gradient features from the multi-scale basic features based on the edge detection unit; Performing enhancement processing on the initial gradient feature based on the edge enhancement unit to obtain a final gradient feature; The final gradient feature and the multi-scale semantic feature are fused based on the alignment unit to obtain a multi-scale edge enhancement feature.

[0009] Optionally, performing enhancement processing on the initial gradient feature based on the edge enhancement unit to obtain a final gradient feature includes: Performing a dilated convolution on the initial gradient features based on the edge enhancement unit to obtain multi-scale intermediate gradient features; The multi-scale intermediate gradient features are spliced ​​and compressed based on the edge enhancement unit to obtain final gradient features.

[0010] Optionally, fusing the final gradient feature with the multi-scale semantic feature based on the alignment unit to obtain a multi-scale edge enhancement feature includes: Based on the alignment unit, the final gradient features and the multi-scale semantic features are spliced ​​and normalized to obtain weights corresponding to the semantic features of each scale, and the weights corresponding to the semantic features of each scale are used to perform weighted processing on the final gradient features and the semantic features of each scale to obtain multi-scale edge enhancement features.

[0011] Optionally, the mask refinement module includes a feature fusion unit, a compensation unit, and a progressive refinement unit, and determines a target mask based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features, including: splicing the multi-scale semantic features and the multi-scale edge enhancement features based on the feature fusion unit, and performing channel pruning processing on the spliced ​​features using attention weights to obtain target features; Performing bilinear interpolation processing on the target feature based on the compensation unit to obtain an initial mask; The initial mask is iteratively refined based on the progressive refinement unit until the number of iterations reaches a preset threshold, thereby obtaining a target mask.

[0012] Optionally, the preset number threshold is determined by the following steps: During the training of the target detection model, the initial mask of the sample is iteratively refined multiple times based on the progressive refinement unit to obtain multiple masks to be evaluated; For each mask to be evaluated, determining a mask distance confidence, a mask intersection-over-union confidence, and a classification confidence between the mask to be evaluated and the actual annotated mask corresponding to the sample initial mask, and determining a mask score corresponding to the mask to be evaluated based on the mask distance confidence, the mask intersection-over-union confidence, and the classification confidence; The number of iterations corresponding to the mask to be evaluated with the highest mask score is determined as the preset number threshold.

[0013] Optionally, the post-processing module includes a mask optimization unit, and before determining the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module, further includes: Determining a mask area of ​​the target mask based on the mask optimization unit, and determining a size of a structural element used for dilation corrosion according to the mask area to obtain the structural element; Based on the mask optimization unit and the structural element, the target mask is eroded and expanded.

[0014] Optionally, the post-processing module includes a bounding box generation unit, and determining a minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module includes: The bounding box generation unit determines a rotation angle of the mask outline in the target mask, and rotates the target mask according to the rotation angle to generate a minimum rotated rectangle containing the rotated mask outline.

[0015] An embodiment of the present application further provides an electronic device, comprising: processor and memory; The processor is used to execute the steps of the autonomous driving target detection method provided in any embodiment of the present application by calling the program or instructions stored in the memory.

[0016] An embodiment of the present application also provides a computer-readable storage medium, which stores a program or instruction, and the program or instruction enables a computer to execute the steps of the autonomous driving target detection method provided by any embodiment of the present application.

[0017] In summary, the present application proposes a method for target detection in autonomous driving, which obtains an image to be detected and inputs it into a pre-trained target detection model. The method first extracts multi-scale semantic features through the feature pyramid module in the model, and then determines the multi-scale edge enhancement features through the boundary extraction module, the image to be detected and the multi-scale semantic features in the model. Then, the target mask is determined through the mask refinement module, the multi-scale semantic features and the multi-scale edge enhancement features in the model. Finally, the minimum rotated rectangle corresponding to the mask contour in the target mask is determined through the post-processing module in the model to obtain the target bounding box and realize target detection. By extracting multi-scale semantic features, this method can It obtains high-resolution semantic information, solves the scale jump problem, and improves the detection accuracy of small targets. The boundary extraction module extracts edge enhancement features, which can improve the boundary continuity in complex scenes and solve the problems of boundary blur and sensitivity to background interference. In addition, this method generates a target mask by combining semantic features and edge enhancement features through a mask refinement module, which can avoid excessive smoothing and improve the positioning accuracy of deformed targets. The post-processing module generates a minimum rotated rectangle, reduces the invalid area in the bounding box, and facilitates the improvement of the processing efficiency of subsequent autonomous driving functions. In addition, the number of parameters of each module of the model in this method is relatively small, which can improve the real-time performance of target detection and meet the real-time requirements of the vehicle embedded platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 This is a flowchart of an autonomous driving target detection method provided in Example 1 of the present application; Figure 2 This is a schematic diagram of a process for extracting semantic features using a feature pyramid module provided in Example 2 of the present application; Figure 3This is a schematic diagram of a process in which a boundary extraction module outputs edge enhancement features, as provided in Example 3 of the present application; Figure 4 This is a schematic diagram of a process for determining a target mask by a mask refinement module provided in the fourth embodiment of the present application; Figure 5 This is a schematic diagram of a process for generating a target bounding box by a target detection model provided in Example 5 of the present application; Figure 6 This is a structural diagram of an electronic device provided in Example 6 of the present application. DETAILED DESCRIPTION

[0020] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.

[0021] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0022] As mentioned in the background technology, in response to the problems in the existing technology, this application proposes an autonomous driving target detection method, which can be applied to the vehicle's autonomous driving system. It detects the target in the image to be detected and then makes a decision based on the detected target bounding box.

[0023] Example 1 Figure 1 This is a flow chart of an autonomous driving target detection method provided in Example 1 of this application. Figure 1 , the autonomous driving target detection method specifically includes: S110: Obtain an image to be detected, and input the image to be detected into a pre-trained object detection model.

[0024] The image to be detected may be an image of the surrounding environment collected by the autonomous driving system in the vehicle. For example, the image to be detected may be an RGB image (resolution ≥ 1920 × 1080) collected by a camera or a depth map generated by a lidar point cloud projection.

[0025] Specifically, after acquiring the image to be detected, the image to be detected can be input into a pre-trained object detection model, which includes a feature pyramid module, a boundary extraction module, a mask refinement module, and a post-processing module.

[0026] In the object detection model, the feature pyramid module is connected to the boundary extraction module, which is then connected to the mask refinement module, which is then connected to the post-processing module. Specifically, after the image to be detected enters the object detection model, it can first enter the feature pyramid module. The output of the feature pyramid module can then enter the boundary extraction module together with the image to be detected. The output of the boundary extraction module and the output of the feature pyramid module can then enter the mask refinement module together. Finally, the output of the mask refinement module enters the post-processing module.

[0027] S120 , extracting multi-scale semantic features from the image to be detected based on the feature pyramid module, and determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features.

[0028] Specifically, after the image to be detected enters the target detection model, it can first pass through the feature pyramid module, which extracts multi-scale semantic features in the image to be detected.

[0029] The feature pyramid module can be used to extract hierarchical structures of features at different resolutions, obtaining shallow high-resolution features and deep low-resolution features, and using the features of different resolutions as multi-scale semantic features. The feature pyramid module can use a lightweight pyramid network, including a small number of convolutional layers and a simple feature fusion structure.

[0030] For example, the feature pyramid module can use multiple convolution kernels of different scales, each of which simultaneously convolves the image to be detected to extract semantic features of different scales. Alternatively, the feature pyramid module can use multiple convolution kernels of the same scale, each of which is connected in series. The first convolution kernel convolves the image to be detected, and the other convolution kernels sequentially convolve the output of the previous convolution kernel to extract semantic features of different scales through the output of each convolution kernel.

[0031] After obtaining multi-scale semantic features through the feature pyramid module, the multi-scale semantic features and the image to be detected can be further input into the boundary extraction module, and the boundary extraction module determines the multi-scale edge enhancement features.

[0032] Among them, the edge extraction module can be used to extract edge information from the image to be detected, and combine multi-scale semantic features to obtain multi-scale edge enhancement features to highlight the edge information in the image to be detected.

[0033] Exemplarily, the edge extraction module can first extract distinct contours from the image to be detected using a preset operator (such as the Sobel operator, Prewitt operator, Roberts operator, Laplacian operator, etc.), output initial edge information, and then blur the image to be detected (e.g., by Gaussian filtering) to obtain its low-frequency component. The low-frequency component is then subtracted from the image to be detected to obtain the high-frequency edge. The high-frequency edge is then superimposed with the initial edge information (e.g., by weighted fusion) to obtain enhanced edge information. Furthermore, the edge extraction module can fuse the enhanced edge information with multi-scale semantic features to obtain multi-scale edge enhancement features.

[0034] S130 : Determine a target mask based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features.

[0035] Specifically, after obtaining multi-scale edge enhancement features, the multi-scale semantic features and multi-scale edge enhancement features can be fed into the mask refinement module, which performs feature fusion and generates a target mask based on the fusion results. The semantic features and edge enhancement features have the same number of scales, and the target mask can be a binary image containing the target edge (i.e., the mask outline).

[0036] The mask refinement module can be used to generate an initial mask and refine the initial mask to obtain a target mask, thereby achieving the purpose of refining the target contour in the mask.

[0037] Specifically, the mask refinement module can first fuse multi-scale semantic features and multi-scale edge enhancement features, obtain an initial mask based on the fusion result, and then the mask refinement module can refine the initial mask to obtain a target mask.

[0038] Exemplarily, the mask refinement module can perform high-pass filtering on the image to be detected to obtain edge reference information, and then use the edge reference information to refine the initial mask, such as superimposing the initial mask with the edge reference information, or calculating the difference between the edge reference information and the initial mask, and using the difference as a residual increment to superimpose it on the initial mask to obtain the target mask.

[0039] It should be noted that, in the process of refining the initial mask to obtain the target mask, iterative refinement can be continuously performed, and the mask finally obtained through multiple iterative refinements is used as the target mask. For example, after refining the initial mask in the above manner, the obtained mask can be further refined, and so on to obtain the target mask.

[0040] S140 : Determine, based on the post-processing module, a minimum rotated rectangle corresponding to the mask outline in the target mask to obtain a target bounding box.

[0041] Among them, the post-processing module can be used to construct a target bounding box according to the mask outline in the target mask to obtain the target detection result.

[0042] Specifically, after obtaining the target mask, the post-processing module may generate a minimum rotated rectangle corresponding to the mask outline in the target mask, and use the minimum rotated rectangle as the target bounding box.

[0043] The mask contour may refer to an edge of an object detected in the target mask, and the minimum rotated rectangle may refer to a minimum rectangle that surrounds the mask contour through rotation.

[0044] Exemplarily, the post-processing module can first locate the mask contour in the target mask, then separate the mask contour from the target mask, and then rotate the separated mask contour so that the central axis of the mask contour is aligned with the central axis of the image to be detected, thereby generating a minimum rectangle surrounding the mask contour, and rotating the minimum rectangle to the initial angle of the mask contour to obtain a minimum rotated rectangle.

[0045] The automatic driving target detection method provided by the embodiment of the present application obtains the image to be detected and inputs it into the pre-trained target detection model. The multi-scale semantic features are first extracted through the feature pyramid module in the model. Then, the multi-scale edge enhancement features are determined through the boundary extraction module in the model, the image to be detected and the multi-scale semantic features. Then, the target mask is determined through the mask refinement module, the multi-scale semantic features and the multi-scale edge enhancement features in the model. Finally, the minimum rotated rectangle corresponding to the mask outline in the target mask is determined through the post-processing module in the model to obtain the target bounding box and realize target detection. This method can obtain the target bounding box by extracting the multi-scale semantic features. High-resolution semantic information solves the scale jump problem and improves the detection accuracy of small targets. The edge enhancement features extracted by the boundary extraction module can improve the boundary continuity in complex scenes and solve the problems of boundary blur and sensitivity to background interference. In addition, this method generates a target mask by combining semantic features and edge enhancement features through the mask refinement module, which can avoid excessive smoothing and improve the positioning accuracy of deformed targets. The minimum rotated rectangle is generated by the post-processing module to reduce the invalid area in the bounding box, which is convenient for improving the processing efficiency of subsequent autonomous driving functions. In addition, the number of parameters of each module of the model in this method is relatively small, which can improve the real-time performance of target detection and meet the real-time requirements of the vehicle embedded platform.

[0046] Example 2 Based on the above embodiments, the feature pyramid module can adopt a bottom-up path and a top-down path to achieve multi-scale feature fusion and semantic detail complementation based on a bidirectional architecture. The method includes the following steps: S210: Obtain an image to be detected, and input the image to be detected into a pre-trained target detection model.

[0047] Among them, the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module and a post-processing module. The feature pyramid module includes a bottom-up path and a top-down path.

[0048] Specifically, the feature pyramid module can be composed of a bottom-up path and a top-down path. The bottom-up path is used to gradually transition from low-level features in the image to be detected to high-level features, extracting global features through gradual downsampling, such as using ResNet-50. The top-down path is used to gradually transfer global information from high-level semantic features to low-level features to enhance semantic understanding of detailed locations.

[0049] S220 , down-sample the image to be detected based on the bottom-up path to obtain a multi-scale first feature, and up-sample the multi-scale first feature based on the top-down path to obtain a multi-scale second feature.

[0050] Specifically, after the image to be detected enters the bottom-up path, it can pass through multiple downsampling stages of the bottom-up path in sequence to generate first features of different scales respectively, and the spatial resolution of the features output between adjacent downsampling stages satisfies a 2-fold decreasing relationship.

[0051] For example, we can take the four downsampling stages as an example. After the image to be detected enters the bottom-up path, the first features of the four scales are generated, which are expressed as , the corresponding spatial resolutions are respectively , the number of channels of the first feature output in each downsampling stage is .

[0052] Furthermore, the multi-scale first features output by the bottom-up path can enter the top-down path, and the high-level features are sequentially upsampled by each upsampling stage in the top-down path to obtain the multi-scale second features.

[0053] For example, the top-most upsampling stage in the top-down path can directly output the first feature with the highest scale as the second feature, and then the remaining upsampling stages sequentially perform 2x bilinear interpolation upsampling on the first feature output by the previous downsampling stage to obtain the second feature. For example, taking four upsampling stages as an example, the second feature output by the top-most upsampling stage is , for other upsampling stages, the second feature of the output is expressed as: ; Where, is the second feature output by the i-th upsampling stage, is the first feature output by the i+1th downsampling stage, Indicates upsampling processing.

[0054] S230: Laterally connect the first features and the second features of each scale based on a top-down path, and perform convolution processing on the multi-scale features after the horizontal connection to obtain multi-scale semantic features.

[0055] After obtaining the second features at multiple scales, the top-down path can further connect the first features and second features at the same scale horizontally, that is, splicing the first features and second features at the same level. Taking the four stages as an example, it is shown in the following formula: ; Where, is the horizontal connection feature of the i-th scale, Indicates that the first feature of the i-th scale is convolved using a 1×1 convolution kernel.

[0056] After connecting the first and second features of the same scale, convolution processing can be performed on the multi-scale features after horizontal connection to eliminate the aliasing effect and obtain multi-scale semantic features. Taking the four stages as an example, it is shown in the following formula: ; Where, is the semantic feature of the i-th scale, Indicates that the 3×3 convolution kernel is used to convolve the horizontal connection features of the i-th scale. The final output is The output multi-scale semantic features not only have high-resolution location information, but also contain deep semantic information, providing multi-scale perception capabilities for subsequent target detection.

[0057] Figure 2 This is a schematic diagram of a process of extracting semantic features using a feature pyramid module provided in the second embodiment of the present application. Figure 2 As shown in Figure 2, the bottom-up path in the feature pyramid module can extract features from low layers to high layers, while the top-down path can extract features from high layers to low layers and connect them laterally with features at the same level (i.e., the same scale). The top-down path can first perform a 2x upsampling, and then add the 2x upsampling result to the 1×1 convolution result of the first feature at the same scale.

[0058] S240 : Determine multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features.

[0059] S250: Determine a target mask based on a mask refinement module, multi-scale semantic features, and multi-scale edge enhancement features.

[0060] S260 : Determine the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module to obtain a target bounding box.

[0061] The autonomous driving target detection method provided in the embodiments of the present application realizes multi-scale feature fusion and semantic detail complementation based on a bidirectional architecture. Through bottom-up and top-down paths, it can extract features at different levels (low to high), covering multi-scale information from details to semantics, and transmit high-level semantic information to low levels, thereby compensating for the semantic deficiencies of shallow features, solving the scale jump problem, and generating a feature pyramid with both high semantics and high resolution, adapting to targets of different sizes, and significantly enhancing the model's perception of complex scenes.

[0062] Example 3 Based on the above embodiments, the boundary extraction module can use a basic feature extraction unit, an edge detection unit, an edge enhancement unit, and an alignment unit to first extract basic features, then extract initial gradient features through the basic features, enhance the initial gradient features through edge enhancement features, align the final gradient features obtained with the semantic features, and obtain edge enhancement features to achieve the purpose of suppressing background noise and enhancing edges. The method includes the following steps: S310: Obtain an image to be detected, and input the image to be detected into a pre-trained object detection model.

[0063] Among them, the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module and a post-processing module; the boundary extraction module includes a basic feature extraction unit, an edge detection unit, an edge enhancement unit and an alignment unit.

[0064] Specifically, the boundary extraction module can be composed of a basic feature extraction unit, an edge detection unit, an edge enhancement unit, and an alignment unit. The image to be detected can be fed into the basic feature extraction unit, the output of which is fed into the edge detection unit, the output of which is fed into the edge enhancement unit, and the output of the edge enhancement unit, along with multi-scale semantic features, is fed into the alignment unit.

[0065] S320: Extract multi-scale semantic features from the image to be detected based on the feature pyramid module.

[0066] S330 , extracting multi-scale basic features from the image to be detected based on the basic feature extraction unit, and extracting initial gradient features from the multi-scale basic features based on the edge detection unit.

[0067] Specifically, after the image to be detected enters the basic feature extraction unit, the basic feature extraction unit can extract multi-scale basic features. Furthermore, the multi-scale basic features enter the edge detection unit, which can use a parameter-adaptive 3×3 convolution kernel to extract initial gradient features from the multi-scale basic features.

[0068] For example, the edge detection unit may first use a 3×3 convolution kernel to perform convolution processing on the basic features of each scale, and then perform nonlinear activation on the convolved features of all scales to obtain the initial gradient features, as shown in the following formula: ; Where, is the initial gradient feature, is the basic feature of the k-th scale, For nonlinear activation function, ReLu (Rectified Linear Unit) can be used. are the learnable convolution kernel parameters, is the number of scales.

[0069] It should be noted that the edge detection unit uses learnable convolution kernel parameters to extract initial gradient features. Compared with edge detection using fixed operators (such as Canny), it can achieve the purpose of dynamically optimizing edges according to the scene and ensure the accuracy of edge detection.

[0070] S340 , enhancing the initial gradient features based on the edge enhancement unit to obtain final gradient features, and fusing the final gradient features with multi-scale semantic features based on the alignment unit to obtain multi-scale edge enhancement features.

[0071] Furthermore, the initial gradient features may be enhanced by an edge enhancement unit to obtain a final gradient feature. For example, the edge enhancement unit may enhance the initial gradient features by using a pre-trained edge enhancer or edge enhancement operator to obtain a final gradient feature.

[0072] In one example, the initial gradient feature is enhanced based on the edge enhancement unit to obtain the final gradient feature, including the following steps: Step 341: Perform a dilated convolution on the initial gradient features based on the edge enhancement unit to obtain multi-scale intermediate gradient features; Step 342: Based on the edge enhancement unit, the multi-scale intermediate gradient features are spliced ​​and compressed to obtain the final gradient features.

[0073] In step 341, the edge enhancement unit may first process the initial gradient features through dilated convolution to expand the receptive field and capture the multi-level edge structure. The edge enhancement unit may use dilated convolution with different expansion rates to process the initial gradient features respectively, thereby obtaining intermediate gradient features of different scales, as shown in the following formula: ; Where, is the intermediate gradient feature of the j-th scale, Represents a convolution operation with a dilation rate of d. Taking 4 dilation rates as an example, , .

[0074] Furthermore, in step 342 , the edge enhancement unit may concatenate and compress the intermediate gradient features of all scales to capture fine-grained edges (through a small dilation rate) and long-range contours (through a large dilation rate) to obtain the final gradient features.

[0075] For example, the edge enhancement unit can first concatenate the intermediate gradient features of all scales, and then use a 1×1 convolution kernel to convolve the concatenated results to complete the compression. Taking the intermediate gradient features of 4 scales as an example, as shown in the following formula: ; Where, is the final gradient feature, Indicates that convolution is performed using a 1×1 convolution kernel. is the intermediate gradient feature of the j-th scale.

[0076] In the above steps 341 and 342, by performing dilated convolution on the initial gradient features and then performing splicing and compression, it is possible to capture multi-level edge structures and obtain final gradient features containing fine-grained edges and long-range contours, thereby further improving the accuracy of target detection.

[0077] After obtaining the final gradient feature, the final gradient feature can be further aligned with the semantic feature through an alignment unit, that is, the final gradient feature is fused with the multi-scale semantic feature to obtain a multi-scale edge enhancement feature.

[0078] For example, the final gradient feature and the semantic feature can be directly concatenated to obtain the edge enhancement feature; or, the final gradient feature and the semantic feature can be weighted using a preset weight to obtain the weighted edge enhancement feature.

[0079] In one example, the final gradient features are fused with multi-scale semantic features based on the alignment unit to obtain multi-scale edge enhancement features, including: Based on the alignment unit, the final gradient features and multi-scale semantic features are spliced ​​and normalized to obtain the weights corresponding to the semantic features of each scale. The weights corresponding to the semantic features of each scale are then used to perform weighted processing on the final gradient features and the semantic features of each scale to obtain multi-scale edge enhancement features.

[0080] Specifically, for each scale of semantic features, the alignment unit can first concatenate and normalize the semantic features with the final gradient features to obtain the weight corresponding to the semantic features. For example, the alignment unit can first concatenate the semantic features with the final gradient features, then convolve the concatenation result using a 3×3 convolution kernel, and normalize the convolution result using an activation function to obtain the corresponding weight. This is shown in the following formula: ; Where, for The corresponding weight, is the semantic feature of the i-th scale, is the final gradient feature, Represents the use of 3×3 convolution kernel for convolution, Indicates channel splicing processing, is a nonlinear activation function, Indicates that a nonlinear activation function is used for processing.

[0081] After obtaining the weights corresponding to the semantic features of each scale, the final gradient features and semantic features can be fused according to the weights corresponding to the semantic features for each semantic feature to obtain edge enhancement features of different scales. Taking the semantic features of four scales as an example, the following formula is shown: ; Where, is the edge enhancement feature of the i-th scale, is the semantic feature of the i-th scale, for The corresponding weight, is the final gradient feature, Indicates element-by-element multiplication. Based on this method, edge enhancement features of four scales can be obtained, namely .

[0082] By concatenating and normalizing the final gradient features with multi-scale semantic features, we can obtain the weights corresponding to the semantic features of different scales. This allows us to suppress background noise in the semantic features by weighted fusion of the semantic features and the final gradient features. At the same time, we can enhance the feature response of edge-related areas, obtain aligned edge enhancement features, and improve the accuracy of target detection.

[0083] Figure 3 FIG. 1 is a schematic diagram of a process of outputting edge enhancement features of a boundary extraction module provided in an embodiment of the present application, such as Figure 3 As shown, the image to be detected can first enter the basic feature extraction unit to extract multi-scale basic features, and then the basic features enter the edge detection unit to extract the initial gradient features. The initial gradient features enter the edge enhancement unit to extract the final gradient features. The final gradient features and semantic features enter the alignment unit together, and the edge enhancement features are obtained through splicing, normalization and weighting operations.

[0084] S350: Determine a target mask based on a mask refinement module, multi-scale semantic features, and multi-scale edge enhancement features.

[0085] S360: Determine the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module to obtain a target bounding box.

[0086] The autonomous driving target detection method provided in the embodiment of the present application replaces fixed operators with learnable convolution kernels, adaptively optimizes gradient responses, realizes dynamic edge detection, and fuses local details with global contours through void convolution to solve the problem of long-range structural fractures. In addition, the semantic features are weighted through spatial attention (i.e., the weights corresponding to the semantic features), suppressing background noise and enhancing edges, which can further improve the accuracy of target detection.

[0087] Example 4 Based on the above embodiments, the mask refinement module can use a feature fusion unit, a compensation unit, and a progressive refinement unit to first fuse semantic features with edge enhancement features, then compensate the fused result and refine the target mask to improve the accuracy of the target outline in the mask. The method includes the following steps: S410: Obtain an image to be detected, and input the image to be detected into a pre-trained object detection model.

[0088] Among them, the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module and a post-processing module; the mask refinement module includes a feature fusion unit, a compensation unit and a progressive refinement unit.

[0089] Specifically, the mask refinement module can be composed of a feature fusion unit, a compensation unit, and a progressive refinement unit. Multi-scale semantic features and multi-scale edge enhancement features can be fed into the feature fusion module, the output of which feeds the compensation unit, and the output of which feeds the progressive refinement unit.

[0090] S420 , extracting multi-scale semantic features from the image to be detected based on the feature pyramid module, and determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features.

[0091] S430: Based on the feature fusion unit, the multi-scale semantic features and the multi-scale edge enhancement features are spliced ​​together, and the attention weights are used to perform channel pruning on the spliced ​​features to obtain target features.

[0092] Specifically, the feature fusion unit can first splice the semantic features and edge enhancement features of the same scale, and then use the activation function to process the attention weights. The processed attention weights are multiplied by the spliced ​​features to dynamically suppress redundant channels, reduce the number of parameters, achieve the purpose of channel pruning, and obtain the target features: ; Where, is the target feature, is the attention weight, is the edge enhancement feature of the i-th scale, is the semantic feature of the i-th scale, Indicates channel splicing, Represents element-wise multiplication.

[0093] S440 , performing bilinear interpolation processing on the target feature based on the compensation unit to obtain an initial mask, and iteratively refining the initial mask based on the progressive refinement unit until the number of iterations reaches a preset threshold, thereby obtaining a target mask.

[0094] Furthermore, the compensation unit may perform bilinear interpolation processing on the target features, that is, interpolation is performed in the feature matrix corresponding to the target features, and an initial mask is determined based on the interpolation result. The initial mask may include edge information of the target.

[0095] Exemplarily, the compensation unit can first perform sub-pixel offset compensation, regard each element in the feature matrix corresponding to the target feature as a coordinate, create sub-pixel coordinates between each coordinate based on a preset offset (which can be obtained through learning), and perform bilinear interpolation on the eigenvalue of each element to obtain the eigenvalue of the sub-pixel coordinate, such as performing two linear interpolations in the horizontal direction and then one linear interpolation in the vertical direction.

[0096] Furthermore, the progressive refinement unit may iteratively refine the initial mask through an iterator until the number of iterations reaches a preset number threshold, thereby obtaining a target mask. The preset number threshold may be an iteration number threshold obtained by training the iterator.

[0097] Exemplarily, the progressive refinement unit can perform bilinear interpolation on the initial mask to iteratively refine the initial mask, thereby obtaining a first refined mask and updating the number of iterations, and then perform bilinear interpolation on the first refined mask to obtain a second refined mask and update the number of iterations, repeating this operation until the number of iterations reaches a preset threshold, and taking the last refined mask as the target mask.

[0098] Figure 4 FIG. 1 is a schematic diagram of a process in which a mask refinement module determines a target mask according to an embodiment of the present application. Figure 4 As shown in the figure, multi-scale semantic features and multi-scale edge enhancement features can enter the feature fusion unit, and after splicing and channel pruning, the target features are obtained. The target features enter the compensation unit, and through sub-pixel offset and bilinear interpolation, the initial mask is obtained. The initial mask then enters the progressive refinement unit, and the target mask is generated through the iterator.

[0099] In an embodiment of the present application, the preset number threshold can be an empirical value, or the preset number threshold can be obtained in advance during the target detection model training process, by generating multiple refined masks and then evaluating each refined mask to determine the iteration number threshold of the iterator.

[0100] In one example, the preset number threshold is determined by the following steps: Step 441: During the object detection model training process, the sample initial mask is iteratively refined multiple times based on the progressive refinement unit to obtain multiple masks to be evaluated; Step 442: For each mask to be evaluated, determine the mask distance confidence, mask intersection-over-union confidence, and classification confidence between the mask to be evaluated and the actual annotated mask corresponding to the sample initial mask. Determine the mask score corresponding to the mask to be evaluated based on the mask distance confidence, mask intersection-over-union confidence, and classification confidence. Step 443: Determine the number of iterations corresponding to the mask to be evaluated with the highest mask score as a preset number threshold.

[0101] In step 441 , during the object detection model training process, the progressive refinement unit may use iterations to perform multiple iterations of refinement on the sample initial mask, and use each refined mask obtained as a mask to be evaluated.

[0102] Furthermore, in step 442, for each mask to be evaluated, the mask to be evaluated can be evaluated from the perspectives of distance, IoU and classification confidence, and multiple indicators are integrated, that is, the mask distance confidence, the mask IoU confidence and the classification confidence are integrated to obtain the mask score corresponding to the mask to be evaluated.

[0103] Among them, the mask distance confidence can reflect the distance between the mask to be evaluated and the actual annotated mask (i.e., label), the mask intersection ratio confidence can reflect the overlapping area between the mask to be evaluated and the actual annotated mask, and the classification confidence can reflect the probability that a point in the mask to be evaluated falls in the actual annotated mask.

[0104] For example, the mask distance confidence can be calculated by the formula: ; ; Where, is the mask to be evaluated, To actually mark the mask, express arrive The Euclidean distance of express arrive The Euclidean distance of express and The bidirectional Hausdorff distance between is the mask distance confidence.

[0105] The confidence of the mask intersection-over-union ratio can be calculated by the formula: ; Where, is the confidence of the mask intersection-union ratio.

[0106] After obtaining the mask distance confidence, mask intersection-over-union confidence, and classification confidence, the mask distance confidence, mask intersection-over-union confidence, and classification confidence can be weighted by pre-assigned weight coefficients to obtain the mask score. As shown in the following formula: ; Where, Score the mask, 、 、 are weight coefficients, The classification confidence can be determined by the number of points in the mask to be evaluated that fall within the actual labeled mask and the total number of points in the mask to be evaluated. 、 、 Set to 32, 2, 2 respectively.

[0107] After obtaining the mask scores of the masks to be evaluated, further in step 443 , the number of iterations corresponding to the mask to be evaluated with the highest mask score is determined as a preset number threshold.

[0108] Through the above steps 441 to 443, each mask to be evaluated is evaluated based on the mask distance confidence, the mask intersection-over-union confidence, and the classification confidence. Considering that the mask distance confidence can measure the boundary continuity of the mask to be evaluated, the mask intersection-over-union confidence and the classification confidence can measure the accuracy of the mask to be evaluated, by fusing multiple indicators, each mask to be evaluated can be accurately evaluated, thereby improving the accuracy of the preset number threshold, and thus improving the accuracy of target detection.

[0109] S450 : Determine the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module to obtain a target bounding box.

[0110] The autonomous driving target detection method provided in the embodiment of the present application can dynamically suppress redundant channels through channel pruning, reduce the number of model parameters, and introduce a learnable offset through sub-pixel offset and bilinear interpolation to improve the accuracy of positioning deformed targets. In addition, by iteratively refining the initial mask to generate a target mask, the boundary can be gradually corrected to avoid over-smoothing.

[0111] Example 5 Based on the above embodiments, the post-processing module can optimize the target mask before determining the minimum rotated rectangle to further refine the target outline in the target mask and improve the accuracy of target detection. The method includes the following steps: S510: Obtain an image to be detected, and input the image to be detected into a pre-trained object detection model.

[0112] Among them, the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module and a post-processing module; the post-processing module includes a mask optimization unit.

[0113] S520 , extracting multi-scale semantic features from the image to be detected based on the feature pyramid module, and determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features.

[0114] S530 : Determine a target mask based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features.

[0115] S540 , determining the mask area of ​​the target mask based on the mask optimization unit, and determining the size of the structural element used for expansion and corrosion through the mask area to obtain the structural element, and performing corrosion and expansion processing on the target mask based on the mask optimization unit and the structural element.

[0116] Considering that there may be noise in the target mask, before generating the minimum rotated rectangle, the target mask may be corroded and expanded to eliminate the noise in the target mask.

[0117] Specifically, the mask optimization unit may determine the mask area of ​​the target mask, and then determine the size of the structural element used for dilation corrosion according to the mask area, as shown in the following formula: ; Where, is the size of the structural element used for expansion corrosion, which can be The circular structural element, is the mask area of ​​the target mask. When you can lock =1 to protect small structures.

[0118] After obtaining the size of the structural element used for expansion and corrosion, the structural element can be constructed according to the size, and then the mask optimization unit can use the structural element to erode and expand the target mask, as shown in the following formula: ; Where, is the target mask, represents the corrosion operation, represents the expansion operation, for The circular structural element, It is the target mask after corrosion and expansion processing.

[0119] S550 : Determine the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module to obtain a target bounding box.

[0120] In a specific embodiment, the post-processing module includes a bounding box generation unit, which determines the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module, including: The bounding box generation unit determines a rotation angle of the mask outline in the target mask, rotates the target mask by the rotation angle, and generates a minimum rotation rectangle containing the rotated mask outline.

[0121] Specifically, the bounding box generation unit may determine a rotation angle of the mask outline, where the rotation angle may be an angle of a central axis of the mask outline relative to a vertical reference axis. Furthermore, the target mask may be rotated according to the rotation angle so that the central axis of the mask outline in the target mask is parallel to the vertical reference axis.

[0122] After rotating the target mask, a minimum rectangle that encloses the mask outline can be generated and used as the minimum rotated rectangle, which is the target bounding box.

[0123] This method can effectively reduce the background area in the bounding box. Compared with the axis-aligned bounding box, it can further improve the accuracy of target detection, thereby reducing the redundant data in the target bounding box and improving the coverage accuracy by 10%, thereby improving the processing efficiency of the subsequent autonomous driving function on the target bounding box.

[0124] The autonomous driving target detection method provided by the present application embodiment eliminates noise by performing erosion and dilation on the target mask, thereby further improving the accuracy of the target bounding box. Furthermore, by dynamically adjusting the structural elements based on the size of the target mask, it is possible to balance noise removal with detail preservation.

[0125] Figure 5 This is a schematic diagram of a process of generating a target bounding box by a target detection model provided in Example 5 of the present application. Figure 5 As shown in the figure, the image to be detected can enter the feature pyramid module to obtain multi-scale semantic features, and then the semantic features and the image to be detected enter the boundary extraction module together to obtain multi-scale edge enhancement features. The multi-scale semantic features and edge enhancement features enter the mask refinement module together to obtain the target mask. After passing through the post-processing module, the target bounding box is obtained.

[0126] Figure 5 The method presented here demonstrates the following technical benefits: 1. Accuracy: In relevant datasets such as GTOT and RGBT234, this method significantly outperforms traditional methods in key metrics such as object boundary integrity and small object recall. It also reduces mask prediction error for partially occluded objects by over 10%, effectively improving the safety of autonomous driving systems. 2. Efficiency: This model, with its minimal parameter design, maintains high frame rates even on resource-constrained in-vehicle platforms, meeting real-time requirements. The model achieved a frame rate exceeding 45 FPS in tests on an RTX3080 graphics card, and is expected to reduce computing resource usage by over 15% in real-world environments. 3. Robustness: In extreme weather (rain, fog, and nighttime) tests on the RGBT234 / RGBT210 datasets, this method achieved a false positive rate reduction of over 10% compared to existing solutions, demonstrating strong environmental adaptability. Furthermore, its dynamic scoring mechanism effectively suppresses background interference, improving detection stability in complex scenarios.

[0127] Example 6 Figure 6 This is a structural diagram of an electronic device provided in Example 6 of this application. Figure 6 As shown, the electronic device 400 includes one or more processors 401 and a memory 402 .

[0128] The processor 401 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.

[0129] Memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and processor 401 may execute the program instructions to implement the autonomous driving target detection method of any embodiment of the present application described above and / or other desired functions. Various contents such as initial extrinsic parameters and thresholds may also be stored in the computer-readable storage medium.

[0130] In one example, electronic device 400 may further include an input device 403 and an output device 404, which are interconnected via a bus system and / or other connection mechanisms (not shown). Input device 403 may include, for example, a keyboard, a mouse, etc. Output device 404 may output various information to the outside, including warning information, braking force, etc. Output device 404 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.

[0131] Of course, to simplify, Figure 4 Only some of the components related to the present application in the electronic device 400 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 400 may further include any other appropriate components according to specific application scenarios.

[0132] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the autonomous driving target detection method provided by any embodiment of the present application.

[0133] The computer program product may be written in any combination of one or more programming languages ​​to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0134] In addition, an embodiment of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps of the autonomous driving target detection method provided by any embodiment of the present application.

[0135] The computer-readable storage medium may be any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0136] It should be noted that the terms used in this application are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates an exception, the words "one", "an", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.

[0137] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0138] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of this application, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.

Claims

1. A method for detecting an object in an autonomous driving system, characterized in that: include: Acquire an image to be detected and input the image to be detected into a pre-trained target detection model, wherein the target detection model includes a feature pyramid module, a boundary extraction module, a mask refinement module, and a post-processing module; Extracting multi-scale semantic features from the image to be detected based on the feature pyramid module, and determining multi-scale edge enhancement features based on the boundary extraction module, the image to be detected, and the multi-scale semantic features; Determining a target mask based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features; A minimum rotated rectangle corresponding to the mask outline in the target mask is determined based on the post-processing module to obtain a target bounding box.

2. The automatic driving target detection method according to claim 1, characterized in that: The feature pyramid module includes a bottom-up path and a top-down path, and extracting multi-scale semantic features in the image to be detected based on the feature pyramid module includes: Downsampling the image to be detected based on the bottom-up path to obtain a multi-scale first feature; Upsampling the multi-scale first features based on the top-down path to obtain multi-scale second features; Based on the top-down path, the first features and the second features of each scale are horizontally connected, and convolution processing is performed on the multi-scale features after the horizontal connection to obtain multi-scale semantic features.

3. The automatic driving target detection method according to claim 1, characterized in that: The boundary extraction module includes a basic feature extraction unit, an edge detection unit, an edge enhancement unit, and an alignment unit. Based on the boundary extraction module, the image to be detected, and the multi-scale semantic features, the multi-scale edge enhancement feature is determined, including: Extracting multi-scale basic features from the image to be detected based on the basic feature extraction unit; Extracting initial gradient features from the multi-scale basic features based on the edge detection unit; Performing enhancement processing on the initial gradient feature based on the edge enhancement unit to obtain a final gradient feature; The final gradient feature and the multi-scale semantic feature are fused based on the alignment unit to obtain a multi-scale edge enhancement feature.

4. The automatic driving target detection method according to claim 3, characterized in that: The initial gradient feature is enhanced based on the edge enhancement unit to obtain a final gradient feature, including: Performing a dilated convolution on the initial gradient features based on the edge enhancement unit to obtain multi-scale intermediate gradient features; The multi-scale intermediate gradient features are spliced ​​and compressed based on the edge enhancement unit to obtain final gradient features.

5. The automatic driving target detection method according to claim 3, characterized in that: The final gradient feature and the multi-scale semantic feature are fused based on the alignment unit to obtain a multi-scale edge enhancement feature, including: Based on the alignment unit, the final gradient features and the multi-scale semantic features are spliced ​​and normalized to obtain weights corresponding to the semantic features of each scale, and the weights corresponding to the semantic features of each scale are used to perform weighted processing on the final gradient features and the semantic features of each scale to obtain multi-scale edge enhancement features.

6. The automatic driving target detection method according to claim 1, characterized in that: The mask refinement module includes a feature fusion unit, a compensation unit, and a progressive refinement unit. The target mask is determined based on the mask refinement module, the multi-scale semantic features, and the multi-scale edge enhancement features, including: splicing the multi-scale semantic features and the multi-scale edge enhancement features based on the feature fusion unit, and performing channel pruning processing on the spliced ​​features using attention weights to obtain target features; Performing bilinear interpolation processing on the target feature based on the compensation unit to obtain an initial mask; The initial mask is iteratively refined based on the progressive refinement unit until the number of iterations reaches a preset threshold, thereby obtaining a target mask.

7. The automatic driving target detection method according to claim 6, characterized in that: The preset number threshold is determined by the following steps: During the training of the target detection model, the initial mask of the sample is iteratively refined multiple times based on the progressive refinement unit to obtain multiple masks to be evaluated; For each mask to be evaluated, determining a mask distance confidence, a mask intersection-over-union confidence, and a classification confidence between the mask to be evaluated and the actual annotated mask corresponding to the sample initial mask, and determining a mask score corresponding to the mask to be evaluated based on the mask distance confidence, the mask intersection-over-union confidence, and the classification confidence; The number of iterations corresponding to the mask to be evaluated with the highest mask score is determined as the preset number threshold.

8. The automatic driving target detection method according to claim 1, characterized in that: The post-processing module includes a mask optimization unit, and before determining the minimum rotated rectangle corresponding to the mask outline in the target mask based on the post-processing module, further includes: Determining a mask area of ​​the target mask based on the mask optimization unit, and determining a size of a structural element used for dilation corrosion according to the mask area to obtain the structural element; Based on the mask optimization unit and the structural element, the target mask is eroded and expanded.

9. The automatic driving target detection method according to claim 1, characterized in that: The post-processing module includes a bounding box generation unit, which determines a minimum rotated rectangle corresponding to a mask outline in the target mask based on the post-processing module, including: The bounding box generation unit determines a rotation angle of the mask outline in the target mask, and rotates the target mask according to the rotation angle to generate a minimum rotated rectangle containing the rotated mask outline.

10. An electronic device, characterized in that: The electronic device comprises: processor and memory; The processor is used to execute the steps of the autonomous driving target detection method as described in any one of claims 1 to 9 by calling the program or instructions stored in the memory.