Target detection method, device and equipment, medium and vehicle

By introducing feature extraction network, spatial attention network and feature fusion network into the object detection model, the detection ability of small targets is enhanced, and the problem of low efficiency of small target detection in the existing technology is solved, and efficient object detection is achieved.

CN120339975APending Publication Date: 2025-07-18BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410077996.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing target detection model has low attention and high computational volume when detecting small targets, resulting in inefficiency.

Method used

The first feature extraction network, the first spatial attention network and the first feature fusion network in the object detection model are used to enhance the recognition ability of the target object, and feature enhancement and fusion are performed for each pixel point in the feature map to improve the detection accuracy and recall rate of small objects.

Benefits of technology

The calculation amount of the target detection model is reduced, the detection efficiency is improved, and the detection accuracy and recall rate of small target objects are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339975A_ABST
    Figure CN120339975A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection method, device and equipment, a medium and a vehicle, and relates to the technical field of image recognition. The target object can be detected by setting the feature extraction network, the space attention network and the feature fusion network in the target detection model, the structure is simple, the calculation amount when the target detection model carries out target detection is greatly reduced, the target detection efficiency is improved, and the user experience is improved. Moreover, the spatial attention network can enable the target detection model to pay attention to the features of most of the objects in the image, and can improve the detection accuracy and recall rate of small target objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image recognition technology, and in particular, to an object detection method, apparatus, device, medium, and vehicle. Background Art

[0002] With the rapid development of deep learning, object detection technology has been widely applied in many fields. Object detection technology can identify objects in images and has the advantages of high real-time performance and high accuracy. For example, object detection technology can be applied to the recognition of objects by vehicles in the field of assisted driving to improve the safety of assisted driving.

[0003] Object detection technology can be implemented through an object detection model. For example, the YOLO (you only look once) model is an efficient object detection model. For example, the YOLOv8 model can identify at least the category and location of objects in an image by browsing the image only once.

[0004] Currently, when an object detection model detects an object in an image, the object detection model can basically accurately detect objects with larger sizes, but pays less attention to small objects with smaller sizes. In the current related technologies, global average pooling and global maximum pooling strategies are used to reduce image information loss, so as to increase the detection of small objects with smaller sizes by the object detection model. However, the existing object detection model has a complex structure and a large amount of computation during object detection, resulting in low efficiency of object detection. Summary of the Invention

[0005] To solve the above technical problems, the present disclosure provides an object detection method, apparatus, device, medium, and vehicle.

[0006] The first aspect of the present disclosure provides an object detection method, including:

[0007] Performing feature extraction on a to-be-detected image through a first feature extraction network in a preset object detection model to obtain a feature map of the to-be-detected image. The to-be-detected image includes a target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model;

[0008] For each pixel point in the feature map, based on the first spatial attention network in the object detection model, at least one pixel point within the preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point, and the features of each pixel point are enhanced by the target pixel points corresponding to each pixel point to obtain the target feature map corresponding to the feature map;

[0009] Based on the fusion result of feature fusion of the target feature map by the first feature fusion network, object recognition is performed to obtain the object detection result of the target object in the image to be detected.

[0010] The second aspect of the present disclosure provides an object detection device, including:

[0011] An extraction module, configured to extract features of an image to be detected through a first feature extraction network in a preset object detection model to obtain a feature map of the image to be detected. The image to be detected includes a target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model;

[0012] A first enhancement module, configured to, for each pixel point in the feature map, based on the first spatial attention network in the object detection model, determine at least one pixel point within the preset range area corresponding to the pixel point as the target pixel point corresponding to the pixel point, and enhance the features of each pixel point by the target pixel points corresponding to each pixel point to obtain the target feature map corresponding to the feature map;

[0013] A recognition module, configured to perform object recognition based on the fusion result of feature fusion of the target feature map by the first feature fusion network to obtain the object detection result of the target object in the image to be detected.

[0014] The third aspect of the present disclosure provides a computer device, including a memory and a processor. Among them, a computer program is stored in the memory, and when the computer program is executed by the processor, the object detection method in the first aspect can be implemented.

[0015] The fourth aspect of the present disclosure provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the object detection method in the first aspect can be implemented.

[0016] The fifth aspect of the present disclosure provides a vehicle, which includes the object detection device in the second aspect and / or the computer device in the third aspect and / or the computer-readable storage medium in the fourth aspect, and can implement the object detection method in the first aspect.

[0017] The technical solution provided by the present disclosure has the following advantages compared with the prior art:

[0018] The present disclosure extracts features from the image to be detected through the first feature extraction network in the preset target detection model, and obtains the feature map of the image to be detected. The image to be detected includes the target object to be recognized. The target detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the recognition ability of the target detection model for the target object. For each pixel point in the feature map, based on the first spatial attention network in the target detection model, at least one pixel point within the preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point, and the feature of each pixel point is enhanced through the target pixel point corresponding to each pixel point, and the target feature map corresponding to the feature map is obtained. Based on the fusion result of feature fusion of the target feature map by the first feature fusion network, target recognition is performed to obtain the target detection result of the target object in the image to be detected. It only needs to set a feature extraction network, a spatial attention network, and a feature fusion network in the target detection model to realize the detection of the target object. The structure is simple, which greatly reduces the calculation amount when the target detection model performs target detection, improves the efficiency of target detection, and the spatial attention network can enable the target detection model to focus on the features of most objects in the image, which can improve the accuracy and recall rate of detecting small target objects. Description of the Drawings

[0019] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained according to these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of a target detection method provided by an embodiment of the present disclosure;

[0022] Figure 2 It is a flowchart of another target detection method provided by an embodiment of the present disclosure;

[0023] Figure 3 It is a flowchart of yet another target detection method provided by an embodiment of the present disclosure;

[0024] Figure 4It is a flowchart of performing object detection by an object detection model provided by an embodiment of the present disclosure;

[0025] Figure 5 It is a flowchart of a method for training an object detection model provided by an embodiment of the present disclosure;

[0026] Figure 6 It is a schematic structural diagram of an object detection device provided by an embodiment of the present disclosure;

[0027] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0028] In order to more clearly understand the above-mentioned objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0029] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0030] It should be understood that the various steps recorded in the method implementation manners of the present disclosure may be executed in different orders and / or executed in parallel. In addition, the method implementation manners may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0031] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0033] The object detection method provided by the embodiments of the present disclosure can be executed by a computer device, which can be understood as any device with processing and computing capabilities. The device can include, but is not limited to, mobile terminals such as smart phones, laptop computers, tablet computers (PADs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed electronic devices such as digital TVs, desktop computers, and smart home devices.

[0034] To better understand the inventive concept of the embodiments of the present disclosure, the technical solutions of the embodiments of the present disclosure will be described below in conjunction with exemplary embodiments.

[0035] Figure 1 It is a flowchart of an object detection method provided by the embodiments of the present disclosure. As Figure 1 shown, the object detection method provided in this embodiment includes the following steps:

[0036] Step 110: Extract features from the to-be-detected image through the first feature extraction network in the preset object detection model to obtain a feature map of the to-be-detected image. The to-be-detected image includes the target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object detection model's ability to recognize the target object.

[0037] In the embodiments of the present disclosure, the computer device can acquire the to-be-detected image. The to-be-detected image may include the target object to be recognized. For example, the to-be-detected image can be acquired through a camera or a webcam.

[0038] A preset object detection model is set in the computer device. The object detection model can at least include a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network can be used to enhance the object detection model's ability to recognize the target object.

[0039] The computer device can input the to-be-detected image into the preset object detection model, and extract features from the to-be-detected image through the first feature extraction network in the object detection model to obtain a feature map of the to-be-detected image.

[0040] Step 120: For each pixel point in the feature map, based on the first spatial attention network in the object detection model, determine at least one pixel point within the preset range area corresponding to the pixel point as the target pixel point corresponding to the pixel point, and enhance the feature of each pixel point through the target pixel point corresponding to each pixel point to obtain the target feature map corresponding to the feature map.

[0041] In the embodiments of the present disclosure, after obtaining the feature map of the image to be measured, based on the first spatial attention network in the object detection model, for each pixel point in the feature map, at least one pixel point within the preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point, and the features of each pixel point are enhanced by the target pixel points corresponding to each pixel point, so as to obtain the target feature map corresponding to the feature map.

[0042] Among them, the spatial attention network can be understood as a model constructed based on the spatial attention mechanism, which can transform various deformation data in space and automatically capture the features of important regions, and can ensure that the same result as the original image before the operation can still be obtained after operations such as cropping, translation, or rotation of the image. In some embodiments, the spatial attention network can adopt the spatial attention network structure of F-SSD.

[0043] The preset range area can be set as needed. For example, it can be a pixel area composed of all adjacent pixel points of the pixel point, and no specific limitation is made here.

[0044] In some embodiments, to enhance the features of each pixel point by the target pixel points corresponding to each pixel point, for each pixel point in the feature map, based on the preset weight value corresponding to the feature map, the features of the target pixel points corresponding to the pixel point are weighted and summed to enhance the features of the pixel point. Thus, the features of the pixel point can include the features of at least one target pixel point within the preset range area of the pixel point.

[0045] Among them, each feature map has a corresponding preset weight value, and the preset weight value can represent the weight value of each pixel point in the feature map. The preset weight value corresponding to each feature map can be set as needed, and no limitation is made here.

[0046] By enhancing the features of each pixel point in the feature map of the image to be measured through the spatial attention network, the features of small targets in the feature map can be strengthened, the probability that the features of small targets are filtered can be reduced, so as to retain more features of small targets, improve the attention of the object detection model to small targets, and enable it to better learn and understand the important features of small targets in the image. A small target can be understood as a target object in the image to be measured with a size smaller than the preset size, and the preset size can be set as needed, and no limitation is made here.

[0047] Step 130: Perform object recognition based on the fusion result of feature fusion of the target feature map to obtain the object detection result of the target object in the image to be measured.

[0048] In an embodiment of the present disclosure, after obtaining the target feature map corresponding to the feature map of the image to be detected, the computer device may perform feature fusion on the target feature map based on the first feature fusion network in the object detection model to obtain a fusion result, and then perform object recognition based on the fusion result to obtain the object detection result of the target object in the image to be detected.

[0049] The object detection result of the image to be detected may include at least one of the category and position information of each target object in the image to be detected.

[0050] In an embodiment of the present disclosure, the first feature extraction network in the preset object detection model is used to extract features from the image to be detected to obtain the feature map of the image to be detected. The image to be detected includes the target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model; for each pixel point in the feature map, based on the first spatial attention network in the object detection model, at least one pixel point within the preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point, and the features of each pixel point are enhanced by the target pixel point corresponding to each pixel point to obtain the target feature map corresponding to the feature map; based on the fusion result of feature fusion of the target feature map by the first feature fusion network, object recognition is performed to obtain the object detection result of the target object in the image to be detected. Only by setting a feature extraction network, a spatial attention network, and a feature fusion network in the object detection model can the detection of the target object be realized. The structure is simple, which greatly reduces the computational amount when the object detection model performs object detection, improves the efficiency of object detection, and the spatial attention network can enable the object detection model to pay attention to the features of most objects in the image, which can improve the accuracy and recall rate of detecting small target objects.

[0051] Figure 2 is a flowchart of an object detection method provided by an embodiment of the present disclosure. As Figure 2 shown, the object detection method provided in this embodiment includes the following steps:

[0052] Step 210: Use the first feature extraction network in the preset object detection model to extract features from the image to be detected to obtain the feature map of the image to be detected. The image to be detected includes the target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model.

[0053] Step 220: For each pixel point in the feature map, based on the first spatial attention network in the object detection model, at least one pixel point within the preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point, and the features of each pixel point are enhanced through the target pixel points corresponding to each pixel point, so as to obtain the target feature map corresponding to the feature map.

[0054] Step 230: Based on the first feature fusion network in the object detection model, feature fusion processing is performed on the target feature map and other feature maps output by other feature extraction networks except the first feature extraction network, so as to obtain the corresponding fusion feature map.

[0055] In the embodiments of the present disclosure, after obtaining the target feature map corresponding to the feature map of the image to be detected, the computer device may perform feature fusion processing on the target feature map and other feature maps output by the feature extraction network based on the first feature fusion network in the object detection model, so as to obtain the corresponding fusion feature map.

[0056] For example, the feature fusion network may be a Feature Pyramid Network (FPN) and / or a Path Aggregation Network (PAN).

[0057] In some embodiments, in the above object detection model, a second feature extraction network may further be included between the first spatial attention network and the feature fusion network. The above-mentioned performing feature fusion processing on the target feature map and other feature maps output by other feature extraction networks except the first feature extraction network based on the feature fusion network in the object detection model to obtain the corresponding fusion feature map may include steps 2301-2302:

[0058] Step 2301: Based on the second feature extraction network in the object detection model, perform upsampling processing and / or downsampling processing on the target feature map to obtain the upsampled image and / or downsampled image corresponding to the target feature map.

[0059] In some embodiments, the computer device may perform upsampling processing on the target feature map to obtain the upsampled image corresponding to the target feature map.

[0060] In some embodiments, the computer device may perform downsampling processing on the target feature map to obtain the downsampled image corresponding to the target feature map.

[0061] In some embodiments, perform upsampling processing and downsampling processing on the target feature map to obtain the upsampled image and downsampled image corresponding to the target feature map.

[0062] Among them, upsampling can increase the resolution of an image, thereby enlarging the image and improving its resolution. Upsampling can be performed by interpolation or filling with new data.

[0063] Downsampling can reduce the resolution of an image, thereby shrinking the image and lowering its resolution. Downsampling can be performed by skipping or merging data points.

[0064] Step 2302: Based on the first feature fusion network, splice the upsampled image and / or downsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks outside the first feature extraction network to obtain the corresponding fused feature map.

[0065] In some embodiments, the computer device can splice the upsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks outside the first feature extraction network to obtain the corresponding fused feature map.

[0066] In some embodiments, the computer device can splice the downsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks outside the first feature extraction network to obtain the corresponding fused feature map.

[0067] In some embodiments, the computer device can splice the upsampled image and downsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks outside the first feature extraction network to obtain the corresponding fused feature map.

[0068] Step 240: Based on the target recognition network in the target detection model, perform target recognition on the fused feature map to obtain the target detection result of the target object in the image to be measured.

[0069] In the embodiments of the present disclosure, after obtaining the fused feature map of the image to be measured, the computer device can perform target recognition on the fused feature map based on the target recognition network in the preset target detection model to obtain the target detection result of the target object in the image to be measured.

[0070] In some embodiments, performing target recognition on the fused feature map to obtain the target detection result of the target object in the image to be measured may include steps 2401-2405:

[0071] Step 2401: Based on the feature differences of each pixel point in the fused feature map, generate multiple detection frames corresponding to each target object included in the fused feature map.

[0072] In the embodiments of the present disclosure, after obtaining the fused feature map of the image to be measured, the object recognition network in the preset object detection model may identify at least one target object in the fused feature map based on the feature differences of each pixel point in the fused feature map, and then generate multiple detection frames corresponding to each target object included in the fused feature map. A detection frame may also be called a bounding box, that is, each target object may have one or more detection frames.

[0073] Step 2402: For each target object in the fused feature map, determine the contour of the target object based on the features of the edge pixels of the target object in the fused feature map.

[0074] In the embodiments of the present disclosure, the object recognition network in the object detection model may determine the contour of each target object in the fused feature map based on the features of the edge pixels of the target object in the fused feature map.

[0075] Step 2403: For each detection frame corresponding to the target object, determine the degree of conformity between the detection frame and the contour of the target object as the confidence level of the detection frame.

[0076] In the embodiments of the present disclosure, the object recognition network in the object detection model may determine the degree of conformity between each detection frame corresponding to the target object and the contour of the target object as the confidence level of the detection frame.

[0077] The confidence level of the detection frame of the target object can be understood as the degree of conformity between the detection frame and the contour of the target object. The closer the detection frame is to the contour of the target object, the greater the confidence level of the detection frame.

[0078] Step 2404: Among the confidence levels of the multiple detection frames corresponding to the target object, determine the detection frame corresponding to the maximum confidence level as the target detection frame of the target object.

[0079] In the embodiments of the present disclosure, since the multiple detection frames of the target object generated by the object recognition network often overlap. Therefore, it is necessary to select a detection frame that is closest to the contour of the target object from at least one detection frame of the target object as the target detection frame of the target object, and remove the redundant detection frames.

[0080] After obtaining the multiple detection frames of each target object and the confidence level of each detection frame, the object recognition network in the object detection model may determine the detection frame corresponding to the maximum confidence level among the confidence levels of the multiple detection frames corresponding to the target object as the target detection frame of the target object, and delete the detection frames other than the target detection frame among the multiple detection frames corresponding to the target object. The target detection frame is the detection frame with the highest degree of conformity to the contour of the target object.

[0081] In some embodiments, step 2404 can be performed by a Non-maximum supression (NMS) algorithm to remove redundant detection boxes of the target object. The principle is to eliminate the detection box results that overlap significantly with the maximum value. For specific details, reference can be made to related technologies and will not be elaborated here.

[0082] Step 2405: Identify the target object in the target detection box to obtain the target detection result of the target object.

[0083] In the embodiments of the present disclosure, after obtaining the target detection box of the target object, the target recognition network in the target detection model can identify the target object in the target detection box to obtain the target detection result of the target object.

[0084] Thus, the features of each pixel point in the feature map can be enhanced, enabling the target detection model to focus on the features of most objects in the image, reducing the missed detection rate and false detection rate of small target objects in the image, improving the accuracy and recall rate of detecting small target objects. It is only necessary to set a feature extraction network, a spatial attention network, and a feature fusion network in the target detection model to achieve the detection of the target object. The structure is simple, greatly reducing the computational complexity when the target detection model performs target detection and improving the efficiency of target detection.

[0085] Figure 3 is a flowchart of a target detection method provided by the embodiments of the present disclosure. As Figure 3 shown, the target detection method provided in this embodiment includes the following steps:

[0086] Step 310: Extract features from the to-be-detected image through a first feature extraction network in a preset target detection model to obtain a feature map of the to-be-detected image. The to-be-detected image includes the target object to be recognized. The target detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the target recognition ability of the target detection model.

[0087] Step 320: For each pixel point in the feature map, based on the first spatial attention network in the target detection model, determine at least one pixel point within the preset range area corresponding to the pixel point as the target pixel point corresponding to the pixel point, and enhance the feature of each pixel point through the target pixel point corresponding to each pixel point to obtain the target feature map corresponding to the feature map.

[0088] Step 330: Based on the second spatial attention network in the target detection model, perform feature enhancement on each pixel point in the target feature map to obtain the target feature map after feature enhancement.

[0089] In the embodiments of the present disclosure, a second spatial attention network may further be included between the second feature extraction network and the first feature fusion network in the target detection model.

[0090] After obtaining the target feature map corresponding to the feature map, the computer device may, based on the second spatial attention network in the target detection model, perform feature enhancement on each pixel point in the target feature map, that is, determine at least one pixel point within a preset range area corresponding to each pixel point in the target feature map as the target pixel point corresponding to the pixel point, and enhance the feature of each pixel point through the target pixel point corresponding to each pixel point to obtain the target feature map after feature enhancement.

[0091] Step 340: Determine the target feature map after feature enhancement as the target feature map for feature fusion by the first feature fusion network.

[0092] Step 350: Based on the first feature fusion network in the target detection model, perform feature fusion processing on the target feature map after feature enhancement and other feature maps output by other feature extraction networks except the first feature extraction network to obtain the corresponding fusion feature map.

[0093] Step 360: Based on the target recognition network in the target detection model, perform target recognition on the fusion feature map to obtain the target detection result of the target object in the image to be detected.

[0094] Therefore, secondary feature enhancement can be performed on each pixel point in the feature map of the image to be detected, so that the target detection model can focus on the features of most objects in the image, reduce the missed detection rate and false detection rate of small target objects in the image, improve the accuracy and recall rate of detecting small target objects, and only need to set a feature extraction network, a spatial attention network, and a feature fusion network in the target detection model to achieve the detection of the target object, with a simple structure, greatly reducing the calculation amount during target detection by the target detection model and improving the efficiency of target detection.

[0095] For example, Figure 4 is a flowchart of a target detection model provided by the embodiments of the present disclosure for target detection, as Figure 4As shown, the image to be tested is input into the target detection model, and along the direction of the arrow, the CBS convolution layer and the C2f module in the first feature extraction network are passed to extract features of the image to be tested, and the feature map A of the image to be tested is output at 401, and then the features of each pixel in the feature map A are enhanced through the first spatial attention network to obtain the target feature map B, and the target feature map B is upsampled based on the second feature extraction network at 402 to obtain the upsampled image B1 corresponding to the target feature map B, and the features of each pixel in the upsampled image B1 are enhanced through the second spatial attention network to obtain the target feature map C, and the target feature map C is upsampled based on the third feature extraction network at 403 to obtain the upsampled image C1 corresponding to the target feature map C, and the feature map A and the upsampled image C1 are feature fused based on the first feature fusion network 404 to obtain the fused feature map R1, and the fused feature map R1 is recognized based on the target recognition network, and the target detection result G1 of the image to be tested is obtained at 405;

[0096] like Figure 4 As shown, the location where the spatial attention network is added in the embodiment of the present application is between the feature extraction network and the feature fusion network. Referring to the setting rules of the first spatial attention network and the second spatial attention network, the third spatial attention network and the fourth spatial attention network can be further set in the structure of the target detection model. After obtaining the fused feature map R1 corresponding to the image to be tested, the fused feature map R1 can be further enhanced through the third spatial attention network to enhance the features of each pixel in the fused feature map R1 to obtain the target feature map D of the image to be tested, and the upsampled image B1 and the target feature map D are fused based on the second feature fusion network 406 to obtain the fused feature map R2, and the fused feature map R2 is subjected to target recognition based on the target recognition network, and the target detection result G2 of the image to be tested is obtained at 407;

[0097] The fused feature map R2 is passed through the fourth spatial attention network to enhance the features of each pixel in the fused feature map R2 to obtain the target feature map E of the image to be tested. The feature map A and the target feature map E are fused based on the third feature fusion network 408 to obtain the fused feature map R3. The fused feature map R3 is used for target recognition based on the target recognition network to obtain the target detection result G3 of the image to be tested at 409.

[0098] Figure 4Among them, CBS represents the CBS convolutional layer, S represents the stride of the CBS convolutional layer, and K represents the number of convolutional kernels of the CBS convolutional layer; C2f represents the C2f module, which is used to convert the output of the CBS convolutional layer into the input of the fully connected layer; SPPF represents the spatial pyramid pooling; Upsample represents the upsampling process, Concat represents the feature fusion, and Detect represents the loss term parameters of the output, which can be used as the object detection result.

[0099] Figure 5 It is a flowchart of a method for training an object detection model provided by an embodiment of the present disclosure. As Figure 5 shown, the object detection model training method provided in this embodiment includes the following steps:

[0100] Step 510: Obtain a preset number of training images, where each training image contains multiple objects to be recognized and the labels of each object to be recognized.

[0101] In the embodiment of the present disclosure, the computer device can obtain a preset number of training images, and each training image contains multiple objects to be recognized and the labels of each object to be recognized.

[0102] The preset number can be set as needed and is not limited here.

[0103] Specifically, after obtaining the preset number of training images, the user can perform labeling operations on the objects to be recognized in each training image, and the computer device can generate the labels of the respective objects to be recognized included in the training image in response to the labeling operations on the objects to be recognized in the training image. The labels can at least include the labels indicating the existence of the objects to be recognized, and can also include the category labels of the objects to be recognized, etc.

[0104] Step 520: Input the preset number of training images into the object detection model, and perform object detection on the training images based on the object detection model to obtain the object detection results of each training image.

[0105] Step 530: For the object detection result of each target object, calculate the gap between the object detection result and the label of the target object based on a preset loss function and the label of the target object. The preset loss function includes a target balance factor, and the target balance factor is used to characterize the loss weight indicating whether the target object exists. The target balance factor is greater than other factors outside the target balance factor in the preset loss function.

[0106] In the embodiment of the present disclosure, after obtaining the object detection results of each target object in each training image, the computer device can calculate the gap between the object detection result and the label of the target object based on a preset loss function and the label of the target object. The smaller the gap, the more accurate the object detection result.

[0107] Among them, the preset loss function includes a target balance factor, which is the loss weight of whether the target object exists. The target balance factor is greater than other factors other than the target balance factor in the preset loss function. Other factors may include the class loss weight of the target object. Thus, during training, the loss of whether the target object exists will increase, and the model will focus more on reducing the target balance factor during training, enabling the model to focus on detecting whether the target object exists and enabling the model to better learn and understand the features of each target object in the image.

[0108] Step 540: Based on the gap, adjust the parameters of the target detection model, and repeat the training of the target detection model until the gap is less than the preset threshold to obtain a trained target detection model.

[0109] In the embodiments of the present disclosure, the computer device can adjust the parameters of the target detection model based on the gap between the target detection result and the label of the target object, and repeat the training of the target detection model until the gap is less than the preset threshold to obtain a trained target detection model.

[0110] At this time, the trained target detection model will focus on detecting whether the target object in the image exists, rather than classifying the target object. Thus, the attention of the target detection model to all target objects can be improved, and further the probability of recognizing small targets can be increased, and the accuracy and recall rate of detecting small target objects can be improved.

[0111] In some embodiments of the present disclosure, the above target detection model may include a YOLO model, such as the YOLOv8 model. The YOLO (you only look once) model only needs to browse the image once to identify the class and location information of each target object in the image, and can output the recognition results of all detected target objects at one time.

[0112] Figure 6 It is a schematic structural diagram of a target detection device provided by an embodiment of the present disclosure. This device can be understood as the above computer device or some functional modules in the above computer device. As Figure 6 shown, the target detection device 600 includes:

[0113] An extraction module 610, configured to extract features from a to-be-detected image through a first feature extraction network in a preset object detection model, so as to obtain a feature map of the to-be-detected image. The to-be-detected image includes a target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model;

[0114] A first enhancement module 620, configured to, for each pixel point in the feature map, based on the first spatial attention network in the object detection model, determine at least one pixel point within a preset range area corresponding to the pixel point as the target pixel point corresponding to the pixel point, and enhance the feature of each pixel point through the target pixel point corresponding to each pixel point, so as to obtain a target feature map corresponding to the feature map;

[0115] A recognition module 630, configured to perform object recognition based on the fusion result of feature fusion of the target feature map by the first feature fusion network, so as to obtain a target detection result of the target object in the to-be-detected image.

[0116] Optionally, the above first enhancement module includes:

[0117] A summation sub-module, configured to, for each pixel point in the feature map, based on a preset weight value corresponding to the feature map, perform weighted summation on the features of the target pixel points corresponding to the pixel point, so as to enhance the feature of the pixel point.

[0118] Optionally, the above recognition module includes:

[0119] A fusion sub-module, configured to perform feature fusion processing on the target feature map and other feature maps output by other feature extraction networks except the first feature extraction network based on the first feature fusion network in the object detection model, so as to obtain a corresponding fusion feature map;

[0120] A recognition sub-module, configured to perform object recognition on the fusion feature map based on the object recognition network in the object detection model, so as to obtain a target detection result of the target object in the to-be-detected image.

[0121] Optionally, a second feature extraction network is further included between the first spatial attention network and the first feature fusion network;

[0122] The above fusion sub-module includes:

[0123] A sampling processing unit, configured to perform upsampling processing and / or downsampling processing on the target feature map based on the second feature extraction network in the object detection model, so as to obtain an upsampled image and / or a downsampled image corresponding to the target feature map;

[0124] The splicing unit is used to splice the upsampled image and / or downsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks except the first feature extraction network based on the first feature fusion network to obtain the corresponding fused feature map.

[0125] Optionally, a second spatial attention network is further included between the second feature extraction network and the first feature fusion network;

[0126] The above-mentioned fusion sub-module includes:

[0127] The feature enhancement unit is used to enhance the features of each pixel point in the target feature map based on the second spatial attention network in the target detection model to obtain the target feature map with enhanced features;

[0128] The first determination unit is used to determine the target feature map for feature fusion by the first feature fusion network as the target feature map with enhanced features.

[0129] Optionally, the above-mentioned recognition sub-module includes:

[0130] The second determination unit is used to determine the contour of the target object based on the features of the edge pixels of the target object in the fused feature map for each target object in the fused feature map;

[0131] The third determination unit is used to determine the degree of conformity between the detection frame and the contour of the target object as the confidence level of the detection frame for each detection frame corresponding to the target object;

[0132] The fourth determination unit is used to determine the detection frame corresponding to the maximum confidence level as the target detection frame of the target object among the confidence levels of multiple detection frames corresponding to the target object;

[0133] The recognition unit is used to recognize the target object in the target detection frame to obtain the target detection result of the target object.

[0134] Optionally, the above-mentioned target detection device includes:

[0135] The acquisition module is used to acquire a preset number of training images, and each training image contains multiple objects to be recognized and the labels of each object to be recognized;

[0136] The detection module is used to input a preset number of training images into the target detection model, and perform target detection on the training images based on the target detection model to obtain the target detection results of each training image;

[0137] A calculation module, configured to calculate, for the object detection result of each target object, the gap between the object detection result and the label of the target object based on a preset loss function and the label of the target object. The preset loss function includes a target balance factor, which is used to represent the loss weight of whether the target object exists. The target balance factor is greater than other factors in the preset loss function other than the target balance factor;

[0138] A training module, configured to adjust the parameters of the object detection model based on the gap, and repeatedly train the object detection model until the gap is less than a preset threshold, so as to obtain a trained object detection model.

[0139] The object detection device provided by the embodiments of the present disclosure can implement the method of any of the above embodiments, and its execution manner and beneficial effects are similar, which will not be elaborated here.

[0140] The embodiments of the present disclosure further provide a computer device, which includes a processor and a memory. Among them, a computer program is stored in the memory. When the computer program is executed by the processor, the method of any of the above embodiments can be implemented, and its execution manner and beneficial effects are similar, which will not be elaborated here.

[0141] Figure 7 is a schematic structural diagram of a computer device provided by the embodiments of the present disclosure. As Figure 7 shown, the computer device 700 may include a processor 710 and a memory 720. Among them, a computer program 721 is stored in the memory 720. When the computer program 721 is executed by the processor 710, the method provided by any of the above embodiments can be implemented, and its execution manner and beneficial effects are similar, which will not be elaborated here.

[0142] Of course, for simplicity, Figure 7 only some of the components related to the present invention in the computer device 700 are shown, and components such as buses, input / output interfaces, input devices, and output devices are omitted. In addition, according to specific application scenarios, the computer device 700 may further include any other appropriate components.

[0143] The embodiments of the present disclosure provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any of the above embodiments can be implemented, and its execution manner and beneficial effects are similar, which will not be elaborated here.

[0144] The above computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0145] The above computer program may be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer device, partially on the user device, executed as an independent software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device or server.

[0146] The embodiments of the present disclosure provide a vehicle. The vehicle includes the above target detection device and / or the above computer device and / or the above computer-readable storage medium, and can implement the method of any of the above embodiments. The execution manner and beneficial effects are similar and will not be elaborated here.

[0147] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0148] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0149] The foregoing is only a specific implementation manner of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather should be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target detection method, characterized in that, Including: Feature extraction is performed on the image to be detected through a first feature extraction network in a preset object detection model, obtaining a feature map of the image to be detected. The image to be detected includes a target object to be recognized. The object detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the object recognition ability of the object detection model. For each pixel point in the feature map, based on the first spatial attention network in the object detection model, at least one pixel point within a preset range area corresponding to the pixel point is determined as the target pixel point corresponding to the pixel point. The feature of each pixel point is enhanced through the target pixel point corresponding to each pixel point, obtaining a target feature map corresponding to the feature map. Based on the fusion result of feature fusion of the target feature map by the first feature fusion network, object recognition is performed to obtain the object detection result of the target object in the image to be detected.

2. The method according to claim 1, wherein The enhancing the feature of each pixel point through the target pixel point corresponding to each pixel point includes: For each pixel point in the feature map, based on the preset weight value corresponding to the feature map, the features of the target pixel points corresponding to the pixel point are weighted and summed to enhance the feature of the pixel point.

3. The method according to claim 1, wherein The performing object recognition based on the fusion result of feature fusion of the target feature map by the first feature fusion network to obtain the object detection result of the target object in the image to be detected includes: Based on the first feature fusion network in the object detection model, feature fusion processing is performed on the target feature map and other feature maps output by other feature extraction networks other than the first feature extraction network, obtaining a corresponding fusion feature map. Based on the object recognition network in the object detection model, object recognition is performed on the fusion feature map to obtain the object detection result of the target object in the image to be detected.

4. The method according to claim 3, wherein A second feature extraction network is further included between the first spatial attention network and the first feature fusion network. The performing feature fusion processing on the target feature map and other feature maps output by other feature extraction networks other than the first feature extraction network based on the first feature fusion network in the object detection model to obtain a corresponding fusion feature map includes: Based on the second feature extraction network in the object detection model, upsampling processing and / or downsampling processing is performed on the target feature map, obtaining an upsampled image and / or a downsampled image corresponding to the target feature map. Based on the first feature fusion network, the upsampled image and / or the downsampled image corresponding to the target feature map and other feature maps output by other feature extraction networks other than the first feature extraction network are spliced to obtain a corresponding fusion feature map.

5. The method according to claim 4, characterized in that A second spatial attention network is further included between the second feature extraction network and the first feature fusion network. Before performing feature fusion processing on the target feature map and other feature maps output by other feature extraction networks other than the first feature extraction network based on the first feature fusion network in the target detection model, the method further includes: Based on the second spatial attention network in the target detection model, perform feature enhancement on each pixel point in the target feature map to obtain a target feature map after feature enhancement; Determine the target feature map after feature enhancement as the target feature map for feature fusion by the first feature fusion network.

6. The method according to claim 3, wherein The performing target recognition on the fused feature map to obtain the target detection result of the target object in the to-be-detected image includes: Based on the feature differences of the pixel points in the fused feature map, generate a plurality of detection frames corresponding to each target object included in the fused feature map; For each target object in the fused feature map, based on the features of the edge pixels of the target object in the fused feature map, determine the contour of the target object; For each detection frame corresponding to the target object, determine the degree of conformity between the detection frame and the contour of the target object as the confidence level of the detection frame; Among the confidence levels of the multiple detection frames corresponding to the target object, determine the detection frame corresponding to the maximum confidence level as the target detection frame of the target object; Perform recognition on the target object in the target detection frame to obtain the target detection result of the target object.

7. The method according to any one of claims 1-6, characterized in that, The training method of the target detection model includes: Obtain a preset number of training images, each of the training images including a plurality of objects to be recognized and labels of each object to be recognized; Input the preset number of training images into the target detection model, and perform target detection on the training images based on the target detection model to obtain the target detection results of the training images; For the target detection result of each target object, calculate the gap between the target detection result and the label of the target object based on a preset loss function and the label of the target object. The preset loss function includes a target balance factor, and the target balance factor is used to represent the loss weight of whether the target object exists. The target balance factor is greater than other factors in the preset loss function other than the target balance factor; Based on the gap, adjust the parameters of the target detection model, and repeat training the target detection model until the gap is less than a preset threshold to obtain a trained target detection model.

8. A target detection device, characterized in that, Includes: An extraction module, configured to perform feature extraction on a to-be-detected image through a first feature extraction network in a preset target detection model to obtain a feature map of the to-be-detected image. The to-be-detected image includes a target object to be recognized. The target detection model at least includes a first feature extraction network, a first spatial attention network, and a first feature fusion network. The first spatial attention network is located between the first feature extraction network and the first feature fusion network, and the first spatial attention network is used to enhance the recognition ability of the target detection model for the target object; A first enhancement module, configured to, for each pixel point in the feature map, based on a first spatial attention network in the object detection model, determine at least one pixel point within a preset range area corresponding to the pixel point as the target pixel point corresponding to the pixel point, and enhance the feature of each pixel point through the target pixel point corresponding to each pixel point, to obtain a target feature map corresponding to the feature map; An identification module, configured to perform object identification based on a fusion result of feature fusion of the target feature map by the first feature fusion network, to obtain a target detection result of a target object in the to-be-detected image.

9. A computer device, characterized in that, Comprising: A memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the object detection method according to any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and when the computer program is executed by a processor, the object detection method according to any one of claims 1-7 is implemented.

11. A vehicle, characterized in that, The vehicle includes the object detection device according to claim 8 and / or the computer device according to claim 9 and / or the computer-readable storage medium according to claim 10, and implements the object detection method according to any one of claims 1-7.