Vehicle driving target detection method, electronic equipment and storage medium

By introducing channel application attention module, pyramid pooling module and feature fusion module into the object detection model, the missed detection and misdetection problems of small object detection are solved, and the safety and decision-making accuracy of autonomous driving are improved.

CN120279526APending Publication Date: 2025-07-08SUZHOU GAIDE PHOTOELECTRIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349878.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, vehicle target detection methods are difficult to effectively identify small targets with sizes less than a certain threshold, resulting in missed or mis-checked, affecting the safety of autonomous driving.

Method used

The object detection model is adopted, including the channel application attention module, pyramid pooling module and feature fusion module, and the detection accuracy of small targets is improved by determining the weight coefficient of the input features, multi-scale pooling and feature fusion.

Benefits of technology

It realizes accurate identification and positioning of small-sized targets, enhances the safety and decision-making capabilities of the autonomous driving system, and effectively warns of potential dangers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279526A_ABST
    Figure CN120279526A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle driving target detection method, electronic equipment and a storage medium. The method comprises the steps that a driving environment image of a target vehicle and a pre-trained target detection model are acquired, the target detection model is used for detecting a target object with the size smaller than a preset size in an input image, and the target detection model comprises a channel application attention module, a pyramid pooling module and a feature fusion module; the channel application attention module is used for determining weight coefficients of input features, the pyramid pooling module is used for pooling the input features on multiple scales, and the feature fusion module is used for fusing the input features; the driving environment image is input into the target detection model to obtain a target detection result corresponding to the driving environment image, the target detection result is at least used for representing whether the driving environment image comprises the target object, potential danger can be effectively warned, and automatic driving safety is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method for target detection of vehicle driving, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of driverless technology, target detection has become an important part of ensuring the safety and reliability of such systems. In the application scenarios of driverless driving, a vehicle must continuously perceive various objects in its surrounding environment, including but not limited to other vehicles, pedestrians, and traffic signs, and make corresponding decisions based on this.

[0003] In related technologies, target detection methods often have difficulty effectively identifying small targets whose sizes are smaller than a certain threshold. This is mainly because small targets occupy a small number of pixels in the image, resulting in less significant features, and are easily submerged by background noise, missed detection, or misjudgment. Especially in the scenario of autonomous driving, failure to timely detect and correctly classify small obstacles such as pedestrians and bicycles may lead to serious safety problems. Summary of the Invention

[0004] The present invention provides a method for target detection of vehicle driving, an electronic device, and a storage medium to solve the problems of missed detection and misdetection of small-size target objects in target detection in related technologies.

[0005] According to one aspect of the present invention, there is provided a method for target detection of vehicle driving, the method comprising:

[0006] Obtaining an image of the driving environment of a target vehicle and a pre-trained target detection model, wherein the target detection model is used to detect target objects in the input image whose sizes are smaller than a preset size; the target detection model includes a channel-wise attention module, a pyramid pooling module, and a feature fusion module, the channel-wise attention module is used to determine the weight coefficients of the input features, the pyramid pooling module is used to perform pooling operations on the input features at multiple scales to capture the global and local context information of the input features, and the feature fusion module is used to fuse the input features;

[0007] Inputting the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, wherein the target detection result is at least used to represent whether the target object is included in the driving environment image.

[0008] According to another aspect of the present invention, there is provided a device for target detection of vehicle driving, the device comprising:

[0009] An acquisition module, configured to acquire an image of the driving environment of a target vehicle and a pre-trained target detection model, wherein the target detection model is used to detect target objects with sizes smaller than a preset size in the input image; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module, the channel application attention module is configured to determine the weight coefficients of the input features, the pyramid pooling module is configured to perform pooling operations on the input features at multiple scales to capture the global and local context information of the input features, and the feature fusion module is configured to fuse the input features;

[0010] A detection result determination module, configured to input the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, wherein the target detection result is at least used to characterize whether the target object is included in the driving environment image.

[0011] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0012] At least one processor; and

[0013] A memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the target detection method for vehicle driving according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the target detection method for vehicle driving according to any embodiment of the present invention when executed.

[0016] The technical solution of the embodiment of the present invention obtains the driving environment image of the target vehicle and a pre-trained target detection model. Since the target detection model is used to detect target objects with sizes smaller than a preset size in the input image; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module. The channel application attention module is used to determine the weight coefficients of the input features, the pyramid pooling module is used to perform pooling operations on the input features at multiple scales to capture the global and local context information of the input features, and the feature fusion module is used to fuse the input features, which can provide sufficient data support and tools for the target detection of vehicle driving; input the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, where the target detection result is at least used to characterize whether the target object is included in the driving environment image, and can accurately identify and locate the target object in the driving environment image, especially small-sized objects, solving the problems of missed detection and false detection of small-sized target objects in target detection in the related art, being able to effectively warn of potential dangers and enhance the safety of autonomous driving.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of a method for target detection of vehicle driving according to Embodiment 1 of the present invention;

[0020] Figure 2 is a flowchart of a method for target detection of vehicle driving according to Embodiment 2 of the present invention;

[0021] Figure 3 is a schematic structural diagram of a target detection model applicable to the embodiments of the present invention;

[0022] Figure 4 is a schematic structural diagram of a dynamic upsampling unit applicable to the embodiments of the present invention;

[0023] Figure 5 is a schematic structural diagram of a feature fusion unit applicable to the embodiments of the present invention

[0024] Figure 6 It is a schematic structural diagram of an object detection device for vehicle driving provided in Embodiment 3 of the present invention;

[0025] Figure 7 It is a schematic structural diagram of an electronic device for implementing the object detection method for vehicle driving applicable to the embodiments of the present invention. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0030] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the users and the authorization of the users should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0031] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure technical solution based on the prompt message.

[0032] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0034] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related regulations.

[0035] Embodiment 1

[0036] Figure 1 A flowchart of a target detection method for vehicle driving is provided for Embodiment 1 of the present invention. This embodiment is applicable to complex environments that require high-precision target detection, especially in the case of driverless or advanced driver assistance systems. This method can be executed by a target detection device for vehicle driving, and the target detection device for vehicle driving can be implemented in the form of hardware and / or software. Optionally, it is implemented through an electronic device, and the electronic device can be a mobile terminal, a PC terminal or a server, etc.

[0037] As Figure 1 shown, the method may specifically include:

[0038] S110. Obtain a driving environment image of a target vehicle and a pre-trained target detection model, where the target detection model is used to detect target objects with a size smaller than a preset size in the input image; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module. The channel application attention module is used to determine the weight coefficient of the input feature, the pyramid pooling module is used to perform pooling operations on the input feature at multiple scales to capture the global context information and local context information of the input feature, and the feature fusion module is used to fuse the input feature.

[0039] Among them, the target vehicle can be understood as the vehicle used for detection, and its driving environment is the object of analysis and detection. The driving environment image can be understood as the external environment information around the target vehicle captured by a camera or other imaging device installed on the target vehicle, which is used to provide visual data for subsequent analysis, helping to identify important information such as road conditions and other traffic participants. The driving environment image can be a high-resolution RGB image. The driving environment image includes but is not limited to targets such as road signs, other vehicles, and pedestrians. The target detection model can be understood as a deep learning model used to identify and locate specific objects (such as pedestrians or vehicles, etc.) in the driving environment image, such as Figure 3 as shown, which is used to process the input image and mark the target objects. The target object can be understood as the object that needs to be recognized, classified, or located in the driving environment image, and the size of the target object is smaller than the preset size and needs to be detected from the driving environment image by a pre-trained target detection model, including but not limited to targets such as road signs, other vehicles, and pedestrians. The preset size can be understood as a preset size threshold, and the threshold is used to define which target objects are regarded as "small-size" targets, that is, target objects with a size smaller than the preset size. The channel application attention module can be understood as assigning different weight coefficients to different channels of the input feature map, which is used to assist the target detection model to focus on important feature information and improve the detection accuracy of key elements. The pyramid pooling module can be understood as a component that performs pooling operations at multiple scales to capture the global and local context information of the input features, allowing the target detection model to understand the image content at different scales and enhancing the detection ability for objects of different sizes. The feature fusion module can be understood as a component responsible for integrating feature information from different levels or different scales. The feature fusion module improves the target detection model's understanding ability of the driving scene and target detection accuracy by merging multi-scale or multi-level information. The global context information can be understood as the overall information provided by all pixels or features in the driving environment image. Exemplarily, in the driving environment, understanding the global context can help distinguish whether a pedestrian is on a safe sidewalk or crossing the road, or determine whether the vehicle in front is changing lanes. The local context information can be understood as the information provided by the pixels or features in a specific area of the driving environment image. Different from the global context information that focuses on the entire image, the local context information focuses on the situation within a specific location or a small area, including the details of the objects within that area and their immediate environment. Exemplarily, in the autonomous driving scenario, understanding the local context around a pedestrian (such as the pedestrian's posture, walking direction, and their relationship with surrounding vehicles or obstacles, etc.) can more accurately predict the pedestrian's behavior and make corresponding driving decisions.

[0040] Based on the above solution, optionally, the training method of the target detection model includes: obtaining a sample driving environment image containing a target object, and determining an expected target object image corresponding to the sample driving environment image and annotated with the target object. Furthermore, training a pre-constructed target detection model using the sample driving environment image, and adjusting the pre-constructed target detection model based on the model output image and the expected target object image to obtain the target detection model. Among them, the methods for obtaining the sample driving environment image and the expected target object image can be the same or different. The target detection model can be a machine learning model, a deep learning model, etc. Further, the training parameters of the target detection model can be configured according to the structure of the target detection model used. The training parameters can include the learning rate, loss function, etc. By adopting this technical solution, by using the sample driving environment image containing the target object and its corresponding expected target object image annotated with the target object to finely adjust the pre-constructed target detection model, the accuracy and robustness of the model in a complex environment can be effectively improved, ensuring that it can accurately identify and locate various target objects. Based on the difference between the output result and the expected result, dynamically adjusting the model parameters (such as the learning rate, loss function, etc.), enabling the model to be continuously optimized until it reaches the best performance, not only enhances the model's ability to adapt to new scenarios, but also improves the autonomous driving system's understanding of the surrounding environment, thus significantly improving driving safety and decision-making efficiency. In addition, flexibly configuring the training parameters can also accelerate the convergence process, reducing training time and resource consumption.

[0041] To enhance the generalization ability of the target detection model, the sample driving environment image can be preprocessed, and then the pre-constructed target detection model can be trained using the preprocessed sample driving environment image. Among them, the preprocessing can include, but is not limited to, at least one of random rotation, translation, flipping, and scaling, etc., to increase the diversity of training data. It is also possible to perform normalization processing on the sample driving environment image so that its mean is 0 and the standard deviation is 1, or other forms of normalization, to facilitate the training of the target detection model.

[0042] S120: Input the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, where the target detection result is at least used to characterize whether the target object is included in the driving environment image.

[0043] Among them, the target detection result can be understood as the result output by the target detection model based on the input image, and can be used to determine whether there are potential dangerous target objects in the driving environment, including, but not limited to, information such as the position and category of the target object.

[0044] Based on the above solution, optionally, when it is detected that the driving environment image includes the target object, the target detection result includes the driving environment image marked with the target object; after obtaining the target detection result corresponding to the driving environment image, it further includes: determining the object type of the target object marked in the driving environment image, generating driving prompt information corresponding to the target object according to the object type, and displaying the driving prompt information. With this technical solution, by accurately identifying and marking the target object in the driving environment image, and then generating targeted driving prompt information according to the object type, the driving safety and driving experience are greatly improved. The marking can help the driver determine the surrounding environment and ensure a rapid response to potential dangers. The driving prompt information based on the object type further enhances the decision-making accuracy, enabling the driver to respond faster and promoting the practicality and reliability of the intelligent driving assistance system.

[0045] Among them, the object type can be understood as the specific classification of the target object, including but not limited to pedestrians, cars, motorcycles, etc. Determining the object type of the target object helps to generate more accurate driving prompt information, and different types of objects may require different response measures. The driving prompt information can be understood as the information generated according to the type of the target object, aiming to provide guidance or warnings for the driver. Generating corresponding driving prompt information according to the different object types (for example, "There is a pedestrian crossing the road ahead" or "There is a motorcycle approaching on the right", etc.) can help the driver make timely and correct responses and improve driving safety.

[0046] The technical solution of the embodiment of the present invention obtains the driving environment image of the target vehicle and a pre-trained target detection model. Since the target detection model is used to detect target objects in the input image whose size is smaller than a preset size; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module. The channel application attention module is used to determine the weight coefficient of the input feature, the pyramid pooling module is used to perform pooling operations on the input feature at multiple scales to capture the global and local context information of the input feature, and the feature fusion module is used to fuse the input feature, which can provide sufficient data support and tools for the target detection of vehicle driving; input the driving environment image into the target detection model to obtain the target detection result corresponding to the driving environment image, where the target detection result is at least used to represent whether the driving environment image includes the target object, and can accurately identify and locate the target object in the driving environment image, especially small-sized objects, solve the problems of missed detection and false detection of small-sized target objects in target detection in the related art, and can effectively warn of potential dangers and enhance the safety of autonomous driving.

[0047] Embodiment Two

[0048] Figure 2 The flowchart of a target detection method for vehicle driving provided in the second embodiment of the present invention. This embodiment further refines how to input the driving environment image into the target detection model to obtain the target detection result corresponding to the driving environment image on the basis of the above embodiment. Optionally, the target detection model further includes a convolution module for extracting features of different regions in the input image. The step of inputting the driving environment image into the target detection model to obtain the target detection result corresponding to the driving environment image includes: inputting the driving environment image into the convolution module for feature extraction to obtain the first features of multiple image regions, where the first features include texture features and semantic features; inputting the first features into the channel application attention module to obtain attention-enhanced features; inputting the attention-enhanced features into the pyramid pooling module to obtain the second features; and inputting the second features into the feature fusion module to obtain the target detection result corresponding to the driving environment image. For the specific implementation manner, reference may be made to the description of this embodiment. Among them, the same or similar technical features as those in the foregoing embodiments will not be described herein again.

[0049] As Figure 2 shown, the method may specifically include:

[0050] S210. Obtain a driving environment image of a target vehicle and a pre-trained target detection model, where the target detection model is used to detect target objects with sizes smaller than a preset size in the input image; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module. The channel application attention module is used to determine the weight coefficients of the input features. The pyramid pooling module is used to perform pooling operations on the input features at multiple scales to capture the global context information and local context information of the input features. The feature fusion module is used to fuse the input features. The target detection model further includes a convolution module for extracting features of different regions in the input image.

[0051] S220. Input the driving environment image into the convolution module for feature extraction to obtain the first features of multiple image regions, where the first features include texture features and semantic features.

[0052] Among them, the image region can be understood as the local pixel region covered by the convolutional module in the driving environment image. Convolutional modules of different sizes (such as 3×3, 5×5, etc.) can extract regional features of different granularities. Exemplarily, in the object detection task, the input driving environment image can be divided into multiple image regions according to the categories of detection features through the convolutional module, which helps the model focus on different parts of the driving environment image for feature extraction and analysis. For example, in the application scenario of an autonomous vehicle, the image may be divided into multiple regions to respectively focus on the road conditions ahead, the vehicles on the side, and pedestrians, etc. The first feature can be understood as the preliminary feature representation extracted from the input driving environment image, including the detailed information about the image (such as texture features) and the higher-level information (such as semantic features). The texture feature can be understood as the visual characteristics such as the surface material and pattern of the target object in the driving environment image, which is convenient for identifying target objects with similar shapes through texture differences. The semantic feature can be understood as the high-order information of the target object in the driving environment image, including but not limited to the category, movement direction, and scene context, etc., which can be used to distinguish the target type.

[0053] An optional implementation manner is that the input driving environment image can be a high-resolution RGB image. In order to extract meaningful features from the driving environment image, this method uses multiple convolutional layers to perform preliminary processing on the input image. Each convolutional module consists of a group of convolutional kernels (Filters or Kernels), which are used to scan each local region of the driving environment image, perform regional segmentation on the driving environment image, and generate corresponding feature maps (Feature Maps). By stacking multiple convolutional layers, the object detection model can capture different levels of features in the driving environment image, from low-level edge and texture information to high-level semantic information.

[0054] S230. Input the first feature into the channel application attention module to obtain an attention-enhanced feature.

[0055] Among them, the attention-enhanced feature can be understood as the feature processed by the channel application attention module, which is used to emphasize the features of the target object. By increasing the emphasis on the first feature and reducing the influence of background noise, the attention-enhanced feature enables the object detection model to more accurately locate and classify target objects in complex scenes.

[0056] An optional implementation manner is that by applying the attention module to some input channels and keeping the remaining channels unchanged, the attention of the object detection model to the key regions is strengthened, the detection effect of small targets and occluded targets is improved, while the computational redundancy is reduced and the effective capture ability of the global context is maintained.

[0057] Based on the above solution, optionally, inputting the first feature into the channel - applied attention module to obtain an attention - enhanced feature includes: inputting the first feature into the channel - applied attention module, determining the mean and variance of the feature values of each image region, determining the similarity according to the mean and the variance, determining the attention weight according to the similarity, and determining the attention - enhanced feature according to the attention weight and the first feature. By adopting this technical solution, the process of inputting the first feature into the channel - applied attention module includes determining the mean and variance of the feature values of each image region, and determining the attention weight according to the mean and the variance. Then, these weights are used to adjust the first feature to form an attention - enhanced feature, allowing the object - detection model to adaptively focus on relevant information according to the specific content of the image, thereby effectively improving the accuracy and efficiency of object detection.

[0058] Among them, the attention weight can be understood as a value assigned to each channel or image region based on the mean and variance, reflecting its importance relative to other parts. The attention weight is used to emphasize or suppress certain features, enabling the model to pay more attention to the key information that helps to complete the task while reducing the influence of irrelevant background noise.

[0059] An optional implementation manner is to define an image region of size 3×3 in the first feature, and calculate the mean and variance of the feature values within this window. These statistics are used to generate more accurate and discriminative attention weights. Then, through normalization processing and the Sigmoid activation function, the calculation results are converted into attention weights. Finally, the attention weights are restored to the shape of the first feature and multiplied by the input first feature to obtain the enhanced attention - enhanced feature.

[0060] Based on the above solution, specifically, the channel - applied attention module determines the similarity according to the mean and the variance based on the following formula:

[0061]

[0062] Among them, represents the mean within the image region w 2 represents the number of pixels within the window;

[0063] represents the variance of each pixel value within the image region; λ represents the regularization constant, a represents the first preset coefficient, and b represents the second preset coefficient.

[0064] Among them, the regularization constant can be understood as a very small positive value used to avoid the situation of the denominator being zero. The regularization constant can ensure numerical stability. Especially when performing division operations, it can prevent the algorithm from exhibiting abnormal behaviors when dealing with values close to zero, thereby ensuring the stability of the model training and inference processes. Exemplarily, the value of λ can be 1e-6. The first preset coefficient can be understood as a preset coefficient used to adjust the attention weights. The second preset coefficient can be understood as another preset coefficient, which is usually used together with the first preset coefficient to fine-tune the behavior of the attention mechanism to ensure that the final attention weights are always positive values. Exemplarily, the value of a can be 4 and the value of b can be 0.5.

[0065] Adopting this technical solution, by setting parameters in the process of calculating attention weights and generating attention-enhanced features, the model can better capture the most important information in the input image, thereby improving the accuracy and efficiency of object detection.

[0066] S240. Input the attention-enhanced feature into the pyramid pooling module to obtain a second feature.

[0067] Among them, the second feature can be understood as the feature representation output from the pyramid pooling module, which contains multi-level information. The second feature integrates information from different scales, providing a more comprehensive perspective to understand the input driving environment image, which is beneficial to improving the accuracy of object detection.

[0068] An optional implementation manner is to perform pooling operations on the input features at multiple scales through the pyramid pooling module to unify feature maps of different sizes into a fixed-size output, enabling the model to simultaneously process large, medium, and small targets, and significantly improving the detection ability for multi-scale targets.

[0069] S250. Input the second feature into the feature fusion module to obtain an object detection result corresponding to the driving environment image, where the object detection result is at least used to characterize whether the target object is included in the driving environment image.

[0070] Inputting the second feature into the feature fusion module is to comprehensively consider the feature information of the driving environment image at different scales, thereby improving the understanding ability of complex scenes and the accuracy of object detection. The finally obtained object detection result can not only characterize whether the target object is included in the driving environment image, but also provide more detailed information about the target object, which is of great significance for ensuring driving safety.

[0071] Based on the above solution, optionally, the feature fusion module includes a dynamic upsampling unit and a feature fusion unit. The dynamic upsampling unit is used to enhance the details of the input features, and the feature fusion unit is used to fuse the input features. The step of inputting the second feature into the spatial feature fusion module to obtain the target detection result corresponding to the driving environment image includes: inputting the upsampled feature into the feature fusion unit to obtain the target detection result corresponding to the driving environment image. By adopting this technical solution, the second feature is input into the dynamic upsampling unit to enhance the details of the feature and generate the upsampled feature. Then, the upsampled feature is passed to the feature fusion unit for multi-scale or multi-level information integration, so as to obtain the target detection result corresponding to the driving environment image. This combines the advantages of detail restoration and feature integration, which helps to improve the accuracy and reliability of target detection.

[0072] Among them, the upsampled feature can be understood as the feature obtained after being processed by the dynamic upsampling unit, which has a higher spatial resolution. The upsampled feature aims to restore or enhance the detail information that may be lost in the second feature map, thereby improving the accuracy of target detection, especially for the detection of small-sized targets.

[0073] Based on the above solution, optionally, the step of inputting the second feature into the dynamic upsampling unit to obtain the upsampled feature includes: inputting the second feature into the dynamic upsampling unit to determine the center coordinates and offsets of the image region through the dynamic upsampling unit, determining the target sampling position according to the center coordinates and the offsets, and determining the upsampled feature according to the target sampling position. By adopting this technical solution, first, the second feature is input into the dynamic upsampling unit. The dynamic upsampling unit will determine the center coordinates and offsets of the image region, allowing the system to accurately locate the region that needs attention and its specific adjustment requirements. Then, according to the center coordinates and offsets, the target sampling position is determined to select which specific positions are the key regions for upsampling. Finally, based on the selected target sampling position, the upsampled feature is generated, which can not only restore the detail information in the feature map but also flexibly adjust the sampling strategy according to actual needs, thereby improving the accuracy and robustness of target detection.

[0074] Among them, the center coordinates can be understood as the geometric center within the image region. The center coordinates can be used to locate the image region that needs to be focused on, ensuring that the details of this region can be accurately restored or enhanced during upsampling. The offset can be understood as the displacement relative to, for example, the center coordinates, which can be in the horizontal direction, vertical direction, or both. The offset is used to fine-tune the target sampling position, so that the dynamic upsampling operation is not limited to simple magnification, but can be flexibly adjusted according to actual needs to better capture key details. The target sampling position can be understood as the specific sampling point position calculated based on the center coordinates and the offset, which can be used to determine which pixels will be focused on when performing dynamic upsampling, thus affecting the quality of the finally generated upsampled features. Reasonably selecting the target sampling position helps to improve the spatial resolution and detail expressiveness of the feature map.

[0075] An optional implementation manner Figure 4 is a schematic structural diagram of the dynamic upsampling unit. The position of the sampling point is dynamically adjusted by the offset. By learning the offset Δx and Δy, the position of the center coordinates is dynamically adjusted: p'(x', y') = p(x, y) + (Δx, Δy), where p'(x', y') is the target sampling position, p(x, y) is the center coordinates, and Δx and Δy are the offsets.

[0076] To achieve more precise dynamic upsampling, the feature information within the local window is introduced at each target sampling position. This step ensures that the upsampled feature map not only retains the details of the original features but also enhances the ability to capture multi-scale features. The specific formula is as follows:

[0077]

[0078] In the formula, F up (x, y) represents the upsampled feature, ε(i, j) represents the local window, w(i, j) represents the weight coefficient, and F(p'(x', y') + ε(i, j)) represents the sampling feature value at the target sampling position.

[0079] Based on the above solution, optionally, the step of inputting the upsampled features into the feature fusion unit to obtain the target detection result corresponding to the driving environment image includes: inputting the upsampled features into the feature fusion unit to determine the average value and the maximum value of the upsampled features through the feature fusion unit, determining weights according to the average value and the maximum value, determining a fused feature map according to the weights and the upsampled features, and determining the target detection result corresponding to the driving environment image according to the fused feature map. With this technical solution, first, the upsampled features are input into the feature fusion unit, and the feature fusion unit calculates the average value and the maximum value of the upsampled features to determine the weights at each position; then, a fused feature map is generated based on the weights and the upsampled features themselves, which can effectively integrate multi-scale information, highlight key features, and reduce the influence of irrelevant information at the same time; finally, the target detection result corresponding to the driving environment image is determined based on the fused feature map, which not only improves the quality of feature representation but also enhances the target detection model's understanding ability of complex scenes and the accuracy of target detection.

[0080] Among them, the fused feature map can be understood as a comprehensive feature representation generated after being processed by the feature fusion unit. As Figure 5 shown in the structural schematic diagram of the feature fusion unit, the fused feature map combines information from different scales or different levels, provides a more comprehensive and detailed scene description than single-scale features, helps improve the target detection model's understanding ability of complex scenes, and improves the target detection result.

[0081] An optional implementation manner is to first calculate the average value μ c and the maximum value max c of each upsampled feature:

[0082]

[0083] max c = max h,w F fusion (h, w, c);

[0084] In the formula, h represents the height identifier of the pixel in the y direction, w represents the width identifier of the pixel in the x direction, c is the channel index, and F fusion (h, w, c) represents the value of the fused feature map at the position and channel.

[0085] Then, a multi-layer perceptron (MLP) is used to generate the weights w c of the channels:

[0086] w c = MLP([μ c , max c );

[0087] Finally, apply the weights of the channels to the fused feature map F att :

[0088] F att = w c ·F fusion 。

[0089] In an alternative implementation, there are usually three different levels of feature maps in the object detection model, corresponding to different detection scales. And each level has different resolutions and numbers of channels. The specific method is as follows:

[0090] Rescale the feature maps to the same resolution. Specifically, first, there are three feature maps X1, X2, X3 with different scales, corresponding to different receptive fields. Let the feature map of the input layer X l (l ∈ (1, 2, 3)) be X l , and adjust the feature maps of other feature layers L n (n ≠ l) to the same scale size as X l . To unify the scales of the feature maps, for feature maps smaller than a given threshold, first use 1×1 point convolution to compress the number of channels to the same as the target level, and then expand its size by interpolation; for feature maps larger than the given threshold, use a 3×3 convolution with a stride of 2 for downsampling, halve its size and adjust the number of channels; if 1 / 4 ratio downsampling is required, first use a max pooling layer with a stride of 2, followed by a 3×3 convolution layer with a stride of 2, so as to achieve the unification of the feature maps in terms of size and number of channels; then, after feature rescaling, adaptively fuse the feature maps of different scales by learning the optimal fusion weights at each position. For each position (i, j), the feature vectors from different levels are weighted and summed according to the learned spatial importance weights . The calculation formula is as follows:

[0091]

[0092] In the formula, is the fused feature map of the l-th layer at the position (i, j). is the fused feature map of the n-th layer at the position (i, j). is the fused feature map of the l-th layer at the position (i, j); finally, the feature map after feature fusion is used for the subsequent prediction. The detection head predicts the positions, sizes, class probabilities, and confidence scores of multiple bounding boxes for each grid cell. Use non-maximum suppression (NMS) to remove redundant overlapping bounding boxes and retain the most likely object detection results.

[0093] In the technical solution of the embodiment of the present invention, first, the convolution module extracts the texture and semantic features of different regions in the input image, laying a foundation for subsequent image processing; then, the channel application attention module dynamically adjusts the weights according to the importance of the features, enhancing the expressiveness of key information and reducing the interference of background noise; subsequently, the pyramid pooling module performs pooling operations at multiple scales to capture global and local context information, ensuring that targets of different sizes can be accurately recognized; finally, the feature fusion module integrates multi-scale features to generate a comprehensive and detailed feature representation, improving the model's understanding ability of complex scenes, not only improving the detection accuracy of small-sized targets, but also enhancing the robustness and adaptability of the system, and improving the safety and reliability of the driving vehicle.

[0094] Embodiment III

[0095] Figure 6 It is a schematic structural diagram of an object detection device for vehicle driving provided by Embodiment III of the present invention. As Figure 6 shown, the device includes: an acquisition module 610 and a detection result determination module 620.

[0096] The acquisition module 610 is configured to acquire the driving environment image of the target vehicle and a pre-trained target detection model, wherein the target detection model is used to detect target objects with sizes smaller than a preset size in the input image; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module, the channel application attention module is configured to determine the weight coefficients of the input features, the pyramid pooling module is configured to perform pooling operations on the input features at multiple scales to capture the global context information and local context information of the input features, and the feature fusion module is configured to fuse the input features; the detection result determination module 620 is configured to input the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, wherein the target detection result is at least used to represent whether the target object is included in the driving environment image.

[0097] In the technical solution of the embodiment of the present invention, a driving environment image of a target vehicle and a pre-trained target detection model are obtained through an acquisition module 610. Since the target detection model is used to detect target objects in the input image that are smaller than a preset size; the target detection model includes a channel application attention module, a pyramid pooling module, and a feature fusion module. The channel application attention module is used to determine the weight coefficients of the input features, the pyramid pooling module is used to perform pooling operations on the input features at multiple scales to capture the global and local context information of the input features, and the feature fusion module is used to fuse the input features, which can provide sufficient data support and tools for the target detection of vehicle driving; the driving environment image is input into the target detection model through a detection result determination module 620 to obtain a target detection result corresponding to the driving environment image, where the target detection result is at least used to characterize whether the target object is included in the driving environment image, and can accurately identify and locate the target object in the driving environment image, especially small-sized objects, solving the problems of missed detection and false detection of small-sized target objects in target detection in the related art, being able to effectively warn of potential dangers, and enhancing the safety of autonomous driving.

[0098] On the basis of the above solution, optionally, the target detection model further includes a convolution module, and the convolution module is used to extract features of different regions in the input image; the detection result determination module includes: a first feature determination sub-module, an attention enhanced feature determination sub-module, a second feature determination module, and a detection result determination sub-module. Among them, the first feature determination sub-module is used to input the driving environment image into the convolution module for feature extraction to obtain first features of multiple image regions, where the first features include texture features and semantic features; the attention enhanced feature determination sub-module is used to input the first features into the channel application attention module to obtain attention enhanced features; the second feature determination module is used to input the attention enhanced features into the pyramid pooling module to obtain second features; the detection result determination sub-module is used to input the second features into the feature fusion module to obtain a target detection result corresponding to the driving environment image.

[0099] On the basis of the above solution, optionally, the attention enhanced feature determination sub-module is used to: input the first features into the channel application attention module, determine the mean and variance of the feature values of each image region, determine the similarity according to the mean and the variance, determine the attention weight according to the similarity, and determine the attention enhanced feature according to the attention weight and the first features.

[0100] On the basis of the above solution, optionally, the channel application attention module determines the similarity according to the mean and the variance based on the following formula:

[0101]

[0102] Among them, represents the mean value within the image region, and w represents the number of pixels within the window; 2 represents the squared difference of each pixel value within the image region; λ represents the regularization constant, a represents the first preset coefficient, and b represents the second preset coefficient. Based on the above solution, optionally, the feature fusion module includes a dynamic upsampling unit and a feature fusion unit. The dynamic upsampling unit is used to enhance the details of the input features, and the feature fusion unit is used to fuse the input features; the detection result determination sub-module includes: an upsampled feature determination unit and a detection result determination unit. Among them, the upsampled feature determination unit is used to input the second feature into the dynamic upsampling unit to obtain an upsampled feature; the detection result determination unit is used to input the upsampled feature into the feature fusion unit to obtain a target detection result corresponding to the driving environment image.

[0103] Based on the above solution, optionally, the upsampled feature determination unit is specifically used to: input the second feature into the dynamic upsampling unit to determine the center coordinates and offsets of the image region through the dynamic upsampling unit, determine the target sampling position according to the center coordinates and the offsets, and determine the upsampled feature according to the target sampling position.

[0104] Based on the above solution, optionally, the detection result determination unit is specifically used to: input the upsampled feature into the feature fusion unit to determine the average value and the maximum value of the upsampled feature through the feature fusion unit, determine the weight according to the average value and the maximum value, determine the fused feature map according to the weight and the upsampled feature, and determine the target detection result corresponding to the driving environment image according to the fused feature map.

[0105] Based on the above solution, optionally, when it is detected that the driving environment image includes the target object, the target detection result includes the driving environment image marked with the target object; the target detection device for vehicle driving further includes: a driving prompt information display module. Among them, the driving prompt information display module is used to, after obtaining the target detection result corresponding to the driving environment image, determine the object type of the target object marked in the driving environment image, generate driving prompt information corresponding to the target object according to the object type, and display the driving prompt information.

[0106] Based on the above solution, optionally, when it is detected that the driving environment image includes the target object, the target detection result includes the driving environment image marked with the target object; the target detection device for vehicle driving further includes: a driving prompt information display module. Among them, the driving prompt information display module is used to, after obtaining the target detection result corresponding to the driving environment image, determine the object type of the target object marked in the driving environment image, generate driving prompt information corresponding to the target object according to the object type, and display the driving prompt information.

[0107] The target detection device for vehicle driving provided by the embodiments of the present invention can execute the target detection method for vehicle driving provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0108] Embodiment 4

[0109] Figure 7 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0110] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0111] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0112] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a target detection method for vehicle driving.

[0113] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above functions defined in the method of the embodiment of the present invention are executed.

[0114] In some embodiments, a target detection method for vehicle driving can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the target detection method for vehicle driving described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute a target detection method for vehicle driving in any other suitable manner (for example, by means of firmware).

[0115] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0117] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0119] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0120] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0122] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A target detection method for vehicle driving, characterized in that, Including: Obtain the driving environment image of the target vehicle and a pre-trained target detection model, where the target detection model is used to detect target objects with sizes smaller than a preset size in the input image; the target detection model includes a channel-wise attention module, a pyramid pooling module, and a feature fusion module. The channel-wise attention module is used to determine the weight coefficients of the input features, the pyramid pooling module is used to perform pooling operations on the input features at multiple scales to capture the global context information and local context information of the input features, and the feature fusion module is used to fuse the input features; Input the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image, where the target detection result is at least used to represent whether the target object is included in the driving environment image.

2. The method according to claim 1, wherein The target detection model further includes a convolution module, and the convolution module is used to extract features of different regions in the input image. The step of inputting the driving environment image into the target detection model to obtain a target detection result corresponding to the driving environment image includes: Input the driving environment image into the convolution module for feature extraction to obtain first features of multiple image regions, where the first features include texture features and semantic features; Input the first features into the channel-wise attention module to obtain attention-enhanced features; Input the attention-enhanced features into the pyramid pooling module to obtain second features; Input the second features into the feature fusion module to obtain a target detection result corresponding to the driving environment image.

3. The method according to claim 2, characterized in that, Inputting the first features into the channel-wise attention module to obtain attention-enhanced features includes: Input the first features into the channel-wise attention module, determine the mean and variance of the feature values of each image region, determine the similarity according to the mean and the variance, determine the attention weights according to the similarity, and determine the attention-enhanced features according to the attention weights and the first features.

4. The method according to claim 3, wherein The channel-wise attention module determines the similarity according to the mean and the variance based on the following formula: Among them, represents the mean value within the image region w 2 represents the number of pixels within the window; represents the squared difference of each pixel value within the image region; λ represents the regularization constant, a represents the first preset coefficient, and b represents the second preset coefficient.

5. The method according to claim 2, characterized in that The feature fusion module includes a dynamic upsampling unit and a feature fusion unit. The dynamic upsampling unit is used to enhance the details of the input features, and the feature fusion unit is used to fuse the input features; The step of inputting the second features into the spatial feature fusion module to obtain a target detection result corresponding to the driving environment image includes: Input the second features into the dynamic upsampling unit to obtain upsampled features; Input the upsampled features into the feature fusion unit to obtain a target detection result corresponding to the driving environment image.

6. The method according to claim 5, wherein Inputting the second features into the dynamic upsampling unit to obtain upsampled features includes: Input the second features into the dynamic upsampling unit to determine the center coordinates and offsets of the image regions through the dynamic upsampling unit, determine the target sampling positions according to the center coordinates and the offsets, and determine the upsampled features according to the target sampling positions.

7. The method according to claim 5, wherein Said inputting the upsampled features into the feature fusion unit to obtain a target detection result corresponding to the driving environment image includes: Inputting the upsampled features into the feature fusion unit to determine the average value and the maximum value of the upsampled features through the feature fusion unit, determining weights according to the average value and the maximum value, determining a fused feature map according to the weights and the upsampled features, and determining a target detection result corresponding to the driving environment image according to the fused feature map.

8. The method according to claim 1, characterized in that When it is detected that the target object is included in the driving environment image, the target detection result includes the driving environment image marked with the target object; After obtaining the target detection result corresponding to the driving environment image, it further includes: Determining the object type of the target object marked in the driving environment image, generating driving prompt information corresponding to the target object according to the object type, and displaying the driving prompt information.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the target detection method for vehicle driving according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the target detection method for vehicle driving according to any one of claims 1-8 when executed.