Image recognition method and device, electronic equipment, storage medium and vehicle

By introducing BEV feature and spatial relationship analysis in object detection, the problem of traditional object detection methods ignoring the spatial relationship between objects and the environment is solved, the accuracy and robustness of object detection are improved, and the overall performance of technologies such as autonomous driving and drone navigation is ensured.

CN120219783APending Publication Date: 2025-06-27XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311817232.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional object detection methods ignore the spatial relationship between objects and environment, resulting in low accuracy and robustness of object detection.

Method used

By processing the sample image, Bird's Eye View (BEV) features are obtained, and the travelable area segmentation, target detection and edge constraints are performed in the defined BEV space respectively, segmentation loss, detection loss and edge constraint loss are determined, model optimization is performed, and image recognition model is obtained.

Benefits of technology

By combining segmentation losses, detection losses and edge constraint losses, the accuracy and robustness of target detection are improved, ensuring that the spatial relationship between objects and the environment is combined, and the overall performance of collision avoidance and planning paths is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219783A_ABST
    Figure CN120219783A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method and device, electronic equipment, a storage medium and a vehicle, and the method comprises the steps: carrying out the processing of a sample image, and obtaining BEV features; according to the BEV features, performing drivable region segmentation, target detection and edge constraint on the sample image in a defined BEV space, and determining segmentation loss, detection loss and edge constraint loss; performing model optimization according to the segmentation loss, the detection loss and the edge constraint loss to obtain an image recognition model; and inputting an obtained to-be-recognized image into the image recognition model to obtain an image recognition result output by the image recognition model, the image recognition result being obtained by fusing a region segmentation result and a target detection result of the to-be-recognized image by the image recognition model. Through segmentation loss, detection loss and edge constraint loss, the spatial relationship between the object and the environment is combined, and the accuracy and robustness of target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image recognition method, apparatus, electronic device, storage medium, and vehicle. Background Art

[0002] In technical fields such as autonomous driving, drone navigation, and security monitoring, target detection is usually required to perform operations such as avoidance and warning. However, traditional target detection methods usually only focus on the recognition and positioning of objects, while ignoring the spatial relationship between the objects and the environment, resulting in low accuracy and robustness of target detection. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides an image recognition method, apparatus, electronic device, storage medium, and vehicle.

[0004] According to a first aspect of an embodiment of the present disclosure, an image recognition method is provided, including:

[0005] Processing a sample image to obtain a BEV feature;

[0006] According to the BEV feature, performing drivable area segmentation, target detection, and edge constraint on the sample image in a defined BEV space respectively, and determining a segmentation loss, a detection loss, and an edge constraint loss;

[0007] According to the segmentation loss, the detection loss, and the edge constraint loss, performing model optimization to obtain an image recognition model;

[0008] Inputting the obtained image to be recognized into the image recognition model, and obtaining an image recognition result output by the image recognition model, where the image recognition result is a fusion of a region segmentation result and a target detection result of the image recognition model for the image to be recognized.

[0009] Optionally, according to the BEV feature, performing drivable area segmentation on the sample image in a defined BEV space, and determining a segmentation loss, including:

[0010] According to the BEV feature, performing predictive segmentation on the drivable area of the sample image in a defined BEV space to obtain a drivable area and a non-drivable area;

[0011] According to the drivable area and its corresponding probability, the non-drivable area and its corresponding probability, determining a segmentation loss for the predictive segmentation of the drivable area through a cross-entropy loss function.

[0012] Optionally, predicting and segmenting the drivable area of the sample image in the defined BEV space according to the BEV feature to obtain a drivable area and a non-drivable area, including:

[0013] Performing image convolution on the sample images corresponding to multiple image acquisition devices according to the time information in the BEV feature to obtain a convolution image;

[0014] Inputting the convolution image into a transformer module for encoding and decoding, and predicting and segmenting the drivable area in the convolution image according to the spatial information in the BEV feature to obtain a drivable area and a non-drivable area.

[0015] Optionally, the detection loss includes a detection classification loss and a detection location loss. Performing object detection on the sample image in the defined BEV space according to the BEV feature to determine the detection loss, including:

[0016] Performing object detection on the sample image in the defined BEV space according to the BEV feature to generate an object detection box;

[0017] Determining the object classification and object location of the target object corresponding to the object detection box;

[0018] Determining the detection classification loss of the object classification and the detection location loss of the object location according to the annotation information in the sample image.

[0019] Optionally, determining the detection classification loss of the object classification and the detection location loss of the object location according to the annotation information in the sample image includes:

[0020] Determining the detection classification loss of the object classification by using a cross-entropy loss function according to the annotation classification of the target object in the annotation information;

[0021] Determining the detection location loss of the object location by using an L1 loss function according to the annotation location of the target object in the annotation information.

[0022] Optionally, performing edge constraint on the sample image in the defined BEV space according to the BEV feature to determine the edge constraint loss, including:

[0023] Performing edge contour recognition on the sample image in the defined BEV space to construct an edge line of the target object in the sample image;

[0024] Constructing an edge gradient field at the edge line;

[0025] Determine the loss value of the L1 loss function where the target detection box falls on the edge of the target object according to the edge gradient field, and obtain the edge constraint loss.

[0026] Optionally, the performing model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model includes:

[0027] Determine the average value of the gradient at the first moment and the variance at the second moment according to the segmentation loss, the detection loss, and the edge constraint loss, where the second moment is the moment of sampling the sample image adjacent to the first moment;

[0028] Determine an adaptive learning rate according to the average value, the variance, and a preset adaptive learning parameter;

[0029] Perform model optimization using gradient descent according to the adaptive learning rate to obtain an image recognition model.

[0030] Optionally, the processing the sample image to obtain a BEV feature includes:

[0031] Perform preprocessing on the sample image to obtain a preprocessed sample image, where the sample image is collected from multiple perspectives in the same scene;

[0032] Perform convolutional feature extraction on the preprocessed sample image to obtain a sample feature image;

[0033] Perform spatial feature extraction on the sample feature image to generate a BEV feature with spatio-temporal information.

[0034] According to a second aspect of the embodiments of the present disclosure, there is provided an image recognition device, including:

[0035] A processing module configured to process a sample image to obtain a BEV feature;

[0036] A determination module configured to perform drivable area segmentation, target detection, and edge constraint on the sample image in a defined BEV space according to the BEV feature, and determine a segmentation loss, a detection loss, and an edge constraint loss;

[0037] An optimization module configured to perform model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model;

[0038] An input module, configured to input the acquired image to be recognized into the image recognition model, and obtain an image recognition result output by the image recognition model, where the image recognition result is obtained by fusing the region segmentation result and the target detection result of the image recognition model for the image to be recognized.

[0039] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0040] A processor;

[0041] A memory for storing processor-executable instructions;

[0042] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspect.

[0043] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method according to any one of the second aspect are implemented.

[0044] According to a fifth aspect of the embodiments of the present disclosure, there is provided a vehicle, including:

[0045] A processor;

[0046] A memory for storing processor-executable instructions;

[0047] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the first aspect.

[0048] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0049] Processing the sample image to obtain BEV features; according to the BEV features, respectively performing drivable area segmentation, target detection and edge constraint on the sample image in the defined BEV space, determining the segmentation loss, the detection loss and the edge constraint loss; according to the segmentation loss, the detection loss and the edge constraint loss, performing model optimization to obtain an image recognition model; inputting the acquired image to be recognized into the image recognition model, and obtaining an image recognition result output by the image recognition model, where the image recognition result is obtained by fusing the region segmentation result and the target detection result of the image recognition model for the image to be recognized. By the segmentation loss, the detection loss and the edge constraint loss, the spatial relationship between the object and the environment is combined with each other, improving the accuracy and robustness of target detection.

[0050] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0051] The drawings herein are incorporated into and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0052] Figure 1 It is a flowchart of an image recognition method shown according to an exemplary embodiment.

[0053] Figure 2 It is a schematic diagram of an image edge point shown according to an exemplary embodiment.

[0054] Figure 3 It is a schematic diagram of an edge gradient field shown according to an exemplary embodiment.

[0055] Figure 4 It is an implementation shown according to an exemplary embodiment Figure 1 of the flowchart of step S11 in

[0056] Figure 5 It is a schematic diagram of an image recognition shown according to an exemplary embodiment.

[0057] Figure 6 It is a block diagram of an image recognition device shown according to an exemplary embodiment.

[0058] Figure 7 It is a block diagram of a terminal device for image recognition shown according to an exemplary embodiment.

[0059] Figure 8 It is a block diagram of a server for image recognition shown according to an exemplary embodiment.

[0060] Figure 9 It is a schematic diagram of a functional block diagram of a vehicle shown according to an exemplary embodiment. Detailed Description of the Embodiments

[0061] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0062] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0063] Before introducing an image recognition method, apparatus, electronic device, storage medium, and vehicle provided by the present disclosure, a method in a related scenario is briefly introduced first. In the related scenario, the recognition and positioning of objects are usually concerned, while the spatial relationship between the object and the environment is ignored. Moreover, the traditional Freespace detection method only determines the area without objects, that is, the drivable area Freespace, in the BEV space, which is independent of the target detection method, resulting in a lack of cooperation between the two, thus affecting the overall performance of collision avoidance and path planning.

[0064] In view of this, the present disclosure provides an image recognition method, aiming to combine the spatial relationship between the object and the environment, improve the accuracy and robustness of target detection, and further ensure the overall performance of collision avoidance and path planning for vehicles, drones, etc.

[0065] Figure 1 is a flowchart of an image recognition method shown according to an exemplary embodiment. Among them, the method can be applied to scenarios such as vehicle autonomous driving, obstacle avoidance driving, automatic parking, etc., and can also be applied to scenarios such as drone navigation and path planning, and can also be applied to scenarios such as security monitoring, as Figure 1 shown, the image recognition method may include the following steps.

[0066] In step S11, the sample image is processed to obtain BEV features.

[0067] In the embodiment of the present disclosure, the sample image may be obtained by multiple image acquisition devices collecting images of the target object in the same scene from different perspectives at the same moment. Among them, in different scenes, the target object may be different. For example, in the vehicle autonomous driving scene, images of lane lines, curbs, other vehicles, and pedestrians may be collected, while in the vehicle reverse driving scene, images of stop lines and railings may be collected, and in the drone flight scene, images of utility poles, plants, and buildings may be collected.

[0068] Furthermore, by processing the sample image, features from different perspectives in the same scene can be extracted, so as to determine the BEV (Bird's Eye View) features. It can be understood that the BEV features are features in the bird's eye view of the scene.

[0069] In step S12, according to the BEV features, the sample image is respectively subjected to drivable area segmentation, target detection, and edge constraint in the defined BEV space, and the segmentation loss, detection loss, and edge constraint loss are determined.

[0070] In the embodiments of the present disclosure, the "defined BEV space" for the sample image may refer to a BEV space defined in advance according to the sample image and specific task requirements. In this BEV space, processing such as drivable area, object detection, and edge constraint can be performed on the sample image, so as to obtain corresponding segmentation loss, detection loss, and edge constraint loss. These loss values can be used to optimize the image recognition model and improve the accuracy and robustness of the model when processing the image to be recognized.

[0071] For example, in the scenario of autonomous driving, the defined BEV space may be a ground plane composed of multiple pixels, and each pixel represents a projection area. By performing processing such as drivable area, object detection, and edge constraint on the sample image within this BEV space, the position information and shape information of objects such as the ground, vehicles, and pedestrians in the bird's-eye view can be obtained, thereby providing important data support for autonomous driving.

[0072] In the embodiments of the present disclosure, binary classification region segmentation can be performed on the regions in the sample image, and then the segmentation loss of the segmented regions is calculated. Among them, binary classification can improve the calculation speed of the model.

[0073] Among them, object detection can be to determine the category and position of all target objects of interest in the sample image through, for example, the R-CNN (Regions with CNN features) model. Due to the different appearances, shapes, and postures of various target objects under different perspectives, and the interference of factors such as illumination and occlusion during imaging, object detection can be to classify, locate, detect, and segment the target objects after adding object detection frames to the target objects, determine the detected target objects, and then determine the detection loss of the detection results according to the standard information for the sample image.

[0074] Among them, the edge of the target object in the sample image can be detected to obtain the edge detection result, and then the edge constraint loss corresponding to the edge detection result is calculated according to the edge in the standard information.

[0075] In step S13, according to the segmentation loss, the detection loss, and the edge constraint loss, model optimization is performed to obtain an image recognition model.

[0076] In the embodiments of the present disclosure, the segmentation loss, the detection loss, and the edge constraint loss are jointly used to guide the optimization of the model, and an optimized image recognition model is obtained. For example, the segmentation loss, the detection loss, and the edge constraint loss can be weighted and averaged first to obtain the total loss. Furthermore, the total loss can be used as the optimization objective, and the model parameters can be updated through optimization algorithms such as gradient descent. It should be noted that during the process of updating the model parameters, the importance of the three loss values needs to be considered simultaneously to ensure that the model does not ignore any aspect of information during the optimization process.

[0077] Through the above steps, the segmentation loss, the detection loss, and the edge constraint loss can be jointly used to guide the optimization of the model, and an optimized image recognition model is obtained. This model can consider three aspects: the drivable area, object detection, and edge constraint simultaneously, so as to achieve a more accurate image recognition result when processing the image to be recognized.

[0078] In step S14, the obtained image to be recognized is input into the image recognition model, and the image recognition result output by the image recognition model is obtained.

[0079] Wherein, the image recognition result is obtained by fusing the region segmentation result and the object detection result of the image recognition model for the image to be recognized.

[0080] In the embodiments of the present disclosure, the image recognition model can perform Freespace prediction and object detection on the input image to be recognized, and then fuse the region segmentation result and the object detection result to obtain the image recognition result, so that the spatial relationship between the object and the environment is combined with each other, improving the accuracy and robustness of object detection.

[0081] The above technical solution processes the sample image to obtain the BEV feature; according to the BEV feature, the drivable area segmentation, object detection, and edge constraint are respectively performed on the sample image in the defined BEV space to determine the segmentation loss, the detection loss, and the edge constraint loss; according to the segmentation loss, the detection loss, and the edge constraint loss, the model is optimized to obtain the image recognition model; the obtained image to be recognized is input into the image recognition model, and the image recognition result output by the image recognition model is obtained, and the image recognition result is obtained by fusing the region segmentation result and the object detection result of the image recognition model for the image to be recognized. Through the segmentation loss, the detection loss, and the edge constraint loss, the spatial relationship between the object and the environment is combined with each other, improving the accuracy and robustness of object detection. Furthermore, the performance of avoiding collisions and path planning for vehicles, drones, etc. is ensured as a whole.

[0082] Optionally, according to the BEV feature, perform drivable area segmentation on the sample image in the defined BEV space, and determine the segmentation loss, including:

[0083] According to the BEV feature, perform predictive segmentation on the drivable area of the sample image in the defined BEV space to obtain a drivable area and a non-drivable area.

[0084] In the embodiments of the present disclosure, a Freespace detection module can be used to predict Freespace in the sample image. Freespace can be an area without obstacles such as pedestrians, vehicles, vegetation, and buildings from the BEV perspective. Among them, the Freespace detection module adopts the segmentation representation of the sample image under the BEV feature, where the drivable area is represented as 1 and the non-drivable area is represented as 0.

[0085] According to the drivable area and its corresponding probability, and the non-drivable area and its corresponding probability, determine the segmentation loss of the predictive segmentation of the drivable area through a cross-entropy loss function.

[0086] In the embodiments of the present disclosure, the cross-entropy loss function can be:

[0087]

[0088] Among them, yi represents the classification of the i-th area in the sample image. If the i-th area is classified as a drivable area, then yi is 1; if the i-th area is classified as a non-drivable area, then yi is 0. pi represents the probability that the i-th area in the sample image is a drivable area. L represents the average value of the segmentation losses of all areas in the sample image.

[0089] Optionally, the step of performing predictive segmentation on the drivable area of the sample image in the defined BEV space according to the BEV feature to obtain a drivable area and a non-drivable area includes:

[0090] According to the time information in the BEV feature, perform image convolution on the sample images corresponding to multiple image acquisition devices to obtain a convolution image.

[0091] In the embodiments of the present disclosure, the sample images collected by multiple image acquisition devices can be preprocessed first, which can include operations such as scaling, cropping, and normalization, so as to unify the sample images collected by multiple image acquisition devices into the same perspective space. Furthermore, by taking into account the time information in the BEV feature, the sample images collected by multiple image acquisition devices can be arranged in chronological order and convolution operations can be performed using a convolutional neural network (CNN).

[0092] Further, the sample images collected by multiple image acquisition devices can be arranged into a three-dimensional tensor in chronological order. Among them, the tensor of the first dimension represents the position of the image, the tensor of the second dimension represents time, and the tensor of the third dimension represents the number of channels of the image. Then, this three-dimensional tensor is used as the input of the convolutional neural network, and the convolutional layer in the CNN is used to perform a convolutional operation on the input data to obtain a convolutional image. During the convolutional operation, different convolutional kernels and convolutional parameters can be set to adjust the intensity and effect of the convolutional operation.

[0093] Through the above method, according to the time information in the BEV feature, the sample images corresponding to multiple image acquisition devices can be subjected to image convolution to obtain a convolutional image. It can better capture the spatial and time information in the image, thereby improving the image recognition accuracy of the model.

[0094] Input the convolutional image into the transformer module for encoding and decoding, and predict and segment the drivable area in the convolutional image according to the spatial information in the BEV feature to obtain a drivable area and a non-drivable area.

[0095] In the embodiment of the present disclosure, the convolutional image is sent into the transformer module for processing. The transformer is a neural network architecture with a self-attention mechanism, which can capture the long-range dependencies and global information in the convolutional image. When processing the convolutional image, the transformer can encode and decode the convolutional image through the self-attention mechanism and multiple encoder-decoder layers to extract feature information with a higher feature level than that in the convolutional image.

[0096] In the embodiment of the present disclosure, combined with the spatial information in the BEV feature, it is predicted whether the area is a passable area. Furthermore, on the basis of combining object detection, it can be accurately determined whether the area is drivable, avoiding incorrect area recognition caused by spatial influence.

[0097] After being processed by the above technical solution through the CNN and the transformer module, a BEV feature can be obtained. The BEV feature can represent the positions and feature information of all pixels in the sample image in the defined BEV space. Through the BEV feature, the spatial relationship and time information in the image can be better understood, thereby improving the image recognition accuracy of the model.

[0098] Optionally, the detection loss includes a detection classification loss and a detection position loss. According to the BEV feature, object detection is performed on the sample image in the defined BEV space to determine the detection loss, including:

[0099] Perform object detection on the sample image within the defined BEV space according to the BEV features to generate object detection bounding boxes.

[0100] In the embodiments of the present disclosure, for example, candidate regions in the sample image are determined, and then sliding window marking is performed on the candidate regions from left to right and from top to bottom to generate object detection bounding boxes for the target objects.

[0101] Determine the object classification and object position of the target object corresponding to the object detection bounding box.

[0102] According to the annotation information in the sample image, determine the detection classification loss of the object classification and the detection position loss of the object position.

[0103] In the embodiments of the present disclosure, for the sample image, the classification and position of the target object can be annotated. Then, according to the classification of the target object in the annotation and the object classification in the detection result, the classification difference is calculated to determine the detection classification loss. According to the position of the target object in the annotation and the object position in the detection result, the position difference is calculated to obtain the detection position loss.

[0104] For example, the detection classification loss can be:

[0105] L cls (p i ,p i *) = -log[p i * × p i + (1 - p i *)(1 - p i )]

[0106] Where L cls (p i ,p i *) is the detection classification loss of the i-th target object in the sample image, p i is the classification probability of the i-th target object in the detection result of the sample image, and p i * is the classification label of the i-th target object annotated in the sample image.

[0107] The detection position loss can be:

[0108] L reg (t i ,t i *) = R(t i - t i *)

[0109] Where L reg (t i ,t i*) The detection location loss of the i-th target object in the sample image, t i is a vector used to represent the 4 parameterized coordinates for predicting the location of the i-th target object. t i * represents the 4 parameterized coordinates for annotating the location of the i-th target object.

[0110] Optionally, determining the detection classification loss of the object classification and the detection location loss of the object location according to the annotation information in the sample image includes:

[0111] According to the annotation classification of the target object in the annotation information, using the cross-entropy loss function to determine the detection classification loss of the object classification.

[0112] In the embodiments of the present disclosure, similar to the classification after region segmentation, the cross-entropy loss function is used to determine the classification loss of the binary classification of the target object.

[0113] According to the annotation location of the target object in the annotation information, using the L1 loss function to determine the detection location loss of the object location.

[0114] In the embodiments of the present disclosure, the detection location loss may respectively include the losses of the position xyz, size length width height lwh, and yaw angle yaw of the detection box. Among them, the L1 loss function can be used to calculate the mean absolute error (MAE) of the position losses of all target objects, that is, calculate the average of the absolute differences between the model prediction value f(x) and the standard true value y to obtain the detection location loss of the object location.

[0115] Optionally, according to the BEV feature, performing edge constraint on the sample image in the defined BEV space to determine the edge constraint loss, including:

[0116] Performing edge contour recognition on the sample image in the defined BEV space to construct the edge line of the target object in the sample image.

[0117] In the embodiments of the present disclosure, taking a vehicle as an example, the radar and camera of the vehicle can collect radar information and image information around the vehicle. Refer to Figure 2 shown. Based on the BEV feature, the edge points facing the vehicle in the sample image can be detected. As Figure 2 highlighted in white, Figure 2 and the gray projection in

[0118] In the embodiments of the present disclosure, for example, the Canny edge detector can be used to identify the edge contour of the target object in the sample image. First, the Gaussian blur algorithm can be used to filter the sample image. Then, with the help of the Sobel filter, the intensity and direction of the target object at the edge can be found, and by applying non-maximum suppression, stronger edges can be isolated and refined into a one-pixel-wide line. Finally, the simplest edges can be isolated using hysteresis, and then the edge line of the target object in the sample image can be constructed based on geometric relationships.

[0119] Construct an edge gradient field at the edge line.

[0120] See Figure 3 As shown, an edge gradient field is constructed in the direction of the edge line towards the camera. The edge gradient field can be a Gaussian gradient field, where the closer to the camera, the weaker the gradient field, and the strongest gradient field is at the edge line.

[0121] According to the edge gradient field, determine the loss value of the L1 loss function where the target detection box falls on the edge of the target object, and obtain the edge constraint loss.

[0122] In the embodiments of the present disclosure, calculate the loss value of the L1 loss function where the target detection box corresponding to each target object in the sample image falls in the edge gradient field. It can be understood that the loss value loss of L1 can refer to the determination of the position where the target detection box of the target object falls in this Gaussian gradient field. In this way, the target detection box of the target object and the actual edge of the target object can be made closer, so that the position estimation of the target object can be more accurate.

[0123] Optionally, the performing model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model includes:

[0124] According to the segmentation loss, the detection loss, and the edge constraint loss, determine the average value of the gradient at the first moment and the variance at the second moment. The second moment is the moment of sampling the sample image adjacent to the first moment.

[0125] In the embodiments of the present disclosure, the average value and variance of the segmentation loss, the detection loss, and the edge constraint loss at the first moment and the second moment can be calculated respectively, and then the average value of the gradient at the first moment and the variance at the second moment can be calculated.

[0126] According to the average value, the variance, and a preset adaptive learning parameter, determine an adaptive learning rate.

[0127] In the embodiments of the present disclosure, the average value, the variance, and a preset adaptive learning parameter are substituted into a preset formula corresponding to the Adam algorithm, and an adaptive learning rate in the preset formula can be calculated.

[0128] According to the adaptive learning rate, gradient descent is used to optimize the model to obtain an image recognition model.

[0129] In the embodiments of the present disclosure, the model is optimized by gradient descent optimization algorithms such as the Adam algorithm.

[0130] Among them, gradient descent can be performed through the following formula:

[0131]

[0132] Among them, θ t is the model parameter at the first moment, θ t-1 is the model parameter at the second moment, η is the adaptive learning rate, is the gradient.

[0133] Further, according to the Adam algorithm it is obtained that and Among them, is the average value of the gradient at the first moment, is the variance of the gradient at the second moment, β1 is the exponential decay rate of the previous estimate in the Adam algorithm, β2 is the exponential decay rate of the previous estimate in the Adam algorithm, m t-2 is the momentum accumulated from the 0th moment to the (t - 2)th moment, m t-1 is the momentum accumulated from the 0th moment to the (t - 1)th moment, and the momentum enables each model optimization to retain the previous optimization direction, v t-2 is the parameter for adjusting the adaptive learning rate with the gradient accumulated from the 0th moment to the (t - 2)th moment, v t-1 is the parameter for adjusting the adaptive learning rate with the gradient accumulated from the 0th moment to the (t - 1)th moment. Thus, it can be obtained that

[0134]

[0135] Optionally, as shown in Figure 4 In step S11, the processing of the sample image to obtain the BEV feature includes:

[0136] In step S111, the sample image is preprocessed to obtain a preprocessed sample image.

[0137] Among them, the sample image is collected from multiple perspectives in the same scene.

[0138] In the embodiments of the present disclosure, the sample images collected from multiple perspectives in the same scene can be the sample images collected by multiple cameras in the same scene at the same moment.

[0139] In the embodiments of the present disclosure, the preprocessing of the sample images may include operations such as data augmentation and normalization, so as to improve the generalization ability of the model.

[0140] Among them, data augmentation can be, for example, horizontal and vertical flipping of the sample images, rotation of the sample images, outward or inward scaling of the sample images. In the embodiments of the present disclosure, when scaling outward, the size of the scaled image is larger than the size of the sample image, and when scaling inward, the size of the scaled image is smaller than the size of the sample image. Cropping the sample images to obtain partial images of the sample images, moving the sample images along the X or Y direction (or both), and adding Gaussian noise to the sample images, either one or more of them.

[0141] Among them, normalization can normalize the pixel values of the pixel points in the sample images. For example, normalize the pixel values of the pixel points in the sample images to the range of (0, 1). Exemplarily, it can traverse and calculate the pixel value of each pixel point minus the minimum pixel value in the sample image to obtain the pixel difference, calculate the maximum pixel difference between the maximum pixel value and the minimum pixel value in the sample image, and then divide the pixel difference by the maximum pixel difference to normalize the pixel values of the pixel points in the sample image.

[0142] In step S112, convolutional feature extraction is performed on the preprocessed sample images to obtain sample feature images.

[0143] In the embodiments of the present disclosure, referring to Figure 5 as shown, a convolutional neural network (CNN) can be used to perform feature extraction on the preprocessed sample images corresponding to the multi-perspective sample images at the current moment t to obtain the multi-camera feature F t , and then generate sample feature images according to the multi-camera feature F t .

[0144] In step S113, spatial feature extraction is performed on the sample feature images to generate BEV features with spatio-temporal information.

[0145] Continue to refer to Figure 5 as shown, the spatial feature extraction of the sample feature images can be performed through a transformer network. The sample feature images corresponding to multiple perspectives in the same scene are input into the transformer network. The transformer network generates the current BEV feature according to the historical BEV feature B corresponding to the previous moment of the current momentt-1 Perform a BEV feature query on the sample feature image at the current moment to obtain the BEV feature B at the current moment t Then, input the BEV feature B at the current moment t into the object detection model and the Freespace detection module respectively to obtain the region segmentation result and the object detection result. By traversing and executing this step, a BEV feature with spatio-temporal information for the environment can be generated. Among them, the transformer network may include multiple Encoders and Decoders.

[0146] The above technical solution can not only improve the generalization ability of the model, but also generate a BEV feature with spatio-temporal information of the environment, so that in the process of Freespace prediction and object detection, a comprehensive understanding of the object and its environment can be achieved, thereby improving the accuracy and robustness of object detection.

[0147] The embodiments of the present disclosure also provide an image recognition device. Refer to Figure 6 the image recognition device shown. The image recognition device includes: a processing module 610, a determination module 620, an optimization module 630, and an input module 640.

[0148] Among them, the processing module 610 is configured to process the sample image to obtain a BEV feature;

[0149] The determination module 620 is configured to perform drivable area segmentation, object detection, and edge constraint on the sample image in the defined BEV space according to the BEV feature, and determine the segmentation loss, the detection loss, and the edge constraint loss;

[0150] The optimization module 630 is configured to perform model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model;

[0151] The input module 640 is configured to input the acquired image to be recognized into the image recognition model to obtain the image recognition result output by the image recognition model, and the image recognition result is a fusion of the region segmentation result and the object detection result of the image recognition model for the image to be recognized.

[0152] Optionally, the determination module 620 includes a segmentation sub-module, which is configured to:

[0153] Perform predictive segmentation on the drivable area of the sample image in the defined BEV space according to the BEV feature to obtain the drivable area and the non-drivable area;

[0154] Determine the segmentation loss for predicting the drivable area through the cross-entropy loss function based on the drivable area and its corresponding probability, and the non-drivable area and its corresponding probability.

[0155] Optionally, the segmentation sub-module is configured to:

[0156] Perform image convolution on the sample images corresponding to multiple image acquisition devices according to the time information in the BEV feature to obtain a convolved image;

[0157] Input the convolved image into a transformer module for encoding and decoding, and predict and segment the drivable area in the convolved image according to the spatial information in the BEV feature to obtain a drivable area and a non-drivable area.

[0158] Optionally, the detection loss includes a detection classification loss and a detection position loss, and the determination module 620 includes a detection sub-module, which is configured to:

[0159] Perform object detection on the sample image in the defined BEV space according to the BEV feature to generate an object detection box;

[0160] Determine the object classification and object position of the target object corresponding to the object detection box;

[0161] Determine the detection classification loss of the object classification and the detection position loss of the object position according to the annotation information in the sample image.

[0162] Optionally, the detection sub-module is configured to:

[0163] Determine the detection classification loss of the object classification by using the cross-entropy loss function according to the annotation classification of the target object in the annotation information;

[0164] Determine the detection position loss of the object position by using the L1 loss function according to the annotation position of the target object in the annotation information.

[0165] Optionally, the determination module 620 includes an identification sub-module, which is configured to:

[0166] Perform edge contour recognition on the sample image in the defined BEV space to construct the edge line of the target object in the sample image;

[0167] Construct an edge gradient field at the edge line;

[0168] Determine the loss value of the L1 loss function where the object detection box falls on the edge of the target object according to the edge gradient field to obtain the edge constraint loss.

[0169] Optionally, the optimization module 630 is configured to:

[0170] Determine the average value of the gradient at the first moment and the variance at the second moment according to the segmentation loss, the detection loss, and the edge constraint loss, where the second moment is the moment adjacent to the first moment when sampling the sample image;

[0171] Determine an adaptive learning rate according to the average value, the variance, and a preset adaptive learning parameter;

[0172] Optimize the model using gradient descent according to the adaptive learning rate to obtain an image recognition model.

[0173] Optionally, the processing module 610 is configured to:

[0174] Perform preprocessing on the sample image to obtain a preprocessed sample image, where the sample image is collected from multiple perspectives in the same scene;

[0175] Extract convolutional features from the preprocessed sample image to obtain a sample feature image;

[0176] Extract spatial features from the sample feature image to generate a BEV feature with spatio-temporal information.

[0177] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0178] Embodiments of the present disclosure further provide an electronic device, including:

[0179] A processor;

[0180] A memory for storing processor-executable instructions;

[0181] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method described in any one of the foregoing embodiments.

[0182] According to the fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the method described in any one of the foregoing embodiments are implemented.

[0183] Figure 7 It is a block diagram of a device 700 for image recognition shown according to an exemplary embodiment. For example, the device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0184] Referring to Figure 7 , the apparatus 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output interface 712, a sensor component 714, and a communication component 716.

[0185] The processing component 702 generally controls the overall operation of the apparatus 700, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 702 may include one or more modules to facilitate interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate interaction between the multimedia component 708 and the processing component 702.

[0186] The memory 704 is configured to store various types of data to support the operation of the apparatus 700. Examples of such data include instructions for any application or method operating on the apparatus 700, contact data, phone book data, messages, pictures, videos, and the like. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0187] The power component 706 provides power to the various components of the apparatus 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 700.

[0188] The multimedia component 708 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0189] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0190] The input / output interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0191] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the device 700. For example, the sensor component 714 can detect the on / off state of the device 700, the relative positioning of components, such as the display and keypad of the device 700. The sensor component 714 can also detect a change in the position of the device 700 or a component of the device 700, the presence or absence of user contact with the device 700, the orientation or acceleration / deceleration of the device 700, and the temperature change of the device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0192] The communication component 716 is configured to facilitate communication, either wired or wirelessly, between the device 700 and other devices. The device 700 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0193] In an exemplary embodiment, the device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above image recognition method.

[0194] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of the device 700 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0195] Figure 8 is a block diagram of a device 800 for image recognition shown according to an exemplary embodiment. For example, the device 800 can be provided as a server. Referring to Figure 8 , the device 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by a memory 832 for storing instructions executable by the processing component 822, such as application programs. The application programs stored in the memory 832 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 822 is configured to execute instructions to perform the above image recognition method.

[0196] The device 800 may further include a power component 826 configured to perform power management of the device 800, a wired or wireless network interface 850 configured to connect the device 800 to a network, and an input / output interface 858. The device 800 can operate based on an operating system stored in the memory 832.

[0197] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-described image recognition method when executed by the programmable device.

[0198] An embodiment of the present disclosure also provides a vehicle, including:

[0199] A processor;

[0200] A memory for storing processor-executable instructions;

[0201] Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of the foregoing embodiments.

[0202] Figure 9 is a block diagram of a vehicle 900 shown according to an exemplary embodiment. For example, the vehicle 900 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 900 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0203] Referring to Figure 9 , the vehicle 900 may include various subsystems. For example, the infotainment system 910, the perception system 920, the decision control system 930, the drive system 940, and the computing platform 950. Among them, the vehicle 900 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 900 may be interconnected by wired or wireless means.

[0204] In some embodiments, the infotainment system 910 may include a communication system, an entertainment system, and a navigation system, etc.

[0205] The perception system 920 may include several sensors for sensing information about the environment around the vehicle 900. For example, the perception system 920 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter wave radar, an ultrasonic radar, and a camera device.

[0206] The decision control system 930 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0207] The drive system 940 can include components that provide motive power for the vehicle 900. In one embodiment, the drive system 940 can include an engine, an energy source, a transmission system, and wheels. The engine can be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.

[0208] Some or all functions of the vehicle 900 are controlled by the computing platform 950. The computing platform 950 can include at least one processor 951 and a memory 952, and the processor 951 can execute instructions 953 stored in the memory 952.

[0209] The processor 951 can be any conventional processor, such as a commercially available CPU. The processor can also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0210] The memory 952 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0211] In addition to the instructions 953, the memory 952 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 952 can be used by the computing platform 950.

[0212] In an embodiment of the present disclosure, the processor 951 can execute the instructions 953 to complete all or part of the steps of the above image recognition method.

[0213] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0214] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image recognition method, characterized in that, Including: Processing a sample image to obtain BEV features; According to the BEV features, performing drivable area segmentation, object detection, and edge constraint on the sample image in the defined BEV space respectively to determine a segmentation loss, a detection loss, and an edge constraint loss; According to the segmentation loss, the detection loss, and the edge constraint loss, performing model optimization to obtain an image recognition model; Inputting the obtained image to be recognized into the image recognition model to obtain an image recognition result output by the image recognition model, where the image recognition result is a fusion of the region segmentation result and the object detection result of the image recognition model for the image to be recognized.

2. The method according to claim 1, wherein According to the BEV features, performing drivable area segmentation on the sample image in the defined BEV space to determine a segmentation loss, including: According to the BEV features, performing predictive segmentation on the drivable area of the sample image in the defined BEV space to obtain a drivable area and a non-drivable area; According to the drivable area and its corresponding probability, and the non-drivable area and its corresponding probability, determining the segmentation loss of the predictive segmentation of the drivable area through a cross-entropy loss function.

3. The method according to claim 2, wherein The step of performing predictive segmentation on the drivable area of the sample image in the defined BEV space according to the BEV features to obtain a drivable area and a non-drivable area includes: According to the time information in the BEV features, performing image convolution on the sample images corresponding to multiple image acquisition devices to obtain a convolution image; Inputting the convolution image into a transformer module for encoding and decoding, and performing predictive segmentation on the drivable area in the convolution image according to the spatial information in the BEV features to obtain a drivable area and a non-drivable area.

4. The method according to claim 1, characterized in that, The detection loss includes a detection classification loss and a detection position loss; According to the BEV features, performing object detection on the sample image in the defined BEV space to determine a detection loss, including: According to the BEV features, performing object detection on the sample image in the defined BEV space to generate an object detection box; Determining the object classification and object position of the target object corresponding to the object detection box; According to the annotation information in the sample image, determining the detection classification loss of the object classification and the detection position loss of the object position.

5. The method according to claim 4, characterized in that, The step of determining the detection classification loss of the object classification and the detection position loss of the object position according to the annotation information in the sample image includes: According to the annotation classification of the target object in the annotation information, using a cross-entropy loss function to determine the detection classification loss of the object classification; According to the annotation position of the target object in the annotation information, using an L1 loss function to determine the detection position loss of the object position.

6. The method according to claim 4, wherein According to the BEV features, performing edge constraint on the sample image in the defined BEV space to determine an edge constraint loss, including: Performing edge contour recognition on the sample image in the defined BEV space to construct an edge line of the target object in the sample image; Construct an edge gradient field at the edge line; According to the edge gradient field, determine the loss value of the L1 loss function where the target detection box falls on the edge of the target object, and obtain the edge constraint loss.

7. The method according to claim 1, characterized in that, The performing model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model includes: According to the segmentation loss, the detection loss, and the edge constraint loss, determine the average value of the gradient at the first moment and the variance at the second moment, where the second moment is the next adjacent moment after the first moment for sampling the sample image, and the second moment is the next adjacent moment after the first moment for sampling the sample image; According to the average value, the variance, and a preset adaptive learning parameter, determine an adaptive learning rate; According to the adaptive learning rate, perform model optimization using gradient descent to obtain an image recognition model.

8. The method according to any one of claims 1-7, characterized in that, The processing the sample image to obtain a BEV feature includes: Perform preprocessing on the sample image to obtain a preprocessed sample image, where the sample image is collected from multiple perspectives in the same scene; Perform convolutional feature extraction on the preprocessed sample image to obtain a sample feature image; Perform spatial feature extraction on the sample feature image to generate a BEV feature with spatio-temporal information.

9. An image recognition device, characterized in that, Includes: A processing module configured to process a sample image to obtain a BEV feature; A determination module configured to perform drivable area segmentation, target detection, and edge constraint on the sample image in a defined BEV space according to the BEV feature, and determine a segmentation loss, a detection loss, and an edge constraint loss; An optimization module configured to perform model optimization according to the segmentation loss, the detection loss, and the edge constraint loss to obtain an image recognition model; An input module configured to input an image to be recognized obtained into the image recognition model to obtain an image recognition result output by the image recognition model, where the image recognition result is a result of fusing the area segmentation result and the target detection result of the image recognition model for the image to be recognized.

10. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-8.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.

12. A vehicle, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions stored in the memory to implement the method according to any one of claims 1-8.