Image detection method, terminal equipment and computer readable storage medium
By segmenting and mapping the image into the coordinate system at the top view angle, the foreground features in the image are extracted and detected, and the problem of redundant information affecting detection efficiency in the prior art is solved, and efficient and accurate image detection is achieved.
Patent Information
- Application Number
- CN202411996936.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
AI Technical Summary
Existing image detection methods need to process a large amount of redundant information, which affects the detection efficiency and effect.
By performing feature segmentation processing on the feature information of the image, the features of the foreground part are extracted and mapped into the coordinate system at the top view angle to detect the target object.
减少了背景部分的冗余信息,提高了检测效率和效果,并通过包含距离信息的特征提高了检测精度。
Smart Images

Figure CN120032097A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular, relates to an image detection method, a terminal device, and a computer-readable storage medium. Background Art
[0002] With the development of image processing technology, its application scope is becoming more and more extensive. For example, image processing technology can be applied to the field of autonomous driving, by performing image detection on the road ahead of the vehicle to perceive obstacles, thereby achieving automatic control of the vehicle.
[0003] Current image detection methods usually detect the entire image, which requires processing a lot of redundant information, affecting the efficiency and effect of image detection. Summary of the invention
[0004] The embodiments of the present application provide an image detection method, a terminal device, and a computer-readable storage medium, which can effectively improve the efficiency and detection effect of image detection.
[0005] In a first aspect, an embodiment of the present application provides an image detection method, comprising:
[0006] Acquire a first image to be processed; wherein the first image is an RGB image;
[0007] Extracting a first feature of a foreground portion of the first image by performing feature segmentation processing on feature information of the first image;
[0008] Mapping the first feature to a preset coordinate system to obtain a second feature; wherein the preset coordinate system is a coordinate system corresponding to the image in a top-down perspective;
[0009] A target object in the first image is detected according to the second feature.
[0010] In the embodiment of the present application, the feature information of the foreground part of the image is extracted, and then the feature information of the foreground part is used for detection, which reduces the redundant information of the background part of the image, not only effectively improves the detection efficiency, but also greatly improves the detection effect. In addition, since the image under the top-down perspective contains distance information, the pixel features of the RGB image are mapped to the coordinate system under the top-down perspective, so that the second feature contains the distance information, and detection based on the second feature helps to improve the detection accuracy.
[0011] In a possible implementation manner of the first aspect, performing feature segmentation processing on feature information of the first image to extract a first feature of a foreground portion of the first image includes:
[0012] Performing feature extraction processing on the first image to obtain a third feature;
[0013] A first feature corresponding to the foreground portion of the first image is extracted from the third feature.
[0014] In the embodiment of the present application, the foreground part is extracted according to the feature information of the image, so as to improve the detection accuracy of the foreground part.
[0015] In a possible implementation manner of the first aspect, extracting a first feature corresponding to a foreground portion of the first image from the third feature includes:
[0016] Performing feature extraction processing at multiple scales on the first image to obtain fourth features corresponding to the multiple scales;
[0017] Fusing the fourth features corresponding to the multiple different scales to obtain a first fused feature;
[0018] Performing segmentation processing on the first fused feature to obtain a fifth feature corresponding to a foreground portion of the first image;
[0019] The first feature is extracted from the third feature according to the fifth feature.
[0020] In the above embodiment, the feature segmentation process is performed using the fusion features of feature information of different scales. Since feature information of different scales can respectively express different image details, the above method can improve the feature segmentation accuracy, thereby facilitating the extraction of accurate foreground image features.
[0021] In a possible implementation manner of the first aspect, extracting the first feature from the third feature according to the fifth feature includes:
[0022] Performing scale conversion on the fifth feature to obtain a sixth feature; wherein the scale of the sixth feature is consistent with the scale of the third feature;
[0023] The third feature is filtered according to the sixth feature to obtain the first feature.
[0024] In the embodiment of the present application, the fifth feature is first converted to the same scale as the third feature, and then feature filtering is performed, which helps to improve the accuracy of feature filtering.
[0025] In a possible implementation manner of the first aspect, mapping the first feature to a preset coordinate system to obtain the second feature includes:
[0026] predicting depth information corresponding to the first feature;
[0027] The first feature is mapped to the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
[0028] In the above manner, using the trained prediction model to predict depth information is not only beneficial to improving the prediction efficiency, but also can ensure the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0029] In a possible implementation manner of the first aspect, predicting the depth information corresponding to the first feature includes:
[0030] Inputting the first feature into the trained prediction model and outputting a first prediction vector; wherein the prediction vector includes a plurality of elements, different elements correspond to different distance values, and each element represents a prediction probability of the distance value corresponding to the element;
[0031] Determine depth information corresponding to the first feature according to the first prediction vector.
[0032] By dividing the distance range into multiple distance intervals for prediction in the above manner, it is helpful to improve the prediction accuracy. In addition, using the trained prediction model to predict the depth information is not only conducive to improving the prediction efficiency, but also can ensure the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0033] In a possible implementation manner of the first aspect, the method further includes:
[0034] Acquire multiple groups of sample data; wherein each group of sample data includes feature information of a sample image and a distance value corresponding to each pixel point;
[0035] Inputting the sample data into the prediction model and outputting a second prediction vector;
[0036] Calculating the loss value of the prediction model according to the second prediction vector and the reference vector; wherein the reference vector is generated according to the distance value corresponding to each pixel point in the sample data;
[0037] If the loss value is greater than a preset threshold, the current prediction model is determined as the trained prediction model;
[0038] If the loss value is less than or equal to a preset threshold, updating the model parameters of the prediction model according to the loss value to obtain the updated prediction model;
[0039] Continue to train the updated prediction model until the trained prediction model is obtained.
[0040] In a possible implementation manner of the first aspect, detecting the target object in the first image according to the second feature includes:
[0041] Acquire a second image; wherein the second image is a point cloud image corresponding to the first image;
[0042] performing feature extraction processing on the second image to obtain a seventh feature;
[0043] Mapping the seventh feature to the preset coordinate system to obtain an eighth feature;
[0044] Performing fusion processing according to the second feature and the eighth feature to obtain a second fusion feature;
[0045] Detect a target object in the first image according to the second fusion feature.
[0046] In an embodiment of the present application, the point cloud data and the RGB image are mapped to the same preset coordinate system, and feature fusion is performed, which is equivalent to combining the pixel features of the RGB image and the depth information of the point cloud data. In this way, it helps to improve the detection accuracy.
[0047] In a second aspect, an embodiment of the present application provides an image detection device, including:
[0048] An acquisition unit, used for acquiring a first image to be processed; wherein the first image is an RGB image;
[0049] an extraction unit, configured to extract a first feature of a foreground portion of the first image by performing feature segmentation processing on feature information of the first image;
[0050] A mapping unit, used to map the first feature into a preset coordinate system to obtain a second feature; wherein the preset coordinate system is a coordinate system corresponding to the image in a top-down perspective;
[0051] A detection unit is used to detect a target object in the first image according to the second feature.
[0052] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, an image detection method as described in any one of the first aspects above is implemented.
[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image detection method as described in any one of the first aspects above is implemented.
[0054] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes the image detection method described in any one of the above-mentioned first aspects.
[0055] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0057] Figure 1 It is a flowchart of the image detection method provided in the embodiment of the present application;
[0058] Figure 2 It is a schematic diagram of the principle of the prediction model provided in the embodiment of the present application;
[0059] Figure 3 is a schematic diagram of an image detection process provided in an embodiment of the present application;
[0060] Figure 4 is a structural block diagram of an image detection device provided in an embodiment of the present application;
[0061] Figure 5 It is a schematic diagram of the structure of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0063] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0064] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0065] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0066] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0067] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0068] With the development of image processing technology, its application scope is becoming more and more extensive. For example, image processing technology can be applied to the field of autonomous driving, by performing image detection on the road ahead of the vehicle to perceive obstacles, thereby achieving automatic control of the vehicle.
[0069] Current image detection methods usually detect the entire image, which requires processing a lot of redundant information, affecting the efficiency and effect of image detection.
[0070] Based on this, an embodiment of the present application provides an image detection method. In the embodiment of the present application, feature information of the foreground part of the image is extracted, and then the feature information of the foreground part is used for detection, which reduces the redundant information of the background part of the image, effectively improves the detection efficiency, and greatly improves the detection effect.
[0071] See also Figure 1 , is a flow chart of an image detection method provided in an embodiment of the present application. As an example but not a limitation, the method may include the following steps:
[0072] S101. Obtain a first image to be processed.
[0073] Wherein, the first image is an RGB image.
[0074] Taking the application scenario of autonomous driving as an example, a captured image of the road in front of the vehicle can be collected by a camera in front of the vehicle, and this captured image serves as the first image to be processed.
[0075] In this application scenario, the image detection method of the embodiments of the present application can be executed by the vehicle's central controller. Specifically, the central controller interacts with the camera to obtain the captured image collected by the camera in real time, denoted as the first image, and then performs image detection processing on the first image through the image detection method of the embodiments of the present application to detect obstacles in the image. Finally, according to the conversion relationship between the image position of the obstacle in the image and the coordinate system of the actual road, the actual position of the obstacle on the actual road is determined, and the vehicle driving is controlled according to this actual position.
[0076] S102. Extract the first feature of the foreground part in the first image by performing feature segmentation processing on the feature information of the first image.
[0077] In one embodiment, S102 may include:
[0078] Perform feature extraction processing on the first image to obtain a third feature;
[0079] Extract the first feature corresponding to the foreground part of the first image from the third feature.
[0080] In the embodiments of the present application, extracting the foreground part according to the feature information of the image can improve the detection accuracy of the foreground part.
[0081] Optionally, the first image can be input into a trained feature extraction network to output a third feature. Wherein, the process of training the feature extraction network may include: inputting a sample image into a detection model to output a detection result; wherein, the detection model includes a feature extraction network and a detection head; calculating the loss value of the detection model according to the true label of the sample image and the detection result; if the loss value is less than a preset value, then determine the feature extraction network of the current detection model as the trained feature extraction network; if the loss value is greater than or equal to the preset value, then update the model parameters of the detection model according to the loss value and continue to train the detection model until the loss value of the detection model is less than the preset value.
[0082] In one implementation manner, the extraction method of the first feature includes:
[0083] Perform feature extraction processing on the first image at multiple different scales to obtain fourth features corresponding to each of the multiple different scales;
[0084] The fourth features corresponding to the multiple different scales are fused to obtain a first fused feature;
[0085] Performing segmentation processing on the first fused feature to obtain a fifth feature corresponding to the foreground portion of the first image;
[0086] The first feature is extracted from the third feature according to the fifth feature.
[0087] Optionally, different parameters may be set for the feature extraction network, each parameter corresponding to a scale, and then the first image is input into the feature extraction networks corresponding to different parameters respectively to obtain fourth features corresponding to multiple different scales.
[0088] For example, the first image is input into a feature extraction network with a scale of 8×, and the feature information of 8× is output, indicating that the feature information is 1 / 8 of the size of the first image. The first image is input into a feature extraction network with a scale of 4×, and the feature information of 4× is output, indicating that the feature information is 1 / 4 of the size of the first image. The first image is input into a feature extraction network with a scale of 16×, and the feature information of 16× is output, indicating that the feature information is 1 / 16 of the size of the first image. In this way, feature information of three scales of 4×, 8×, and 16× can be obtained.
[0089] Optionally, the fourth features of different scales may be upsampled or downsampled, transformed into feature information of the same scale, and fused.
[0090] For example, the 4× fourth feature is upsampled to obtain 8× feature information; the 16× fourth feature is downsampled to obtain 8× feature information; and then the obtained 8× feature information is fused to obtain the first fused feature. The fusion method may include superposition or concatenation.
[0091] Optionally, the fourth features of different scales may be fused through a feature fusion network. For example, a feature pyramid network (FPN) may be used to input the fourth features corresponding to different scales into the FPN network and output the first fused feature.
[0092] Optionally, the first fusion feature can be input into a trained segmentation network to output a fifth feature. The segmentation network can adopt a neural network model or an algorithm model with image segmentation capability, which is not specifically limited in the embodiment of the present application.
[0093] It is understandable that in practical applications, a feature segmentation model can be used to obtain the fifth feature of the first image. The feature segmentation model may include a feature extraction network, a feature fusion network, and a segmentation network. The feature extraction network is used to perform feature extraction processing on the first image at multiple different scales to obtain fourth features corresponding to multiple different scales; the feature fusion network is used to fuse the fourth features corresponding to multiple different scales to obtain a first fused feature; the segmentation network is used to segment the first fused feature to obtain the fifth feature corresponding to the foreground part of the first image. The first image is input into the segmentation model, and the fifth feature is output through the feature extraction network, the feature fusion network, and the segmentation network in sequence.
[0094] In one implementation, the method of extracting the first feature according to the fifth feature includes:
[0095] Performing scale conversion on the fifth feature to obtain a sixth feature; wherein the scale of the sixth feature is consistent with the scale of the third feature;
[0096] The third feature is feature filtered according to the sixth feature to obtain the first feature.
[0097] As described in the above embodiment, during the feature fusion process, the fourth feature may be upsampled or downsampled, resulting in the scale of the fifth feature being inconsistent with the third feature. In the embodiment of the present application, the fifth feature is first converted to the same scale as the third feature, and then feature filtering is performed, which helps to improve the accuracy of feature filtering.
[0098] Optionally, the scale conversion method may be upsampling or downsampling.
[0099] Optionally, the feature filtering method may be: setting the features of the third feature that are in different positions from the fifth feature to 0, and retaining the features of the third feature that are in the same position as the fifth feature. It is understood that the same position here means the same pixel coordinates.
[0100] In the above embodiment, the feature segmentation process is performed using the fusion features of feature information of different scales. Since feature information of different scales can respectively express different image details, the above method can improve the feature segmentation accuracy, thereby facilitating the extraction of accurate foreground image features.
[0101] In another embodiment, the method of extracting the first feature may include:
[0102] Performing feature extraction processing on the first image to obtain a third feature;
[0103] Segment the first image according to the third feature to obtain an object detection frame of a foreground portion of the first image;
[0104] The first feature is extracted from the third feature according to the target detection box.
[0105] In one implementation, target detection can be obtained through a segmentation model. The segmentation model includes a feature extraction network and a segmentation network. The feature extraction network is used to perform feature extraction processing on the first image, and the segmentation network is used to perform segmentation processing on the first image according to the third feature to obtain a target detection frame of the foreground part of the first image. Specifically, the first image is input into the segmentation model, and the target detection frame is output.
[0106] In another implementation, the segmentation process may include:
[0107] Randomly generate multiple candidate detection boxes in the first image;
[0108] Detecting the image category to which the image in each candidate detection frame belongs according to the third feature in each candidate detection frame; wherein the image category includes a first category and a second category, the first category represents the foreground, and the second category represents the background;
[0109] Multiple candidate detection frames are filtered according to the image category corresponding to each candidate detection frame to obtain the target detection frame.
[0110] The randomly generated candidate detection boxes may have different sizes and random positions.
[0111] Optionally, the third feature in each candidate detection frame can be input into a trained classification model to output the category to which the image in each candidate detection frame belongs. The classification model can adopt a neural network model or an algorithm model with classification capabilities (such as a clustering model, a binary classification model, etc.).
[0112] For example, the classification model can use the first stage network in the Region Proposal Network (RPN). Among them, the working principle of RPN includes two stages: in the first stage, it is determined whether there is an object in each random detection frame based on the input feature information, and then the random detection frames where no objects exist are filtered out. In the second stage, the category of the objects in the remaining random detection frames is determined, and the target detection frame is output. By applying the first stage network in RPN to the embodiment of the present application, it is possible to determine whether the image in the candidate detection frame is the foreground or the background.
[0113] In one implementation, a method of filtering candidate detection frames includes:
[0114] Obtain a plurality of candidate detection frames whose image categories are the first category from the candidate detection frames to obtain a first detection frame;
[0115] Perform deduplication processing on the first detection box to obtain the deduplicated target detection box.
[0116] Optionally, delete the candidate detection boxes with the image category of the second category and retain the candidate detection boxes with the image category of the first category.
[0117] Since there may be duplicate candidate detection boxes randomly generated, optionally, in some implementation manners, deduplication processing can be performed on the target detection box. Specifically, calculate the intersection over union (IoU) between every two candidate detection boxes; if the IoU between two candidate detection boxes is greater than a preset value, delete any one of the two candidate detection boxes.
[0118] In the embodiments of the present application, the image is split into multiple local regions by using randomly generated candidate detection boxes, and then foreground detection is performed on each local region separately. This method has high calculation efficiency, helps to quickly detect the foreground part in the image, and thus is beneficial to improving the efficiency of image detection.
[0119] S103, map the first feature to a preset coordinate system to obtain a second feature.
[0120] Among them, the preset coordinate system is the coordinate system corresponding to the image in the top-down view.
[0121] For example, in the embodiments of the present application, the preset coordinate system can adopt the coordinate system of a birds-eye view (BEV). The BEV space is a 3D perception method that converts the traditional 2D image view of autonomous driving into a birds-eye view. Through algorithm correction and change, the BEV space can convert the 2D image captured by the camera into a top-down view based on the top-down view, so as to achieve 3D perception. This method has an important impact on the cognition and structure of the environment in autonomous driving and can improve the perception and decision-making ability of the autonomous driving system.
[0122] However, since the first image is a two-dimensional planar image, to convert it to the preset coordinate system, the depth information of the pixel points needs to be obtained. To solve this problem, in one embodiment, S103 may include:
[0123] Predict the depth information corresponding to the first feature;
[0124] According to the pixel coordinates and depth information corresponding to the first feature, map the first feature to the preset coordinate system.
[0125] In one implementation manner, the prediction method of the depth information includes:
[0126] Input the first feature into the trained prediction model to output a first prediction vector; wherein, the prediction vector includes multiple elements, different elements correspond to different distance values, and each element represents the prediction probability of the distance value corresponding to the element.
[0127] Determine the depth information corresponding to the first feature according to the first prediction vector.
[0128] Specifically, the working principle of the prediction model is as follows: First, construct a frustum map using image features, and divide a certain distance range into multiple intervals at a preset interval. For example, divide the range from 1 to 60 meters into 118 intervals at intervals of 0.5 meters each. Correspondingly, the prediction model will predict the probability that the input feature belongs to each interval and output a prediction vector. For example, the prediction vector [1, 0, 0..., 0] indicates that the probability that the input feature belongs to the distance interval of 0 - 0.5 meters is 1; the prediction vector [0, 1, 0..., 0] indicates that the probability that the input feature belongs to the distance interval of 0.5 - 1 meter is 1.
[0129] Exemplarily, refer to Figure 2 which is a schematic diagram of the principle of the prediction model provided by an embodiment of the present application. As Figure 2 shown, Image Features F(u, v) represents the features of the pixel at the u-th row and v-th column in the image. Depth Distributions D(u, v) represents the feature distribution corresponding to F(u, v). As Figure 2 shown, in the feature distribution, the distance range is divided into D distance intervals, and the predicted values corresponding to different distance intervals are different. Frustum Features G(u, v) represents the frustum map corresponding to F(u, v).
[0130] Optionally, the depth information of the first feature can be determined according to the distance interval corresponding to the element with the largest value in the first prediction vector. For example, if the element with the largest value 0.8 in the prediction vector [0.8, 0.2, 0..., 0] corresponds to the distance interval of 0 - 0.5, then the depth information of the first feature is determined to be 0.5 or 0.25.
[0131] Optionally, the depth information of the first feature can be determined according to the distance intervals corresponding to the non-zero elements in the first prediction vector. For example, in the prediction vector [0.8, 0.2, 0..., 0], the non-zero element 0.8 corresponds to the distance interval of 0 - 0.5, and 0.2 corresponds to the distance interval of 0.5 - 1. The depth information of the first feature can be determined to be 0.8×0.5 + 0.2×1 = 0.6 meters according to the weighted data.
[0132] It can be understood that the first feature corresponding to each pixel point in the image is traversed, and the depth information of the first feature corresponding to each pixel point is predicted in the above manner.
[0133] Optionally, the training process of the prediction model may include:
[0134] Acquire multiple groups of sample data; wherein each group of sample data includes feature information of a sample image and a distance value corresponding to each pixel point;
[0135] Inputting sample data into the prediction model and outputting a second prediction vector;
[0136] Calculating the loss value of the prediction model according to the second prediction vector and the reference vector; wherein the reference vector is generated according to the distance value corresponding to each pixel point in the sample data;
[0137] If the loss value is greater than a preset threshold, the current prediction model is determined as the trained prediction model;
[0138] If the loss value is less than or equal to the preset threshold, the model parameters of the prediction model are updated according to the loss value to obtain an updated prediction model;
[0139] Continue to train the updated prediction model until a trained prediction model is obtained.
[0140] By dividing the distance range into multiple distance intervals for prediction in the above manner, it is helpful to improve the prediction accuracy. In addition, using the trained prediction model to predict the depth information is not only conducive to improving the prediction efficiency, but also can ensure the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0141] S104: Detect a target object in the first image according to the second feature.
[0142] In one embodiment, S104 may include:
[0143] Acquire a second image; wherein the second image is a point cloud image corresponding to the first image;
[0144] Performing feature extraction processing on the second image to obtain a seventh feature;
[0145] Mapping the seventh feature to the preset coordinate system to obtain an eighth feature;
[0146] Perform fusion processing according to the second feature and the eighth feature to obtain a second fusion feature;
[0147] The target object in the first image is detected according to the second fused feature.
[0148] Optionally, the manner of mapping the seventh feature to a preset coordinate system may include: performing dimensionality reduction processing on the seventh feature to obtain an eighth feature having a dimension conforming to the preset coordinate system.
[0149] For example, the seventh feature (i.e., the feature of the point cloud image) usually includes 5 dimensions, such as 1 (number of clusters) × 128 (number of channels) × 180 (length of the BEV space) × 180 (width of the BEV space) × 2 (height). Multiply the height information by the number of channels to get 1 × 256 × 180 × 180, which is reduced to 4 dimensions, realizing the mapping to the BEV space.
[0150] In the embodiments of the present application, the point cloud data and the RGB image are mapped to the same preset coordinate system and feature fusion is performed, which is equivalent to combining the pixel features of the RGB image and the depth information of the point cloud data. In this way, it helps to improve the detection accuracy.
[0151] Exemplarily, refer to Figure 3 , which is a schematic diagram of the image detection process provided by the embodiments of the present application. By way of example and not limitation, as Figure 3 shown, in the upper processing flow, first perform feature segmentation processing on the feature information of the first image to extract the first feature of the foreground part in the first image; then map the first feature to the BEV space to obtain the second feature. In the lower processing flow, first extract the features in the second image to obtain the seventh feature; then map the seventh feature to the BEV space to obtain the eighth feature. Then perform fusion processing on the second feature and the eighth feature to obtain the second fusion feature; finally, detect the target object in the first image according to the second fusion feature.
[0152] In the embodiments of the present application, the feature information of the foreground part in the image is extracted, and then the foreground part's feature information is used for detection, reducing the redundant information of the background part in the image. This not only effectively improves the detection efficiency but also greatly enhances the detection effect. In addition, mapping the pixel features of the RGB image and the depth features of the point cloud image to the same coordinate system and performing fusion, and using the fusion features for image detection can combine various types of feature information, thus helping to improve the detection accuracy.
[0153] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0154] Corresponding to the image detection method described in the above embodiments, Figure 4 is a structural block diagram of the image detection device provided by the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0155] Referring to Figure 4 , the device includes:
[0156] An acquisition unit 41, configured to acquire a first image to be processed; wherein, the first image is an RGB image.
[0157] The extraction unit 42 is used to extract the first feature of the foreground part of the first image by performing feature segmentation processing on the feature information of the first image.
[0158] The mapping unit 43 is used to map the first feature into a preset coordinate system to obtain a second feature; wherein the preset coordinate system is a coordinate system corresponding to the image in a top-down perspective.
[0159] The detection unit 44 is configured to detect a target object in the first image according to the second feature.
[0160] Optionally, the extraction unit 42 is further configured to:
[0161] Performing feature extraction processing on the first image to obtain a third feature;
[0162] A first feature corresponding to the foreground portion of the first image is extracted from the third feature.
[0163] Optionally, the extraction unit 42 is further configured to:
[0164] Performing feature extraction processing at multiple scales on the first image to obtain fourth features corresponding to the multiple scales;
[0165] Fusing the fourth features corresponding to the multiple different scales to obtain a first fused feature;
[0166] Performing segmentation processing on the first fused feature to obtain a fifth feature corresponding to a foreground portion of the first image;
[0167] The first feature is extracted from the third feature according to the fifth feature.
[0168] Optionally, the extraction unit 42 is further configured to:
[0169] Performing scale conversion on the fifth feature to obtain a sixth feature; wherein the scale of the sixth feature is consistent with the scale of the third feature;
[0170] The third feature is filtered according to the sixth feature to obtain the first feature.
[0171] Optionally, the mapping unit 43 is further configured to:
[0172] Predicting depth information corresponding to the first feature;
[0173] The first feature is mapped to the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
[0174] Optionally, the mapping unit 43 is further configured to:
[0175] Inputting the first feature into the trained prediction model and outputting a first prediction vector; wherein the prediction vector includes a plurality of elements, different elements correspond to different distance values, and each element represents a prediction probability of the distance value corresponding to the element;
[0176] Determine depth information corresponding to the first feature according to the first prediction vector.
[0177] Optionally, the mapping unit 43 is further configured to:
[0178] Acquire multiple groups of sample data; wherein each group of sample data includes feature information of a sample image and a distance value corresponding to each pixel point;
[0179] Inputting the sample data into the prediction model and outputting a second prediction vector;
[0180] Calculating the loss value of the prediction model according to the second prediction vector and the reference vector; wherein the reference vector is generated according to the distance value corresponding to each pixel point in the sample data;
[0181] If the loss value is greater than a preset threshold, the current prediction model is determined as the trained prediction model;
[0182] If the loss value is less than or equal to a preset threshold, updating the model parameters of the prediction model according to the loss value to obtain the updated prediction model;
[0183] Continue to train the updated prediction model until the trained prediction model is obtained.
[0184] Optionally, the detection unit 44 is further used for:
[0185] Acquire a second image; wherein the second image is a point cloud image corresponding to the first image;
[0186] performing feature extraction processing on the second image to obtain a seventh feature;
[0187] Mapping the seventh feature to the preset coordinate system to obtain an eighth feature;
[0188] Performing fusion processing according to the second feature and the eighth feature to obtain a second fusion feature;
[0189] Detect a target object in the first image according to the second fusion feature.
[0190] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0191] in addition, Figure 4 The image detection device shown may be a software unit, a hardware unit, or a combination of software and hardware, which is built into an existing terminal device, or may be integrated into the terminal device as an independent accessory, or may exist as an independent terminal device.
[0192] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0193] Figure 5 Schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 5 As shown, the terminal device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one is shown in the figure) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, and when the processor 50 executes the computer program 52, the steps in any of the above-mentioned image detection method embodiments are implemented.
[0194] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 5 It is only an example of the terminal device 5 and does not constitute a limitation on the terminal device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0195] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0196] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as a hard disk or memory of the terminal device 5. In other embodiments, the memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the terminal device 5. The memory 51 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or is to be output.
[0197] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0198] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0199] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0200] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0201] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0202] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0203] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An image detection method, characterized in that: include: Acquire a first image to be processed; wherein the first image is an RGB image; Extracting a first feature of a foreground portion of the first image by performing feature segmentation processing on feature information of the first image; Mapping the first feature to a preset coordinate system to obtain a second feature; wherein the preset coordinate system is a coordinate system corresponding to the image in a top-down perspective; A target object in the first image is detected according to the second feature.
2. The image detection method according to claim 1, characterized in that: The extracting a first feature of a foreground portion of the first image by performing feature segmentation processing on the feature information of the first image includes: Performing feature extraction processing on the first image to obtain a third feature; A first feature corresponding to the foreground portion of the first image is extracted from the third feature.
3. The image detection method according to claim 2, characterized in that: The step of extracting a first feature corresponding to a foreground portion of the first image from the third feature includes: Performing feature extraction processing at multiple scales on the first image to obtain fourth features corresponding to the multiple scales; Fusing the fourth features corresponding to the multiple different scales to obtain a first fused feature; Performing segmentation processing on the first fused feature to obtain a fifth feature corresponding to a foreground portion of the first image; The first feature is extracted from the third feature according to the fifth feature.
4. The image detection method according to claim 3, characterized in that: The extracting the first feature from the third feature according to the fifth feature includes: Performing scale conversion on the fifth feature to obtain a sixth feature; wherein the scale of the sixth feature is consistent with the scale of the third feature; The third feature is filtered according to the sixth feature to obtain the first feature.
5. The image detection method according to any one of claims 1 to 4, characterized in that: Mapping the first feature to a preset coordinate system to obtain a second feature includes: Predicting depth information corresponding to the first feature; The first feature is mapped to the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
6. The image detection method according to claim 5, characterized in that: The predicting the depth information corresponding to the first feature includes: Inputting the first feature into the trained prediction model and outputting a first prediction vector; wherein the prediction vector includes a plurality of elements, different elements correspond to different distance values, and each element represents a prediction probability of the distance value corresponding to the element; Determine depth information corresponding to the first feature according to the first prediction vector.
7. The image detection method according to claim 6, characterized in that: The method further comprises: Acquire multiple groups of sample data; wherein each group of sample data includes feature information of a sample image and a distance value corresponding to each pixel point; Inputting the sample data into the prediction model and outputting a second prediction vector; Calculating the loss value of the prediction model according to the second prediction vector and the reference vector; wherein the reference vector is generated according to the distance value corresponding to each pixel point in the sample data; If the loss value is greater than a preset threshold, the current prediction model is determined as the trained prediction model; If the loss value is less than or equal to a preset threshold, updating the model parameters of the prediction model according to the loss value to obtain the updated prediction model; Continue to train the updated prediction model until the trained prediction model is obtained.
8. The image detection method according to any one of claims 1 to 6, characterized in that: The detecting the target object in the first image according to the second feature comprises: Acquire a second image; wherein the second image is a point cloud image corresponding to the first image; performing feature extraction processing on the second image to obtain a seventh feature; Mapping the seventh feature to the preset coordinate system to obtain an eighth feature; Performing fusion processing according to the second feature and the eighth feature to obtain a second fusion feature; Detect a target object in the first image according to the second fusion feature.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.