Image detection method, terminal equipment and computer readable storage medium
By segmenting the image, extracting and mapping the foreground features to the coordinate system under the top view angle, the problem of redundant information processing affecting detection efficiency in the prior art is solved, and efficient and accurate image detection is achieved.
Patent Information
- Application Number
- CN202411994408.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
Existing image detection methods need to process a large amount of redundant information, which affects the detection efficiency and effect.
By segmenting the image, the features of the foreground part are extracted and mapped into the coordinate system at the top view angle, reducing background redundant information and improving detection accuracy.
It effectively improves the efficiency and effect of image detection, reduces redundant information processing, and improves detection accuracy.
Smart Images

Figure CN120047663A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to an image detection method, a terminal device, and a computer-readable storage medium. Background Art
[0002] With the development of image processing technology, its application scope is becoming more and more extensive. For example, image processing technology can be applied to the field of autonomous driving. By performing image detection on the captured image of the road ahead of the vehicle, obstacles can be perceived, thereby realizing automatic control of the vehicle.
[0003] The current image detection methods usually perform overall detection on the image, and need to process a large amount of redundant information, which affects the efficiency and detection effect of image detection. Summary of the Invention
[0004] The embodiments of this application provide an image detection method, a terminal device, and a computer-readable storage medium, which can effectively improve the efficiency and detection effect of image detection.
[0005] In a first aspect, the embodiments of this application provide an image detection method, including:
[0006] Obtain a first image to be processed; wherein, the first image is an RGB image;
[0007] Through performing image segmentation processing on the first image, extract a first feature of the foreground part in the first image;
[0008] Map the first feature to a preset coordinate system to obtain a second feature; wherein, the preset coordinate system is the coordinate system corresponding to the image under a top-down perspective;
[0009] Detect a target object in the first image according to the second feature.
[0010] In the embodiments of this application, the feature information of the foreground part in the image is extracted, and then the feature information of the foreground part is used for detection, reducing the redundant information of the background part in the image. This not only effectively improves the detection efficiency, but also greatly improves the detection effect. In addition, since the image under a top-down perspective contains distance information, mapping the pixel features of the RGB image to the coordinate system under a top-down perspective makes the second feature contain distance information. Detecting according to the second feature helps to improve the detection accuracy.
[0011] In a possible implementation manner of the first aspect, the step of through performing image segmentation processing on the first image, extract a first feature of the foreground part in the first image, includes:
[0012] Perform feature extraction processing on the first image to obtain a third feature;
[0013] Segment the first image according to the third feature to obtain the target detection box of the foreground part in the first image;
[0014] Extract the first feature from the third feature according to the target detection box.
[0015] In a possible implementation manner of the first aspect, the segmenting the first image according to the third feature to obtain the target detection box of the foreground part in the first image includes:
[0016] Randomly generate a plurality of candidate detection boxes in the first image;
[0017] Detect the image category to which the image in each candidate detection box belongs according to the third feature in each candidate detection box; wherein, the image category includes a first category and a second category, the first category represents the foreground, and the second category represents the background;
[0018] Filter the plurality of candidate detection boxes according to the image category corresponding to each candidate detection box to obtain the target detection box.
[0019] In the embodiments of the present application, the candidate detection boxes generated randomly are used to split the image into multiple local regions, and then the foreground of each local region is detected respectively. This method has high computational efficiency, helps to quickly detect the foreground part in the image, and thus is beneficial to improving the efficiency of image detection.
[0020] In a possible implementation manner of the first aspect, the filtering the plurality of candidate detection boxes according to the image category corresponding to each candidate detection box to obtain the target detection box includes:
[0021] Obtain the candidate detection boxes with the image category of the first category among the plurality of candidate detection boxes to obtain the first detection box;
[0022] Perform duplicate removal processing on the first detection box to obtain the target detection box after duplicate removal.
[0023] Through the above method, duplicate detection boxes can be removed, which helps to improve the detection effect and at the same time reduces the redundancy of subsequent data processing.
[0024] In a possible implementation manner of the first aspect, the mapping the first feature to a preset coordinate system to obtain the second feature includes:
[0025] Predict the depth information corresponding to the first feature;
[0026] Map the first feature to the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
[0027] In the above manner, predicting depth information using the trained prediction model not only helps improve the prediction efficiency, but also ensures the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0028] In a possible implementation manner of the first aspect, predicting the depth information corresponding to the first feature includes:
[0029] Inputting the first feature into the trained prediction model to output a first prediction vector; wherein, the prediction vector includes multiple elements, different elements correspond to different distance values, and each element represents the prediction probability of the distance value corresponding to the element.
[0030] Determining the depth information corresponding to the first feature according to the first prediction vector.
[0031] In the above manner, dividing the distance range into multiple distance intervals for prediction helps improve the prediction accuracy. In addition, predicting depth information using the trained prediction model not only helps improve the prediction efficiency, but also ensures the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0032] In a possible implementation manner of the first aspect, the method further includes:
[0033] Obtaining multiple groups of sample data; wherein, each group of the sample data includes the feature information of a sample image and the distance value corresponding to each pixel point.
[0034] Inputting the sample data into the prediction model to output a second prediction vector.
[0035] Calculating the loss value of the prediction model according to the second prediction vector and a reference vector; wherein, the reference vector is generated according to the distance value corresponding to each pixel point in the sample data.
[0036] If the loss value is greater than a preset threshold, determining the current prediction model as the trained prediction model.
[0037] If the loss value is less than or equal to the preset threshold, updating the model parameters of the prediction model according to the loss value to obtain the updated prediction model.
[0038] Continuing to train the updated prediction model until the trained prediction model is obtained.
[0039] In a possible implementation manner of the first aspect, detecting the target object in the first image according to the second feature includes:
[0040] Obtain a second image; wherein, the second image is a point cloud image corresponding to the first image;
[0041] Perform feature extraction processing on the second image to obtain a seventh feature;
[0042] Map the seventh feature into the preset coordinate system to obtain an eighth feature;
[0043] Perform fusion processing according to the second feature and the eighth feature to obtain a second fusion feature;
[0044] Detect a target object in the first image according to the second fusion feature.
[0045] In the embodiments of the present application, mapping point cloud data and RGB images into the same preset coordinate system and performing feature fusion is equivalent to combining the pixel features of the RGB image and the depth information of the point cloud data. In this way, it helps to improve the detection accuracy.
[0046] In a second aspect, an image detection device provided by an embodiment of the present application includes:
[0047] An acquisition unit, configured to acquire a first image to be processed; wherein, the first image is an RGB image;
[0048] An extraction unit, configured to extract a first feature of the foreground part in the first image by performing feature segmentation processing on the feature information of the first image;
[0049] A mapping unit, configured to map the first feature into a preset coordinate system to obtain a second feature; wherein, the preset coordinate system is a coordinate system corresponding to an image in a top-down view;
[0050] A detection unit, configured to detect a target object in the first image according to the second feature.
[0051] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image detection method described in any item of the first aspect above is implemented.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image detection method described in any item of the first aspect above is implemented.
[0053] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute the image detection method described in any one of the above first aspects.
[0054] It can be understood that the beneficial effects of the above second aspect to fifth aspect can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 is a schematic flowchart of the image detection method provided by an embodiment of the present application;
[0057] Figure 2 is a schematic diagram of the principle of the prediction model provided by an embodiment of the present application;
[0058] Figure 3 is a schematic diagram of the image detection process provided by an embodiment of the present application;
[0059] Figure 4 is a structural block diagram of the image detection device provided by an embodiment of the present application;
[0060] Figure 5 is a schematic diagram of the structure of the terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0062] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0063] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0064] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0065] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0066] Referring to "one embodiment" or "some embodiments" described in the specification of this application means that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0067] With the development of image processing technology, its application scope is becoming more and more extensive. For example, image processing technology can be applied to the field of autonomous driving. By performing image detection on the captured images of the road ahead of the vehicle, obstacles can be perceived, thereby realizing automatic control of the vehicle.
[0068] Current image detection methods usually perform overall detection on the image, which requires processing a large amount of redundant information, affecting the efficiency and detection effect of image detection.
[0069] Based on this, the embodiments of this application provide an image detection method. In the embodiments of this application, the feature information of the foreground part in the image is extracted, and then the feature information of the foreground part is used for detection, reducing the redundant information of the background part in the image, not only effectively improving the detection efficiency, but also greatly enhancing the detection effect.
[0070] See Figure 1 , which is a schematic flow chart of the image detection method provided by the embodiments of this application. By way of example and not limitation, the method may include the following steps:
[0071] S101. Obtain a first image to be processed.
[0072] Wherein, the first image is an RGB image.
[0073] Taking the application scenario of autonomous driving as an example, a captured image of the road in front of the vehicle can be collected by a camera in front of the vehicle, and this captured image serves as the first image to be processed.
[0074] In this application scenario, the image detection method of the embodiments of the present application can be executed by the vehicle's central controller. Specifically, the central controller interacts with the camera to obtain the captured image collected by the camera in real time, denoted as the first image, and then performs image detection processing on the first image through the image detection method of the embodiments of the present application to detect obstacles in the image. Finally, according to the conversion relationship between the image position of the obstacle in the image and the coordinate system of the actual road, the actual position of the obstacle on the actual road is determined, and the vehicle driving is controlled according to this actual position.
[0075] S102. Extract the first feature of the foreground part in the first image by performing image segmentation processing on the first image.
[0076] In one embodiment, S102 may include:
[0077] Perform feature extraction processing on the first image to obtain a third feature;
[0078] Perform segmentation processing on the first image according to the third feature to obtain a target detection box of the foreground part in the first image;
[0079] Extract the first feature from the third feature according to the target detection box.
[0080] In one implementation, the target detection can be obtained through a segmentation model. Wherein, the segmentation model includes a feature extraction network and a segmentation network. The feature extraction network is used to perform feature extraction processing on the first image, and the segmentation network is used to perform segmentation processing on the first image according to the third feature to obtain a target detection box of the foreground part in the first image. Specifically, the first image is input into the segmentation model, and the target detection box is output.
[0081] In another implementation, the segmentation processing may include:
[0082] Randomly generate a plurality of candidate detection boxes in the first image;
[0083] Detect the image category to which the image in each candidate detection box belongs according to the third feature in each candidate detection box; wherein, the image categories include a first category and a second category, the first category represents the foreground, and the second category represents the background;
[0084] Filter multiple candidate detection boxes according to the image categories corresponding to each candidate detection box to obtain target detection boxes.
[0085] Among them, the sizes of randomly generated candidate detection boxes can be different and the positions can be random.
[0086] Optionally, the third feature in each candidate detection box can be input into a trained classification model to output the category to which the image in each candidate detection box belongs. Among them, the classification model can adopt a neural network model or an algorithm model with classification ability (such as a clustering model, a binary classification model, etc.).
[0087] For example, the classification model can adopt the network in the first stage of the Region Proposal Network (RPN). Among them, the working principle of RPN includes two stages: in the first stage, it is judged whether there is an object in each random detection box according to the input feature information, and then the random detection boxes without objects are filtered out. In the second stage, the category of the object in the remaining random detection boxes is judged, and the target detection box is output. Applying the network in the first stage of RPN to the embodiments of the present application can judge whether the image in the candidate detection box is a foreground or a background.
[0088] In one implementation manner, the method for filtering candidate detection boxes includes:
[0089] Obtain the candidate detection boxes with the image category being the first category among the multiple candidate detection boxes to obtain the first detection box;
[0090] Perform duplicate removal processing on the first detection box to obtain the target detection box after duplicate removal.
[0091] Optionally, delete the candidate detection boxes with the image category being the second category and retain the candidate detection boxes with the image category being the first category.
[0092] Since there may be duplicate candidate detection boxes randomly generated, optionally, in some implementation manners, duplicate removal processing can be performed on the target detection box. Specifically, calculate the intersection over union (IoU) between every two candidate detection boxes; if the IoU between two candidate detection boxes is greater than a preset value, then delete any one of the two candidate detection boxes.
[0093] In the embodiments of the present application, the image is split into multiple local regions by using randomly generated candidate detection boxes, and then foreground detection is performed on each local region respectively. This method has high computational efficiency, helps to quickly detect the foreground part in the image, and thus is beneficial to improving the efficiency of image detection.
[0094] In another embodiment, the method for extracting the first feature can include:
[0095] Perform feature extraction processing on the first image to obtain the third feature;
[0096] Extract the first feature corresponding to the foreground part of the first image from the third feature.
[0097] In the embodiments of the present application, extracting the foreground part according to the feature information of the image can improve the detection accuracy of the foreground part.
[0098] Optionally, the first image can be input into the trained feature extraction network to output the third feature. Among them, the process of training the feature extraction network can include: inputting the sample image into the detection model to output the detection result; where the detection model includes a feature extraction network and a detection head; calculating the loss value of the detection model according to the true label and the detection result of the sample image; if the loss value is less than the preset value, determining the feature extraction network of the current detection model as the trained feature extraction network; if the loss value is greater than or equal to the preset value, updating the model parameters of the detection model according to the loss value and continuing to train the detection model until the loss value of the detection model is less than the preset value.
[0099] In one implementation manner, the extraction method of the first feature includes:
[0100] Perform feature extraction processing on the first image at multiple different scales to obtain the fourth features corresponding to the respective different scales;
[0101] Fuse the fourth features corresponding to the respective different scales to obtain the first fused feature;
[0102] Perform segmentation processing on the first fused feature to obtain the fifth feature corresponding to the foreground part in the first image;
[0103] Extract the first feature from the third feature according to the fifth feature.
[0104] Optionally, different parameters can be set for the feature extraction network, each parameter corresponding to a scale, and then the first image is respectively input into the feature extraction networks corresponding to different parameters to obtain the fourth features corresponding to the respective different scales.
[0105] For example, input the first image into the feature extraction network with a scale of 8×, and output the feature information of 8×, indicating that the feature information is 1 / 8 of the size of the first image. Input the first image into the feature extraction network with a scale of 4×, and output the feature information of 4×, indicating that the feature information is 1 / 4 of the size of the first image. Input the first image into the feature extraction network with a scale of 16×, and output the feature information of 16×, indicating that the feature information is 1 / 16 of the size of the first image. In this way, the feature information of three scales of 4×, 8×, and 16× can be obtained.
[0106] Optionally, the fourth features of different scales can be upsampled or downsampled to transform into feature information of the same scale and then fused.
[0107] For example, the fourth feature of 4× is upsampled to obtain feature information of 8×; the fourth feature of 16× is downsampled to obtain feature information of 8×; then the obtained feature information of 8× is fused to obtain a first fused feature. Among them, the fusion method can include stacking or concatenation.
[0108] Optionally, different-scale fourth features can be fused through a feature fusion network. For example, a Feature Pyramid Networks (FPN) network can be adopted, and the fourth features corresponding to different scales are input into the FPN network to output a first fused feature.
[0109] Optionally, the first fused feature can be input into a trained segmentation network to output a fifth feature. The segmentation network can adopt a neural network model or an algorithm model with image segmentation ability, and the embodiments of the present application do not make specific limitations thereto.
[0110] It can be understood that in practical applications, a feature segmentation model can be used to obtain the fifth feature of the first image. The feature segmentation model can include a feature extraction network, a feature fusion network, and a segmentation network. Among them, the feature extraction network is used to perform feature extraction processing on the first image at multiple different scales to obtain the fourth features corresponding to different scales respectively; the feature fusion network is used to fuse the fourth features corresponding to different scales respectively to obtain a first fused feature; the segmentation network is used to perform segmentation processing on the first fused feature to obtain the fifth feature corresponding to the foreground part in the first image. The first image is input into the segmentation model, and passes through the feature extraction network, the feature fusion network, and the segmentation network in sequence to output the fifth feature.
[0111] In one implementation manner, the method for extracting the first feature according to the fifth feature includes:
[0112] Performing scale conversion on the fifth feature to obtain a sixth feature; wherein the scale of the sixth feature is the same as that of the third feature;
[0113] Performing feature filtering on the third feature according to the sixth feature to obtain the first feature.
[0114] As described in the above embodiments, in the process of feature fusion processing, the fourth feature may be upsampled or downsampled, resulting in the scale of the fifth feature being inconsistent with that of the third feature. In the embodiments of the present application, the fifth feature is first converted to the same scale as the third feature and then feature filtering is performed, which helps to improve the accuracy of feature filtering.
[0115] Optionally, the scaling conversion method can be upsampling or downsampling.
[0116] Optionally, the feature filtering method can be: setting the features in the third feature that are different in position from the fifth feature to 0, and retaining the features in the third feature that are the same in position as the fifth feature. It can be understood that the same position here refers to the same pixel coordinates.
[0117] In the above embodiments, the fusion feature of feature information of different scales is used for feature segmentation processing. Since the feature information of different scales can respectively express different image details, the above method can improve the segmentation accuracy of features, thereby helping to extract the features of the accurate foreground image.
[0118] S103. Map the first feature to a preset coordinate system to obtain a second feature.
[0119] Wherein, the preset coordinate system is the coordinate system corresponding to the image in the top-down view.
[0120] For example, in the embodiments of the present application, the preset coordinate system can adopt the coordinate system of a birds-eye view (BEV). The BEV space is a 3D perception method that converts the perspective of a traditional autonomous driving 2D image into a birds-eye view. Through algorithm correction and change, the BEV space can convert the 2D image captured by the camera into a top-down view based on the top-down perspective, thereby realizing 3D perception. This method has an important impact on the perception and structure of the environment in autonomous driving and can improve the perception and decision-making capabilities of the autonomous driving system.
[0121] However, since the first image is a two-dimensional planar image, to convert it to the preset coordinate system, the depth information of the pixel points needs to be obtained. To solve this problem, in one embodiment, S103 may include:
[0122] Predict the depth information corresponding to the first feature;
[0123] Map the first feature to the preset coordinate system according to the pixel coordinates and depth information corresponding to the first feature.
[0124] In one implementation, the method for predicting depth information includes:
[0125] Input the first feature into a trained prediction model to output a first prediction vector; wherein, the prediction vector includes multiple elements, different elements correspond to different distance values, and each element represents the prediction probability of the distance value corresponding to the element;
[0126] Determine the depth information corresponding to the first feature according to the first prediction vector.
[0127] Specifically, the working principle of the prediction model includes: First, a frustum map is constructed using image features, and a certain distance range is divided into multiple intervals at a preset interval. For example, the range from 1 to 60 meters is divided into 118 intervals at intervals of 0.5 meters each. Correspondingly, the prediction model predicts the probability that the input features belong to each interval and outputs a prediction vector. For example, the prediction vector [1, 0, 0..., 0] indicates that the probability that the input features belong to the distance interval of 0 - 0.5 meters is 1; the prediction vector [0, 1, 0..., 0] indicates that the probability that the input features belong to the distance interval of 0.5 - 1 meter is 1.
[0128] Exemplarily, referring to Figure 2 , which is a schematic diagram of the principle of the prediction model provided by the embodiments of the present application. As Figure 2 shown, Image Features F(u, v) represents the features of the pixel at the u-th row and v-th column in the image. Depth Distributions D(u, v) represents the feature distribution corresponding to F(u, v). As Figure 2 shown, in the feature distribution, the distance range is divided into D distance intervals, and the predicted values corresponding to different distance intervals are different. Frustum Features G(u, v) represents the frustum map corresponding to F(u, v).
[0129] Optionally, the depth information of the first feature can be determined according to the distance interval corresponding to the element with the largest value in the first prediction vector. For example, in the prediction vector [0.8, 0.2, 0..., 0], the element 0.8 with the largest value corresponds to the distance interval of 0 - 0.5, then the depth information of the first feature is determined to be 0.5 or 0.25.
[0130] Optionally, the depth information of the first feature can be determined according to the distance intervals corresponding to the non-zero elements in the first prediction vector. For example, in the prediction vector [0.8, 0.2, 0..., 0], among the non-zero elements, 0.8 corresponds to the distance interval of 0 - 0.5, and 0.2 corresponds to the distance interval of 0.5 - 1. The depth information of the first feature can be determined to be 0.8×0.5 + 0.2×1 = 0.6 meters according to the weighted data.
[0131] It can be understood that each first feature corresponding to each pixel point in the image is traversed, and the depth information of each first feature corresponding to each pixel point is predicted in the above manner.
[0132] Optionally, the training process of the prediction model can include:
[0133] Obtain multiple sets of sample data; where each set of sample data includes the feature information of a sample image and the distance value corresponding to each pixel point;
[0134] Input the sample data into the prediction model to output a second prediction vector;
[0135] Calculate the loss value of the prediction model according to the second prediction vector and the reference vector; wherein, the reference vector is generated according to the distance value corresponding to each pixel point in the sample data;
[0136] If the loss value is greater than the preset threshold, determine the current prediction model as the trained prediction model;
[0137] If the loss value is less than or equal to the preset threshold, update the model parameters of the prediction model according to the loss value to obtain an updated prediction model;
[0138] Continue to train the updated prediction model until a trained prediction model is obtained.
[0139] In the above manner, dividing the distance range into multiple distance intervals for prediction helps to improve the prediction accuracy. In addition, using the trained prediction model to predict depth information not only helps to improve the prediction efficiency, but also can ensure the prediction accuracy because the prediction accuracy of the trained prediction model meets the standard.
[0140] S104. Detect the target object in the first image according to the second feature.
[0141] In one embodiment, S104 may include:
[0142] Obtain a second image; wherein, the second image is a point cloud image corresponding to the first image;
[0143] Perform feature extraction processing on the second image to obtain a seventh feature;
[0144] Map the seventh feature to the preset coordinate system to obtain an eighth feature;
[0145] Perform fusion processing according to the second feature and the eighth feature to obtain a second fusion feature;
[0146] Detect the target object in the first image according to the second fusion feature.
[0147] Optionally, the manner of mapping the seventh feature to the preset coordinate system may include: performing dimensionality reduction processing on the seventh feature to obtain an eighth feature that conforms to the dimension of the preset coordinate system.
[0148] For example, the seventh feature (i.e., the feature of the point cloud image) usually includes 5 dimensions, such as 1 (number of clusters) × 128 (number of channels) × 180 (length of the BEV space) × 180 (width of the BEV space) × 2 (height). Multiply the height information by the number of channels to obtain 1 × 256 × 180 × 180, and reduce it to 4 dimensions to achieve the mapping to the BEV space.
[0149] In the embodiments of the present application, the point cloud data and the RGB image are mapped to the same preset coordinate system and feature fusion is performed, which is equivalent to combining the pixel features of the RGB image and the depth information of the point cloud data. In this way, it helps to improve the detection accuracy.
[0150] Exemplarily, referring to Figure 3 , which is a schematic diagram of the image detection process provided by the embodiments of the present application. By way of example and not limitation, as Figure 3 shown, in the upper processing flow, first, the feature information of the first image is subjected to feature segmentation processing to extract the first feature of the foreground part in the first image; then the first feature is mapped to the BEV space to obtain the second feature. In the lower processing flow, first, the features in the second image are extracted to obtain the seventh feature; then the seventh feature is mapped to the BEV space to obtain the eighth feature. Then the second feature and the eighth feature are subjected to fusion processing to obtain the second fusion feature; finally, the target object in the first image is detected according to the second fusion feature.
[0151] In the embodiments of the present application, the feature information of the foreground part in the image is extracted, and then the feature information of the foreground part is used for detection, reducing the redundant information of the background part in the image. This not only effectively improves the detection efficiency but also greatly enhances the detection effect. In addition, the pixel features of the RGB image and the depth features of the point cloud image are mapped to the same coordinate system and fused, and the fused features are used for image detection, which can combine various types of feature information, thus helping to improve the detection accuracy.
[0152] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0153] Corresponding to the image detection method described in the above embodiments, Figure 4 is a structural block diagram of the image detection device provided by the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0154] Referring to Figure 4 , the device includes:
[0155] An acquisition unit 41, configured to acquire a first image to be processed; wherein, the first image is an RGB image.
[0156] An extraction unit 42, configured to extract the first feature of the foreground part in the first image by performing image segmentation processing on the first image.
[0157] A mapping unit 43 for mapping the first feature into a preset coordinate system to obtain a second feature; wherein, the preset coordinate system is the coordinate system corresponding to the image in the top-down view.
[0158] A detection unit 44 for detecting a target object in the first image according to the second feature.
[0159] Optionally, the extraction unit 42 is further configured to:
[0160] Perform feature extraction processing on the first image to obtain a third feature;
[0161] Perform segmentation processing on the first image according to the third feature to obtain a target detection frame for the foreground part in the first image;
[0162] Extract the first feature from the third feature according to the target detection frame.
[0163] Optionally, the extraction unit 42 is further configured to:
[0164] Randomly generate a plurality of candidate detection frames in the first image;
[0165] Detect the image category to which the image in each candidate detection frame belongs according to the third feature in each candidate detection frame; wherein, the image categories include a first category and a second category, the first category represents the foreground, and the second category represents the background;
[0166] Filter the plurality of candidate detection frames according to the image category corresponding to each candidate detection frame to obtain the target detection frame.
[0167] Optionally, the extraction unit 42 is further configured to:
[0168] Obtain candidate detection frames with the image category being the first category among the plurality of candidate detection frames to obtain a first detection frame;
[0169] Perform duplicate removal processing on the first detection frame to obtain the target detection frame after duplicate removal.
[0170] Optionally, the mapping unit 43 is further configured to:
[0171] Predict the depth information corresponding to the first feature;
[0172] Map the first feature into the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
[0173] Optionally, the mapping unit 43 is further configured to:
[0174] Input the first feature into the trained prediction model to output a first prediction vector; wherein, the prediction vector includes multiple elements, different elements correspond to different distance values, and each element represents the prediction probability of the distance value corresponding to the element.
[0175] Determine the depth information corresponding to the first feature according to the first prediction vector.
[0176] Optionally, the mapping unit 43 is further configured to:
[0177] Obtain multiple groups of sample data; wherein, each group of the sample data includes the feature information of a sample image and the distance value corresponding to each pixel point.
[0178] Input the sample data into the prediction model to output a second prediction vector.
[0179] Calculate the loss value of the prediction model according to the second prediction vector and the reference vector; wherein, the reference vector is generated according to the distance value corresponding to each pixel point in the sample data.
[0180] If the loss value is greater than a preset threshold, determine the current prediction model as the trained prediction model.
[0181] If the loss value is less than or equal to the preset threshold, update the model parameters of the prediction model according to the loss value to obtain the updated prediction model.
[0182] Continue to train the updated prediction model until the trained prediction model is obtained.
[0183] Optionally, the detection unit 44 is further configured to:
[0184] Obtain a second image; wherein, the second image is the point cloud image corresponding to the first image.
[0185] Perform feature extraction processing on the second image to obtain a seventh feature.
[0186] Map the seventh feature to the preset coordinate system to obtain an eighth feature.
[0187] Perform fusion processing according to the second feature and the eighth feature to obtain a second fusion feature.
[0188] Detect the target object in the first image according to the second fusion feature.
[0189] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought, for details, please refer to the method embodiment part, and will not be elaborated here.
[0190] In addition, Figure 4 the illustrated image detection device may be a software unit, a hardware unit, or a unit combining software and hardware built into an existing terminal device, may also be integrated into the terminal device as an independent attachment, or may exist as an independent terminal device.
[0191] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example for illustration. In practical applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, can also be physically present separately for each unit, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0192] Figure 5 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 5 shown, the terminal device 5 in this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the foregoing image detection method embodiments are implemented.
[0193] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 5 merely an example of the terminal device 5, does not constitute a limitation on the terminal device 5, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0194] The processor 50 may be a Central Processing Unit (CPU), and the processor 50 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0195] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as the hard disk or memory of the terminal device 5. In other embodiments, the memory 51 may also be an external storage device of the terminal device 5, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal device 5. Further, the memory 51 may also include both the internal storage unit and the external storage device of the terminal device 5. The memory 51 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as the program code of the computer program, etc. The memory 51 may also be used to temporarily store data that has been output or is to be output.
[0196] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.
[0197] The embodiments of the present application provide a computer program product, and when the computer program product runs on a terminal device, the terminal device can implement the steps in the above various method embodiments when executed.
[0198] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0199] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0200] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0201] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0202] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An image detection method, characterized in that: include: Acquire a first image to be processed; wherein the first image is an RGB image; Extracting a first feature of a foreground portion of the first image by performing image segmentation processing on the first image; Mapping the first feature to a preset coordinate system to obtain a second feature; wherein the preset coordinate system is a coordinate system corresponding to the image in a top-down perspective; A target object in the first image is detected according to the second feature.
2. The image detection method according to claim 1, characterized in that: The extracting a first feature of a foreground portion of the first image by performing image segmentation processing on the first image includes: Performing feature extraction processing on the first image to obtain a third feature; Segment the first image according to the third feature to obtain an object detection frame of a foreground part of the first image; The first feature is extracted from the third feature according to the target detection frame.
3. The image detection method according to claim 2, characterized in that: The step of segmenting the first image according to the third feature to obtain an object detection frame of a foreground portion of the first image includes: Randomly generating a plurality of candidate detection frames in the first image; Detecting the image category to which the image in each candidate detection frame belongs according to the third feature in each candidate detection frame; wherein the image category includes a first category and a second category, the first category represents the foreground, and the second category represents the background; The plurality of candidate detection frames are filtered according to the image category corresponding to each of the candidate detection frames to obtain the target detection frame.
4. The image detection method according to claim 3, characterized in that: The filtering of the plurality of candidate detection frames according to the image category corresponding to each of the candidate detection frames to obtain the target detection frame includes: Obtain a plurality of candidate detection frames whose image categories are the first category from the candidate detection frames to obtain a first detection frame; The first detection frame is deduplicated to obtain the deduplicated target detection frame.
5. The image detection method according to any one of claims 1 to 4, characterized in that: Mapping the first feature to a preset coordinate system to obtain a second feature includes: Predicting depth information corresponding to the first feature; The first feature is mapped to the preset coordinate system according to the pixel coordinates corresponding to the first feature and the depth information.
6. The image detection method according to claim 5, characterized in that: The predicting the depth information corresponding to the first feature includes: Inputting the first feature into the trained prediction model and outputting a first prediction vector; wherein the prediction vector includes a plurality of elements, different elements correspond to different distance values, and each element represents a prediction probability of the distance value corresponding to the element; Determine depth information corresponding to the first feature according to the first prediction vector.
7. The image detection method according to claim 6, characterized in that: The method further comprises: Acquire multiple groups of sample data; wherein each group of sample data includes feature information of a sample image and a distance value corresponding to each pixel point; Inputting the sample data into the prediction model and outputting a second prediction vector; Calculating the loss value of the prediction model according to the second prediction vector and the reference vector; wherein the reference vector is generated according to the distance value corresponding to each pixel point in the sample data; If the loss value is greater than a preset threshold, the current prediction model is determined as the trained prediction model; If the loss value is less than or equal to a preset threshold, updating the model parameters of the prediction model according to the loss value to obtain the updated prediction model; Continue to train the updated prediction model until the trained prediction model is obtained.
8. The image detection method according to any one of claims 1 to 6, characterized in that: The detecting the target object in the first image according to the second feature comprises: Acquire a second image; wherein the second image is a point cloud image corresponding to the first image; performing feature extraction processing on the second image to obtain a seventh feature; Mapping the seventh feature to the preset coordinate system to obtain an eighth feature; Performing fusion processing according to the second feature and the eighth feature to obtain a second fusion feature; Detect a target object in the first image according to the second fusion feature.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.