Monocular vision occlusion attribute identification method and device, storage medium and electronic equipment
By converting three-dimensional point cloud data into two-dimensional distance images and fusing them with two-dimensional visual images, the occlusion attributes are automatically identified using deep learning technology, which solves the problem of high efficiency and low manual verification cost in monocular 3D perception, and realizes automated and efficient occlusion attribute recognition.
Patent Information
- Application Number
- CN202510328729.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, monocular 3D perception algorithms require manual verification of occlusion properties in autonomous driving, resulting in high verification cost and low efficiency.
Two-dimensional visual images are acquired through monocular vision sensors and vehicle-mounted radars to obtain three-dimensional point cloud data, convert the three-dimensional point cloud data into two-dimensional distance images, and fuse it with the two-dimensional visual image, and automatically identify occlusion attributes using deep learning and image processing technology.
The automatic identification of occlusion attributes is realized, which reduces the tedious process of manual verification, improves efficiency and eliminates subjectivity and inconsistency, and improves the accuracy and robustness of target detection.
Smart Images

Figure CN120339764A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence. Specifically, the embodiments of the present application relate to a method, device, storage medium, and electronic device for identifying monocular vision occlusion attributes. Background Art
[0002] With the rapid development of artificial intelligence, the field of autonomous driving has also made great leaps. Target perception is a very important task requirement in the field of autonomous driving. The planning and control of vehicles are inseparable from perception. Currently, perception is mainly divided into lidar perception and camera-based perception. Monocular 3D perception has attracted increasing attention in the industrial and academic communities due to its advantages in terms of installation cost and immunity to external environmental interference.
[0003] Currently, most monocular 3D perception algorithms are based on 2D object detection methods. Based on the original 2D detection, 3D information is directly learned to obtain 3D results. This method of obtaining 3D detection results through a 2D detection head is relatively simple and can be applied in any 2D object framework. Usually, since accurate distance information cannot be obtained through camera devices, true value labels often need to be provided by lidar in visual annotation. Therefore, joint visual-point cloud annotation is required.
[0004] However, the lidar in the true value vehicle is installed at a relatively high height and can often detect more targets than vision, that is, there are many targets that are invisible in the vision system. Therefore, multi-view visual image verification needs to be performed manually, and labels indicating whether vision is occluded are marked on the point cloud annotation box. The manual verification process is often time-consuming and laborious, resulting in a large amount of manual verification costs. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, storage medium, and electronic device for identifying monocular vision occlusion attributes, so as to at least solve the technical problems of high verification costs and low efficiency caused by manually determining occlusion attributes in related technologies.
[0006] According to one aspect of the embodiments of the present application, a method for identifying monocular vision occlusion attributes is provided, including: obtaining a two-dimensional visual image of a vehicle environment through a monocular vision sensor, and obtaining three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; the two-dimensional visual image includes visual information of multiple two-dimensional targets; the three-dimensional point cloud data includes distance information of multiple three-dimensional targets relative to the in-vehicle radar; converting the three-dimensional point cloud data into a two-dimensional distance image, and performing fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature; classifying the fusion feature to obtain the occlusion attribute of each three-dimensional target among the multiple three-dimensional targets in the two-dimensional visual image; the occlusion attribute is used to eliminate the three-dimensional detection frame corresponding to the three-dimensional target with a specified attribute during the visual and point cloud joint annotation process.
[0007] According to another aspect of the embodiments of the present application, a device for identifying monocular vision occlusion attributes is further provided, including: an acquisition module, configured to obtain a two-dimensional visual image of a vehicle environment through a monocular vision sensor, and obtain three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; the two-dimensional visual image includes visual information of multiple two-dimensional targets; the three-dimensional point cloud data includes distance information of multiple three-dimensional targets relative to the in-vehicle radar; a fusion module, configured to convert the three-dimensional point cloud data into a two-dimensional distance image, and perform fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature; an attribute classification module, configured to classify the fusion feature to obtain the occlusion attribute of each three-dimensional target among the multiple three-dimensional targets in the two-dimensional visual image; the occlusion attribute is used to eliminate the three-dimensional detection frame corresponding to the three-dimensional target with a specified attribute during the visual and point cloud joint annotation process.
[0008] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is further provided, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] According to still another aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0010] According to still another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to execute the steps in any one of the above method embodiments through the computer program.
[0011] Through the present application, three-dimensional point cloud data is converted into a two-dimensional distance image, and fused with a two-dimensional visual image to obtain a fused feature. The construction of the fused feature can utilize both the appearance information of the visual image and the depth information of the point cloud data, thereby more accurately determining the occlusion attribute of the target. Compared with the prior art, the embodiments of the present application automatically complete the occlusion attribute recognition process, avoiding the cumbersome and potential errors of manual comparison, eliminating the subjectivity and inconsistency of manual judgment, realizing the automation of occlusion attribute recognition, solving the technical problems of high verification cost and low efficiency caused by manually determining the occlusion attribute in the related art, and greatly improving the efficiency. Description of the Drawings
[0012] Figure 1 is a schematic diagram of an application scenario of a monocular vision occlusion attribute recognition method according to an embodiment of the present application;
[0013] Figure 2 is a schematic flowchart of an optional monocular vision occlusion attribute recognition method according to an embodiment of the present application;
[0014] Figure 3 is a schematic structural diagram of an optional monocular vision occlusion attribute recognition network according to an embodiment of the present application;
[0015] Figure 4 is a schematic block diagram of an optional monocular vision occlusion attribute recognition device according to an embodiment of the present application;
[0016] Figure 5 is a schematic block diagram of a computer system of an optional electronic device according to an embodiment of the present application. Detailed Embodiments
[0017] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0018] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0019] According to one aspect of the embodiments of the present application, a method for identifying monocular vision occlusion attributes is provided. Optionally, in this embodiment, the above-mentioned method for identifying monocular vision occlusion attributes can be but is not limited to being applied to a hardware environment such as Figure 1 as shown, including a monocular vision sensor 102, a vehicle-mounted radar 104, and a computer device 106. Among them, the monocular vision sensor 102 and the vehicle-mounted radar 104 are respectively arranged on the same vehicle, and the monocular vision sensor 102 and the vehicle-mounted radar 104 are respectively in communication with the computer device 106 through a network. The monocular vision sensor 102 collects a two-dimensional visual image of the vehicle environment where the vehicle is located, and transmits the collected two-dimensional visual image to the computer device 106 through the network. The vehicle-mounted radar 104 simultaneously collects three-dimensional point cloud data of the same vehicle environment where the vehicle is located, and transmits the collected three-dimensional point cloud data to the computer device 106 through the network. The computer device 106 converts the three-dimensional point cloud data into a two-dimensional distance image, and performs fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature; the computer device 106 classifies the fusion feature to obtain the occlusion attribute of each three-dimensional target among multiple three-dimensional targets in the two-dimensional visual image; the occlusion attribute is used to remove the three-dimensional detection box corresponding to the three-dimensional target with the specified attribute during the visual and point cloud joint annotation process.
[0020] The above-mentioned network can include but is not limited to at least one of the following: a wired network, a wireless network. The above-mentioned wired network can include but is not limited to at least one of the following: a wide area network, a metropolitan area network, a local area network. The above-mentioned wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The computer device 106 can be but is not limited to a PC (Personal Computer), a mobile phone, a tablet computer, a vehicle-mounted terminal, etc. The vehicle-mounted radar 104 can be but is not limited to a lidar, a millimeter wave radar or other radar types.
[0021] The monocular vision occlusion attribute recognition method according to the embodiments of the present application can be executed by a computer device 106. Figure 2 It is a schematic flowchart of an optional monocular vision occlusion attribute recognition method according to the embodiments of the present application. As Figure 2 shown, the process of this method can include the following steps:
[0022] Step S202, obtain a two-dimensional visual image of the vehicle environment through a monocular vision sensor, and obtain three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; the two-dimensional visual image includes visual information of multiple two-dimensional targets; the three-dimensional point cloud data includes distance information of multiple three-dimensional targets relative to the in-vehicle radar.
[0023] Among them, the monocular vision occlusion attribute recognition method in this embodiment can be applied to the field of autonomous driving, especially in a monocular vision 3D target detection system, and can solve the problem of target occlusion in vehicle environment perception.
[0024] In autonomous driving technology, the perception of the vehicle's surrounding environment is the key to achieving safe driving. Generally, the vehicle is equipped with multiple sensors, such as a monocular vision sensor (camera) and an in-vehicle radar (such as a lidar), to obtain environmental information. The monocular vision sensor can capture a two-dimensional image of the front scene, while the in-vehicle radar can provide three-dimensional coordinate information of the objects in the environment, including distance, azimuth, etc. However, due to the limitations of monocular vision, it cannot directly obtain the depth information of the object, and although the in-vehicle radar can provide accurate three-dimensional information, its cost is high and it is affected by environmental conditions. The embodiments of the present application use three-dimensional point cloud data to supplement or enhance the recognition of targets in the two-dimensional visual image to combine the advantages of both.
[0025] A monocular vision sensor is a sensor that captures images using only one camera. It can obtain a two-dimensional visual image of the front scene of the vehicle, but cannot directly measure the depth information of the object. The two-dimensional visual image obtained by collecting the vehicle environment through the monocular vision sensor includes visual information of multiple two-dimensional targets, such as color, shape, position, etc., but lacks depth information, that is, the actual distance of the target from the camera. In an autonomous driving system, the two-dimensional visual image is used to identify vehicles, pedestrians, road signs, etc.
[0026] On-vehicle radar is a sensor that measures the distance to a target by emitting laser light and receiving the reflected laser light. For example, Light Detection and Ranging (LiDAR). On-vehicle radar can acquire three-dimensional point cloud data of objects in the environment. Each point in the three-dimensional point cloud data represents the precise distance information of that point relative to the on-vehicle radar. The three-dimensional point cloud data acquired by the on-vehicle radar includes multiple points in the vehicle environment, and each point has its three-dimensional coordinate information, which can provide the precise distance, azimuth, and height information of the target relative to the radar. Optionally, the computer device acquires a two-dimensional visual image of the vehicle environment where the vehicle is located through a monocular vision sensor disposed on the vehicle, and acquires three-dimensional point cloud data of the same vehicle environment where the vehicle is located through the on-vehicle radar disposed on the same vehicle.
[0027] Step S204, convert the three-dimensional point cloud data into a two-dimensional distance image, and perform fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature.
[0028] Among them, in the scenario of monocular vision label occlusion attribute recognition, the 3D information of the target, including position, size, and occlusion state, depends on the three-dimensional point cloud data provided by the lidar. However, there are information asymmetry and fusion problems between lidar data and the two-dimensional visual images captured by the monocular camera. This is mainly because the two-dimensional image lacks depth information, while although the three-dimensional point cloud data contains depth information, its field of view may exceed the viewing angle range of the camera, resulting in some targets being visible in the three-dimensional data but not visible in the two-dimensional image. In this case, it is necessary to manually perform multi-view visual image verification and label the visual occlusion attribute on the point cloud annotation box. The manual verification process is often time-consuming and laborious, bringing a large amount of manual verification costs. Therefore, in order to improve the accuracy and efficiency of target detection in the autonomous driving system, the embodiments of this application convert the three-dimensional point cloud data into a two-dimensional distance image, obtain a fusion feature by fusing the features of the two-dimensional visual image and the two-dimensional distance image, and then judge the occlusion attribute of the target according to the fusion feature, so as to eliminate the detection frames of invisible targets in the joint annotation of vision and point cloud, reduce the manual verification work, and improve the data quality and detection efficiency.
[0029] The embodiments of the present application provide a processing method capable of effectively fusing 3D point cloud data and 2D visual images. Specifically, by processing the 3D point cloud data, it is converted into a 2D distance image (also known as a Range image). Each pixel point in the 2D distance image represents the distance from the point in the corresponding scene to the vehicle-mounted radar. Then, the Range image (containing the distance information of the 3D point cloud data) and the 2D visual image (containing the appearance information of the target) are fused to obtain a fused feature, where the fused feature contains visual information (such as color, edge, texture) and depth information (the distance from the target to the vehicle-mounted radar). In the embodiments of the present application, the depth information of the 3D point cloud data is converted into a visible format of a 2D distance image, which is convenient for fusion and analysis with a monocular visual image (i.e., a 2D visual image). Through a fusion algorithm, the information in the two images is combined, which can improve the accuracy and robustness of target detection, especially when dealing with the problem of target occlusion.
[0030] In some embodiments, converting the 3D point cloud data into a 2D distance image is used to map 3D spatial information onto a 2D image, facilitating subsequent computer vision algorithm processing. The 3D point cloud data can be converted into a 2D distance image by using a polar coordinate conversion method, a depth map generation method, or other methods.
[0031] For example, the specific process of the polar coordinate conversion method is as follows: First, preprocess the 3D point cloud data collected by the vehicle-mounted radar, including removing noise points, ground point filtering, and point cloud registration, etc., to ensure the accuracy and consistency of the point cloud data. Convert the preprocessed 3D point cloud data from the Cartesian coordinate system to the polar coordinate system, where the position information (x, y, z) of each point cloud point is converted to (ρ, θ, z). ρ represents the radial distance from the point to the center of the vehicle-mounted radar, θ represents the horizontal angle, and z represents the vertical height. In the polar coordinate system, according to the ρ (radial distance) value, project the 3D point cloud data onto a 2D plane to form a 2D distance image with the radar center as the origin. Each pixel of the 2D distance image corresponds to a direction angle, and its value represents the nearest point distance or the minimum distance value in that direction. In this way, a 2D distance image containing environmental depth information can be constructed.
[0032] For example, the specific generation process of the depth map generation method is as follows: First, perform necessary registration and filtering on the 3D point cloud data of the vehicle-mounted radar to ensure the data quality. Project the 3D point cloud data onto an image plane that matches the camera's perspective. Based on the projection result, construct a depth map, where the value of each pixel represents the distance information of the pixel in the real world. The depth map has the same resolution and perspective as the 2D image captured by the camera, facilitating the fusion and comparison between the two. The set depth map is the 2D distance image.
[0033] Optionally, the computer device converts the three-dimensional point cloud data into a two-dimensional distance image by using a polar coordinate conversion method, a depth map generation method, or other methods, and performs a fusion process on the two-dimensional distance image and the two-dimensional visual image by using a fusion algorithm (such as a weighted average method, a principal component analysis (PCA) method, a wavelet transform method, etc.) to obtain a fusion feature.
[0034] Step S206: Classify the fusion feature to obtain the occlusion attribute of each three-dimensional object in the multiple three-dimensional objects in the two-dimensional visual image; the occlusion attribute is used to remove the three-dimensional detection frame corresponding to the three-dimensional object with the specified attribute during the visual and point cloud joint annotation process.
[0035] Among them, the occlusion attribute refers to the characteristic that a three-dimensional object is partially or completely obscured in a two-dimensional visual image due to reasons such as other objects, light conditions, or sensor viewing angle limitations. The occlusion attribute is used to determine the occlusion state of the three-dimensional detection frame of the three-dimensional object in the two-dimensional visual image. The occlusion attribute can be divided into multiple levels, such as completely visible, partially occluded, severely occluded, invisible, etc. The specified attribute refers to one or several attributes in the occlusion attribute that affect target perception. For example, the specified attribute refers to the "invisible" attribute.
[0036] Combined with the current image processing technology, an occlusion classification method based on deep learning, a fusion feature classification method based on projection transformation and occlusion detection, and other methods can be used. Among them, the occlusion classification method based on deep learning uses a deep learning network (such as a convolutional neural network CNN) to learn and classify the fusion feature; the deep learning network determines the occlusion state of each three-dimensional object in the two-dimensional visual image by extracting features from the fusion feature, constructing a classifier, etc. In the training stage of the deep learning network, a classification network for the Range image and the two-dimensional image based on deep learning is constructed. The occlusion attribute of each three-dimensional object in the manually annotated three-dimensional point cloud data in the monocular vision system is used as the ground truth label. The historical point cloud data and the historical two-dimensional image data are collected through the monocular vision sensor and the vehicle-mounted radar on the ground truth vehicle in the same vehicle environment respectively. The historical point cloud data, the historical two-dimensional image data, and the ground truth label are input into the classification network respectively; the classification network converts the historical point cloud data into a historical two-dimensional distance image (i.e., a historical Range image), fuses the historical Range image and the historical two-dimensional image data to obtain a historical fusion feature, and judges the occlusion attribute based on the historical fusion feature, outputs the predicted occlusion attribute corresponding to each three-dimensional object in the monocular vision system, calculates the error between the predicted occlusion attribute of each three-dimensional object and the corresponding ground truth label, and updates the network weights through optimization algorithms such as gradient descent. Repeat the above steps until the loss converges or reaches the preset number of iterations to obtain a trained classification network for the Range image and the two-dimensional image based on deep learning.
[0037] The fusion feature classification method based on projection transformation and occlusion detection projects 3D point cloud data onto a 2D image plane to obtain a projection image, combines the 2D visual image and the projection image, and uses an occlusion detection algorithm to judge the occlusion attribute of each 3D object. Specifically, the 3D point cloud data is projected onto the 2D image plane using camera parameters to obtain a projection image, the projection image and the 2D visual image are matched and compared, image processing techniques (such as edge detection, morphological processing, etc.) are used to detect the occlusion area, and according to the occlusion detection result, each 3D object is classified and labeled to obtain its occlusion state in the 2D visual image.
[0038] The 3D detection box corresponding to a 3D object refers to a box in 3D space that represents the position, size, and orientation of the object. Different from the 2D detection box, the 3D detection box has depth information and can more accurately describe the position and size of the object in the real world. In the applications of autonomous driving and robot vision, the 3D detection box is the basis for object detection and recognition based on point cloud data and can provide accurate 3D information of the object. Through the judgment of the occlusion attribute, those 3D detection boxes that are invisible in the visual image can be excluded when jointly annotating visual and point cloud data, thereby improving the quality of the annotated data and the accuracy of subsequent object detection.
[0039] Optionally, the computer device adopts a pre-constructed classification network, such as a convolutional neural network (CNN) or a recurrent neural network (RNN) combined with an attention mechanism. The classification network is used to learn the occlusion attribute of the object from the fusion features. The computer device predicts the fusion features through the constructed classification network to obtain the probability distribution of the occlusion attribute of each 3D object among multiple 3D objects in the 2D visual image, and determines the occlusion attribute of the object according to the category with the highest probability, such as "fully visible", "partially occluded", "severely occluded", or "invisible".
[0040] Through the embodiments provided by the present application, the 3D point cloud data is converted into a 2D distance image and fused with the 2D visual image to obtain fusion features. The construction of the fusion features can utilize both the appearance information of the visual image and the depth information of the point cloud data, thereby more accurately judging the occlusion attribute of the object. Compared with the prior art, the embodiments of the present application automatically complete the occlusion attribute recognition process, avoid the cumbersome and potential errors of manual comparison, eliminate the subjectivity and inconsistency of manual judgment, realize the automation of occlusion attribute recognition, solve the technical problems of high verification cost and low efficiency caused by manually determining the occlusion attribute in the related art, and greatly improve the efficiency.
[0041] In an exemplary embodiment, converting the 3D point cloud data into a 2D distance image includes:
[0042] Project the three-dimensional point cloud data onto the two-dimensional image plane corresponding to the monocular vision sensor to obtain a two-dimensional projection map; according to the visible range of the two-dimensional vision image, crop the two-dimensional projection map so that the visible range of the cropped two-dimensional projection map is consistent with the visible range of the two-dimensional vision image, and obtain a two-dimensional distance image.
[0043] Among them, the two-dimensional image plane refers to the plane where the image captured by the monocular vision sensor exists. In this embodiment, projecting the three-dimensional point cloud data onto the two-dimensional image plane corresponding to the monocular vision sensor can form a two-dimensional projection map.
[0044] Due to the limitation of the field of view (FOV) of the monocular vision sensor, some three-dimensional point cloud data may fall outside the blind area of the image. In order to make the information of the two-dimensional projection map correspond to that of the two-dimensional vision image, it is necessary to crop the two-dimensional projection map to ensure that its visible range is consistent with that of the monocular vision sensor.
[0045] Optionally, the computer device converts each point in the three-dimensional point cloud data from the vehicle coordinate system (or world coordinate system) to the camera coordinate system of the monocular vision sensor, projects the three-dimensional points in the camera coordinate system onto the two-dimensional image plane, and during the projection process, retains the depth information (i.e., the z-axis value) of the three-dimensional point cloud for constructing the two-dimensional distance image. The depth value of each projection point can be used as the grayscale value or color value of the point in the two-dimensional image. Secondly, the computer device obtains the visible range of the monocular vision sensor, which is represented by the size of the monocular vision sensor or the width and height of the two-dimensional vision image. According to the visible range of the monocular vision sensor, crop the projected two-dimensional image to ensure that only the area that the monocular vision sensor can actually see is retained. Finally, for the cropped two-dimensional projection map, the computer device uses the retained depth information to construct a two-dimensional distance image, and the value of each pixel represents the distance from the point at that position in the three-dimensional space to the sensor. It can be represented by a grayscale image, where the grayscale value is in a certain proportional relationship with the distance; or use pseudo-color coding to map different distance ranges to different colors.
[0046] Through this embodiment, considering the limited visible range of the monocular vision sensor, directly using the complete two-dimensional projection map may contain a large amount of point cloud information that is invisible in the vision image. Therefore, in this embodiment, the two-dimensional projection map is cropped according to the visible range of the monocular vision image, only the points that fall within the effective area of the image are retained, and the point cloud data outside the FOV is excluded, ensuring that the finally processed two-dimensional distance image only contains the information of the targets within the visible range. Compared with manual verification, the automated projection and cropping process can quickly and accurately identify which point cloud targets are visible in the vision image and which are occluded, significantly improving the efficiency and accuracy of occlusion attribute recognition.
[0047] In an exemplary embodiment, a two-dimensional distance image and a two-dimensional visual image are fused to obtain a fused feature, including:
[0048] Feature extraction is performed on the two-dimensional distance image to obtain two-dimensional spatial data features; feature extraction is performed on the two-dimensional visual image to obtain two-dimensional image detection features; the feature channels of the two-dimensional spatial data features are the same as those of the two-dimensional image detection features, and the size of the two-dimensional spatial data features is the same as that of the two-dimensional image detection features; the two-dimensional spatial data features and the two-dimensional image detection features are subjected to feature superposition processing according to the same feature channels to obtain a spliced target feature; and the spliced target feature is subjected to fusion processing to obtain a fused feature.
[0049] Among them, by converting three-dimensional point cloud data into a two-dimensional distance image and performing feature extraction on it, two-dimensional spatial data features are obtained, and the two-dimensional spatial data features are feature vectors reflecting the attributes of the target in terms of spatial position, size, shape, etc. The extracted two-dimensional spatial data features are used as part of the fusion processing and combined with visual features to judge the occlusion attribute of the target.
[0050] The two-dimensional visual image captured by a monocular camera is used for the preliminary detection of the target, and two-dimensional image detection features are obtained through feature vectors extracted by technologies such as CNN. The two-dimensional image detection features can identify and locate the target in the image and are feature vectors for target detection and classification, such as edges, colors, textures, etc., providing appearance information for occlusion attribute recognition.
[0051] In order to enable the effective fusion of the two-dimensional spatial data features and the two-dimensional image detection features, in this embodiment, the feature channels of the two-dimensional spatial data features are the same as those of the two-dimensional image detection features. In a convolutional neural network, the feature channel refers to the depth of the feature map, which usually represents different feature types or filters for capturing multi-faceted information of the image.
[0052] The features of the two-dimensional distance image and the features of the two-dimensional visual image are spliced under the condition of the same number of channels to obtain a spliced target feature containing more information, where splicing refers to the process of splicing feature vectors from different sources in a certain dimension (usually the channel dimension) to form a comprehensive feature vector containing more information. In this embodiment, the purpose of performing feature superposition processing on the two-dimensional spatial data features and the two-dimensional image detection features is to integrate depth information and appearance information and provide more comprehensive data support for subsequent occlusion attribute classification.
[0053] Optionally, Figure 3 is a schematic structural diagram of an optional monocular vision occlusion attribute recognition network according to an embodiment of the present application, as Figure 3As shown, the computer device extracts features from the two-dimensional distance image through a first feature extraction network (referring to an object detection model based on a convolutional neural network (CNN), such as YOLO (You Only Look Once), SSD (Single Shot Detector), or Faster R-CNN (Faster Region-based Convolutional Neural Network), etc.) to obtain two-dimensional spatial data features, and extracts features from the two-dimensional visual image through a second feature extraction network (the same as the first feature extraction network) to obtain two-dimensional image detection features, so that the feature channels of the obtained two-dimensional spatial data features are consistent with the feature channels of the two-dimensional image detection features, and the size of the two-dimensional spatial data features is consistent with the size of the two-dimensional image detection features; then, through the feature superposition module, the aligned two-dimensional spatial data features and two-dimensional image detection features are subjected to feature superposition processing according to the same feature channels, and a feature splicing operation is used to merge the two-dimensional spatial data features and two-dimensional image detection features in the channel dimension to obtain the spliced target features; finally, the feature fusion module performs fusion and feature extraction processing on the spliced target features to obtain the fused features, and the classification network classifies the fused features to obtain the occlusion attributes of each three-dimensional target in the two-dimensional visual image. In this embodiment, the occlusion attributes of the monocular visual image are autonomously learned by the neural network, avoiding the matching error problem caused by the inaccuracy of the two-dimensional projection boxes obtained by three-dimensional projection in methods such as IoU (Intersection over Union) comparison, which is simpler, more accurate, and has stronger robustness.
[0054] Through this embodiment, it is ensured that the two-dimensional spatial data features (extracted from the two-dimensional distance image) are consistent with the two-dimensional image detection features (extracted from the two-dimensional visual image) in terms of feature channels and size, ensuring that the two types of features can be processed and understood by the neural network at the same level, and avoiding information loss or fusion difficulties caused by feature mismatch; by superimposing the two-dimensional spatial data features and two-dimensional image detection features on the same feature channels, the spliced target features are obtained, and then the spliced target features are subjected to fusion processing to generate the fused features. This process realizes the cross-modal fusion between the depth information in the point cloud data and the visual information in the 2D image data. Feature superposition enables the neural network to "see" both the depth and appearance attributes of the target simultaneously, while fusion processing combines these attributes through the deep learning ability of the neural network to form a more comprehensive understanding of the target occlusion state. This fusion not only considers the positional relationship of the target in space but also combines the appearance features of the target, such as texture and shape, so as to more accurately judge whether the target is occluded and the degree of occlusion.
[0055] In an exemplary embodiment, extracting features from the two-dimensional distance image to obtain two-dimensional spatial data features includes:
[0056] Perform object recognition on the two-dimensional distance image to obtain the two-dimensional projection box corresponding to each three-dimensional object; determine the two-dimensional spatial data features according to the two-dimensional projection box corresponding to each three-dimensional object; the two-dimensional projection box corresponding to each three-dimensional object is the two-dimensional projection box of the three-dimensional detection box corresponding to each three-dimensional object on the two-dimensional image plane.
[0057] Among them, the two-dimensional spatial data features include the position and size of the two-dimensional projection box obtained by projecting the three-dimensional detection box corresponding to each three-dimensional object onto the two-dimensional image plane corresponding to the monocular vision sensor. The position information of the two-dimensional projection box includes the center coordinates of the box (x and y coordinates in the image plane), and the size information includes the width and height of the box. Information such as the position and size of the two-dimensional projection box can be used to locate the objects in the two-dimensional distance image and combined with the two-dimensional image detection features in the fusion algorithm to identify the occlusion attributes of the objects.
[0058] Optionally, as Figure 3 shown, the computer device inputs the two-dimensional distance image into the first feature extraction network. The first feature extraction network performs the positioning and recognition of the target area, outputs the bounding box of each detected three-dimensional object on the image, that is, the two-dimensional projection box, and at the same time gives the object category and confidence level within the two-dimensional projection box. Extract the position parameters of the two-dimensional projection box, including the center coordinates and size parameters of the two-dimensional projection box, such as width w and height h. According to the position and size of the two-dimensional projection box corresponding to each three-dimensional object and the corresponding depth information features, construct the two-dimensional spatial data features. Among them, the two-dimensional spatial data features are represented in the form of a feature map. The feature map corresponding to the two-dimensional spatial data features includes heatmap (classification feature map, Gaussian heatmap), x (x-direction offset), y (y-direction offset), w (width), h (height), and other feature maps. The scale of the feature map corresponding to the two-dimensional spatial data features is B×1×H×W. Among them, the heatmap is used to represent the object category confidence level, the x and y direction offsets are used to refine the position of the two-dimensional projection box, the width and height are used to describe the size of the two-dimensional projection box, and B in the scale B×1×H×W represents the batch size, that is, the number of images processed at one time; H and W represent the height and width of the feature map, reflecting the ability of the first feature extraction network to capture the details of the image.
[0059] Through this embodiment, target detection is directly performed on the two-dimensional distance image, and the two-dimensional projection box of each target obtained accurately reflects the visible part of the target in the monocular vision image. Different from the traditional target detection based only on RGB images, this method utilizes the depth information contained in the distance image and can more accurately determine whether the target is occluded and the degree of occlusion. After obtaining the two-dimensional projection box, the two-dimensional spatial data features of each target are further extracted and constructed. These features not only include the position and size, but may also contain the depth information of the target, as well as other information helpful for classification (such as target category, confidence, etc.). By representing these features in the form of a feature map, the occlusion state of the target can be converted into a data form that can be processed by the neural network.
[0060] In an exemplary embodiment, feature extraction is performed on the two-dimensional visual image to obtain two-dimensional image detection features, including:
[0061] Feature extraction is performed on the two-dimensional visual image through a two-dimensional feature encoding module to obtain a first target feature. The two-dimensional feature encoding module includes a backbone network, a feature pyramid network, and a path aggregation network. The backbone network is used to perform feature extraction on the two-dimensional visual image to obtain a first feature. The feature pyramid network is used to perform semantic enhancement processing on the first feature to obtain a second feature. The path aggregation network is used to perform semantic supplementation on the second feature to obtain a first target feature. Target detection is performed on the first target feature through a target detection head to obtain a two-dimensional detection box corresponding to each two-dimensional target. According to the two-dimensional detection box corresponding to each two-dimensional target, the two-dimensional image detection features are determined.
[0062] Among them, as Figure 3 shown, the second feature extraction network includes a two-dimensional feature encoding module and a target detection head, and the second feature extraction network is used to perform multi-scale feature extraction on the two-dimensional visual image. As Figure 3As shown in the figure, the two-dimensional feature encoding module includes a backbone network, a Feature Pyramid Network (FPN), and a Path Aggregation Network (PAN). These networks work together to gradually enhance the semantic information and localization accuracy of features from the original image to the high-level feature map, providing high-quality input features for subsequent object detection and attribute recognition. In this embodiment, the backbone network is the basic part of the two-dimensional feature encoding module and usually adopts a deep convolutional neural network (such as ResNet, DLA, etc.). It is responsible for performing preliminary feature extraction on the input two-dimensional visual image, generating low-level feature maps, that is, the first feature, which contains the basic texture and edge information of the image and is the starting point for subsequent feature enhancement and fusion. The Feature Pyramid Network (FPN) is based on the backbone network and, through a top-down approach, fuses high-level strong semantic features with low-level fine localization features to generate feature maps of different scales, that is, the second feature, which can improve the detection ability of the two-dimensional feature encoding module for objects of different sizes and is a common enhancement means in object detection algorithms. The Path Aggregation Network (PAN) is a bottom-up pyramid structure added on the basis of the FPN to supplement the information of low-level features, further enhancing the semantic and localization accuracy of features and generating the final first target feature. It fuses the preliminary features of the backbone network, the semantic enhancement information of the FPN, and the semantic supplementary information of the PAN for subsequent object detection and occlusion attribute recognition. The PAN ensures that the two-dimensional feature encoding module has both rich semantic information at the high level and accurate position information at the low level, improving the robustness of detection.
[0063] The object detection head refers to the part responsible for locating and identifying target objects from the feature map. For example, the object detection head can be CenterPoint (center detection algorithm). It receives the first target feature processed by the two-dimensional feature encoding module as input and, through a series of operations such as convolution, pooling, and fully connected layers, finally outputs the position information of the two-dimensional bounding box of each target on the two-dimensional image, usually including the center coordinates and size parameters of the two-dimensional detection box, such as width w and height h. At the same time, it also gives the category and confidence of these targets and determines the two-dimensional image detection features based on the two-dimensional detection box corresponding to each two-dimensional target. Among them, the two-dimensional image detection features are represented in the form of a feature map, and the feature map corresponding to the two-dimensional image detection features also includes feature maps such as heatmap (classification feature map, Gaussian heatmap), x (offset in the x direction), y (offset in the y direction), w (width), h (height), etc. The scale of the feature map corresponding to the two-dimensional image detection features is also B×1×H×W. It can be understood that the two-dimensional image detection features include the positions and sizes of the two-dimensional detection boxes of each two-dimensional target among multiple two-dimensional targets.
[0064] In some embodiments, in the process of performing feature superposition processing on two-dimensional spatial data features and two-dimensional image detection features according to the same feature channels to obtain spliced target features, the two-dimensional spatial data features and the two-dimensional image detection features are spliced on the corresponding channels. Specifically, the heatmap of the two-dimensional distance image is spliced with the heatmap of the two-dimensional visual image, x is spliced with x, y is spliced with y, w is spliced with w, and h is spliced with h to form the spliced target features.
[0065] Optionally, as Figure 3 shown, the computer device inputs the two-dimensional visual image into the two-dimensional feature encoding module. The backbone network in the two-dimensional feature encoding module extracts features from the two-dimensional visual image to obtain the first feature. The feature pyramid network performs semantic enhancement processing on the first feature to obtain the second feature. The path aggregation network performs semantic supplementation on the second feature to obtain the first target feature. Then, the computer device performs object detection on the first target feature through the object detection head to obtain the position and size information of the two-dimensional detection box corresponding to each two-dimensional object, and at the same time gives the object category and confidence level within the two-dimensional detection box, and extracts the position parameters of the two-dimensional detection box, including the center coordinates and size parameters of the two-dimensional detection box, such as the width w and the height h. According to the position and size of the two-dimensional detection box corresponding to each two-dimensional object, two-dimensional image detection features are constructed. The two-dimensional image detection features are represented in the form of a feature map. The feature map corresponding to the two-dimensional image detection features also includes feature maps such as heatmap (classification feature map, Gaussian heatmap), x (x-direction offset), y (y-direction offset), w (width), h (height), etc. The scale of the feature map corresponding to the two-dimensional image detection features is also B×1×H×W.
[0066] Through this embodiment, in the two-dimensional feature encoding module, not only the backbone network is used for feature extraction, but also the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN) are introduced for feature enhancement and supplementation. Among them, FPN enhances the semantic information of features through a top-down path, while PAN supplements the localization information through a bottom-up path. The combination of the two enables the first target feature to not only contain the detailed information of the target, but also enhances the neural network's recognition ability for multi-scale targets, enabling the neural network to effectively process targets of different sizes. Even under occlusion and complex background conditions, it can maintain a high detection accuracy, reduce the occlusion attribute recognition error caused by the change of target size, and thus be more accurate and robust when processing occlusion attribute recognition; the target detection head performs target detection on the first target feature to obtain the two-dimensional detection frame corresponding to each two-dimensional target. This process utilizes the efficiency and accuracy of deep learning, avoiding the possible false detection and missed detection problems in traditional methods. The target detection head can quickly locate the target and provide accurate bounding box information for subsequent attribute classification. In summary, the combination of the backbone network, FPN, and PAN, as well as the efficient utilization of the target detection head, ensures the rapid generation of two-dimensional detection boxes from images. This step is the basis for occlusion attribute recognition and improves the execution efficiency of the overall solution.
[0067] In an exemplary embodiment, the spliced target feature is fused to obtain a fused feature, including:
[0068] Feature extraction is performed on the spliced target feature through grouped convolution of a preset size to obtain a second target feature; the number of groups of the grouped convolution is equal to half of the number of feature channels of the spliced target feature; feature fusion processing is performed on the second target feature through a convolutional network of a preset depth to obtain a fused feature.
[0069] Among them, when using point cloud data and 2D image joint annotation, it is necessary to manually verify the occlusion attribute of the target, which is time-consuming and costly. By introducing grouped convolution, high-dimensional spliced feature maps can be processed more efficiently, reducing the demand for computing resources, accelerating the training and inference speed of the neural network, thereby reducing the necessity of manual verification and reducing costs. In addition, grouped convolution helps to maintain the locality of features, which is particularly important for accurately identifying occlusion situations because occlusion attributes are often determined by local details in the image. Grouped convolution is a network layer that performs convolution operations on the input feature map in groups, and each group only convolves a specific part of the input, reducing the number of parameters while also being able to maintain the spatial correlation of features. In this embodiment, the size of the grouped convolution is preset. For example, as Figure 3As shown, the size of the grouped convolution can be 3×3. By grouping the concatenated target features, grouped convolution helps to retain the local information of the concatenated target features while reducing the computational burden. Since the dimension of the concatenated target features is relatively high, the use of grouped convolution can effectively reduce the complexity of the neural network and ensure that the training and inference speeds of the neural network are not affected when processing the fused multi-layer features.
[0070] In this embodiment, the number of groups of the grouped convolution is equal to half of the number of feature channels of the concatenated target features. The reasons for this setting are as follows: First, in a deep learning network, the convolutional layer usually contains a large number of parameters, which not only increases the training difficulty of the neural network but also raises the demand for computing resources. By setting the number of groups of the grouped convolution to half of the number of feature channels, the number of parameters can be effectively reduced, and the computational complexity can be lowered. Grouped convolution allows the neural network to perform convolution operations without completely mixing all input channels. Each group only processes a part of the channels of the input feature map, which can significantly reduce the size of the weight matrix and thus lower the computational cost. Second, in image processing tasks, especially in object detection and occlusion attribute recognition, local features are crucial for judging the attributes of specific regions. Setting the number of groups to half of the number of feature channels means that the neural network can still retain most of the local information of the input features without losing key details due to excessive grouping. In this way, grouped convolution can perform efficient feature extraction and processing while maintaining local features. In a deep learning network, the convolutional layer usually contains a large number of parameters, which not only increases the training difficulty of the neural network but also raises the demand for computing resources. By setting the number of groups of the grouped convolution to half of the number of feature channels, the number of parameters can be effectively reduced, and the computational complexity can be lowered. Grouped convolution allows the neural network to perform convolution operations without completely mixing all input channels. Each group only processes a part of the channels of the input feature map, which can significantly reduce the size of the weight matrix and thus lower the computational cost. Finally, in image processing tasks, especially in object detection and occlusion attribute recognition, local features are crucial for judging the attributes of specific regions. Setting the number of groups to half of the number of feature channels means that the neural network can still retain most of the local information of the input features without losing key details due to excessive grouping. In this way, grouped convolution can perform efficient feature extraction and processing while maintaining local features. Reducing the number of parameters also helps to avoid overfitting of the neural network. In the processing of image features, overfitting is a common problem, especially when the training data is limited. By setting the number of groups to half of the number of channels, the neural network will not overly rely on specific feature combinations during the learning process but form a more generalized feature representation, which helps to improve the performance of the neural network on unseen data.
[0071] The problem with manually verifying occlusion attributes lies in the low processing efficiency and the difficulty in ensuring consistency. A convolutional network with a preset depth can abstract and fuse features layer by layer, extracting deep information about the target from a multi-scale perspective, including the texture, shape, and relative position of the occluded area. This deep feature processing method improves the accuracy and robustness of occlusion attribute recognition, reduces misjudgments and missed detections caused by complex occlusion situations, thus solving the problems existing in manual verification and providing a more automatic and accurate occlusion attribute recognition solution. A convolutional network is a neural network mainly used to process data with a grid structure (such as images), and automatically extracts and learns the features of the input data through convolutional operations. In this embodiment, the depth of the convolutional network is preset. For example, the depth of the convolutional network can be 3×3. The convolutional network is used to perform deeper feature fusion processing on the second target feature extracted by grouped convolution to obtain a fused feature. The depth design of the convolutional network is to ensure that the neural network can understand more complex image structures. Through the stacking of multiple convolutional layers, the neural network can learn high-level features of the target, such as texture, shape, and occlusion status.
[0072] Optionally, as Figure 3 shown, the computer device determines the number of feature channels for stitching the target feature. Assume that the stitched target feature has N channels. These channels are evenly divided into N / 2 groups, each group containing N / 2 channels, forming the input for grouped convolution. A convolution kernel of a preset size, such as a 3×3 kernel, is independently applied to each group for convolution operations. During the convolution process, each group only interacts with the channels within the group and does not mix with the channels of other groups. The computer device merges all the feature maps after the above grouped convolution to form the second target feature. The number of channels of the second target feature is still N, but the features of each channel have been convolved, enhancing the ability to describe the target. The computer device uses the second target feature as the input and enters a convolutional network with a preset depth. This network consists of a series of convolutional layers, such as 3 convolutional layers (Group Conv 3×3, Conv 3×3, Conv 3×3), for performing deep feature fusion. In each convolutional layer, a convolution kernel is applied to the input feature map for convolution operations to extract more advanced features. By using an activation function (such as ReLU), the non-linear representation ability of the features is enhanced after the convolution operation. Finally, through the output layer of the convolutional network, the processed feature maps are integrated to obtain a fused feature.
[0073] In this embodiment, grouped convolution is used to extract features from the spliced target features, and the number of groups is equal to half of the number of feature channels of the spliced target features. This specific selection of the number of groups ensures that while reducing the computational complexity and parameters, it is still possible to process complex image features; a deep convolutional network is used to fuse the second target features obtained by the grouped convolution to obtain more representative and descriptive fused features.
[0074] In an exemplary embodiment, the above method further includes:
[0075] During the process of joint visual and point cloud annotation, the 3D detection boxes corresponding to the 3D objects with the occlusion attribute being the specified attribute are removed, and the remaining 3D objects are matched and fused with the 2D detection boxes corresponding to each of the multiple 2D objects to obtain the joint visual and point cloud annotation data corresponding to each 2D object.
[0076] Among them, the identification of the occlusion attribute helps to filter out the visible objects in the vision, and the matching and fusion of the 3D detection box and the 2D detection box are based on removing the invisible objects, and by automatically pairing the remaining 3D object detection boxes with multiple 2D object detection boxes, the efficient fusion of visual and point cloud data is realized. This process not only eliminates the data inconsistency caused by occlusion, but also makes each 2D object obtain the joint visual and point cloud annotation data, that is, the joint visual and point cloud annotation data, by supplementing the missing depth information in the 2D detection box. This data fusion method significantly improves the robustness and accuracy of object detection.
[0077] Optionally, during the process of joint visual and point cloud annotation, the computer device obtains the occlusion attribute of each 3D object, removes the 3D detection boxes corresponding to the 3D objects with the occlusion attribute being "invisible", and matches the detection boxes corresponding to the remaining 3D objects with the 2D detection boxes corresponding to each of the multiple 2D objects in the 2D visual image, so that each 3D object corresponds to its projection in the 2D visual image (i.e., the 2D object), obtaining the association relationship between each 2D object and its corresponding 3D object. According to this association relationship, the position and size of the 2D detection box of the 2D object with the association relationship and the depth of the 3D detection box of the 3D object are fused to obtain the joint visual and point cloud annotation data corresponding to each 2D object, thereby completing the conversion from the point cloud annotation ground truth (i.e., the 3D detection box) to the monocular visual annotation ground truth (i.e., the 2D detection box).
[0078] Through this embodiment, by combining the visual features of the 2D image and the geometric features of the 3D point cloud, richer object descriptions are provided, and the position, size and shape of the object can be identified more accurately, especially in complex or occluded scenes, which helps to reduce false detections and missed detections and improve the accuracy of object detection.
[0079] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0081] According to another aspect of the embodiments of this application, a monocular vision occlusion attribute recognition device is further provided. This monocular vision occlusion attribute recognition device can be used to implement the monocular vision occlusion attribute recognition method provided in the above embodiments, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0082] Figure 4 is a structural block diagram of an optional monocular vision occlusion attribute recognition device according to the embodiments of this application. As Figure 4 shown in, this monocular vision occlusion attribute recognition device includes:
[0083] An acquisition module 402, configured to acquire a two-dimensional visual image of the vehicle environment through a monocular vision sensor, and acquire three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; the two-dimensional visual image includes visual information of multiple two-dimensional targets; the three-dimensional point cloud data includes distance information of multiple three-dimensional targets relative to the in-vehicle radar;
[0084] A fusion module 404, configured to convert the three-dimensional point cloud data into a two-dimensional distance image, and perform fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature;
[0085] An attribute classification module 406 is configured to classify the fused features to obtain the occlusion attributes of each three-dimensional object in the multiple three-dimensional objects in the two-dimensional visual image; the occlusion attributes are used to exclude the three-dimensional detection frames corresponding to the three-dimensional objects with the specified attributes during the joint annotation process of vision and point cloud.
[0086] It should be noted that the acquisition module 402 in this embodiment may be configured to execute the above step S202, the fusion module 404 in this embodiment may be configured to execute the above step S204, and the attribute classification module 406 in this embodiment may be configured to execute the above step S206.
[0087] Through the embodiment provided by this application, the three-dimensional point cloud data is converted into a two-dimensional distance image and fused with the two-dimensional visual image to obtain fused features. The construction of the fused features can utilize both the appearance information of the visual image and the depth information of the point cloud data, thereby more accurately determining the occlusion attributes of the target. Compared with the prior art, the embodiment of this application automatically completes the occlusion attribute recognition process, avoiding the cumbersome and potential errors of manual comparison, eliminating the subjectivity and inconsistency of manual judgment, realizing the automation of occlusion attribute recognition, solving the technical problems of high verification cost and low efficiency caused by manually determining the occlusion attributes in the related art, and greatly improving the efficiency.
[0088] In an exemplary embodiment, the fusion module 404 is further configured to project the three-dimensional point cloud data onto the two-dimensional image plane corresponding to the monocular vision sensor to obtain a two-dimensional projection map; and crop the two-dimensional projection map according to the visible range of the two-dimensional visual image so that the visible range of the cropped two-dimensional projection map is consistent with the visible range of the two-dimensional visual image to obtain a two-dimensional distance image.
[0089] In an exemplary embodiment, the attribute classification module 406 is further configured to extract features from the two-dimensional distance image to obtain two-dimensional spatial data features; extract features from the two-dimensional visual image to obtain two-dimensional image detection features; the feature channels of the two-dimensional spatial data features are consistent with the feature channels of the two-dimensional image detection features, and the size of the two-dimensional spatial data features is consistent with the size of the two-dimensional image detection features; perform feature superposition processing on the two-dimensional spatial data features and the two-dimensional image detection features according to the same feature channels to obtain a spliced target feature; and perform fusion processing on the spliced target feature to obtain a fused feature.
[0090] In an exemplary embodiment, the two-dimensional spatial data features include the positions and sizes of two-dimensional projection frames obtained by projecting the three-dimensional detection frames corresponding to each three-dimensional object onto the two-dimensional image plane corresponding to the monocular vision sensor; the attribute classification module 406 is further configured to perform object recognition on the two-dimensional distance image to obtain the two-dimensional projection frames corresponding to each three-dimensional object; and determine the two-dimensional spatial data features according to the two-dimensional projection frames corresponding to each three-dimensional object; the two-dimensional projection frame corresponding to each three-dimensional object is the two-dimensional projection frame of the three-dimensional detection frame corresponding to each three-dimensional object on the two-dimensional image plane.
[0091] In an exemplary embodiment, the two-dimensional image detection features include the positions and sizes of the two-dimensional detection frames of each two-dimensional object among a plurality of two-dimensional objects; the attribute classification module 406 is further configured to extract features from the two-dimensional vision image through the two-dimensional feature encoding module to obtain a first target feature; the two-dimensional feature encoding module includes a backbone network, a feature pyramid network, and a path aggregation network. The backbone network is configured to extract features from the two-dimensional vision image to obtain a first feature; the feature pyramid network is configured to perform semantic enhancement processing on the first feature to obtain a second feature; the path aggregation network is configured to perform semantic supplementation on the second feature to obtain a first target feature; perform object detection on the first target feature through an object detection head to obtain the two-dimensional detection frames corresponding to each two-dimensional object; and determine the two-dimensional image detection features according to the two-dimensional detection frames corresponding to each two-dimensional object.
[0092] In an exemplary embodiment, the attribute classification module 406 is further configured to extract features from the spliced target feature through grouped convolution with a preset size to obtain a second target feature; the number of groups of the grouped convolution is equal to half of the number of feature channels of the spliced target feature; and perform feature fusion processing on the second target feature through a convolutional network with a preset depth to obtain a fused feature.
[0093] In an exemplary embodiment, the attribute classification module 406 is further configured to, during the visual and point cloud joint annotation process, remove the three-dimensional detection frames corresponding to the three-dimensional objects whose occlusion attribute is a specified attribute, and match and fuse the remaining three-dimensional objects with the two-dimensional detection frames corresponding to each two-dimensional object among the plurality of two-dimensional objects to obtain the visual and point cloud joint annotation data corresponding to each two-dimensional object.
[0094] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: all the above-mentioned modules are located in the same processor; or, the above-mentioned various modules are separately located in different processors in any combination form.
[0095] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein when the program runs, it executes the steps in any one of the above method embodiments.
[0096] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0097] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to execute the steps in any one of the above method embodiments through the computer program. In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0098] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0099] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes a computer program / instructions, and the computer program / instructions include program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, it executes various functions provided by the embodiments of the present application. The above serial numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0100] Figure 5 Schematically shows a block diagram of the computer system structure of the electronic device for implementing the embodiments of the present application. As Figure 5 shown, the computer system 500 includes a CPU (Central Processing Unit) 501, which can execute various appropriate actions and processes according to the program stored in the ROM 502 or the program loaded from the storage part 508 into the RAM 503. In the random access memory 503, various programs and data required for system operation are also stored. The central processing unit 501, the read-only memory 502, and the random access memory 503 are connected to each other through a bus 504. The I / O (Input / Output) interface 505 is also connected to the bus 504.
[0101] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 510 as required so that a computer program read from it can be installed into the storage section 508 as required.
[0102] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit 501, various functions defined in the system of the present application are executed.
[0103] It should be noted that Figure 5 The computer system 500 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present application.
[0104] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general computing device. They can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0105] The above are only the preferred embodiments of the present application, and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application should be included in the protection scope of the present application.
Claims
1. A monocular vision occlusion property recognition method, characterized in that Including: Obtaining a two-dimensional visual image of the vehicle environment through a monocular vision sensor, and obtaining three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; The two-dimensional visual image includes visual information of multiple two-dimensional targets; The three-dimensional point cloud data includes distance information of multiple three-dimensional targets relative to the in-vehicle radar; Converting the three-dimensional point cloud data into a two-dimensional distance image, and performing fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature; Classifying the fusion feature to obtain the occlusion attribute of each three-dimensional target in the multiple three-dimensional targets in the two-dimensional visual image; The occlusion attribute is used to eliminate the three-dimensional detection frame corresponding to the three-dimensional target with the specified attribute during the visual and point cloud joint annotation process.
2. The method according to claim 1, wherein The converting the three-dimensional point cloud data into a two-dimensional distance image includes: Projecting the three-dimensional point cloud data onto the two-dimensional image plane corresponding to the monocular vision sensor to obtain a two-dimensional projection map; According to the visible range of the two-dimensional visual image, cropping the two-dimensional projection map so that the visible range of the cropped two-dimensional projection map is consistent with the visible range of the two-dimensional visual image to obtain the two-dimensional distance image.
3. The method according to claim 1, wherein The performing fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature includes: Performing feature extraction on the two-dimensional distance image to obtain two-dimensional spatial data features; performing feature extraction on the two-dimensional visual image to obtain two-dimensional image detection features; the feature channels of the two-dimensional spatial data features are consistent with the feature channels of the two-dimensional image detection features, and the size of the two-dimensional spatial data features is consistent with the size of the two-dimensional image detection features; Performing feature superposition processing on the two-dimensional spatial data features and the two-dimensional image detection features according to the same feature channels to obtain a spliced target feature; performing fusion processing on the spliced target feature to obtain the fusion feature.
4. The method according to claim 3, characterized in that, The two-dimensional spatial data features include the positions and sizes of two-dimensional projection frames obtained by projecting the three-dimensional detection frames corresponding to each three-dimensional target onto the two-dimensional image plane corresponding to the monocular vision sensor; The performing feature extraction on the two-dimensional distance image to obtain two-dimensional spatial data features includes: Performing target recognition on the two-dimensional distance image to obtain two-dimensional projection frames corresponding to each three-dimensional target; determining the two-dimensional spatial data features according to the two-dimensional projection frames corresponding to each three-dimensional target; the two-dimensional projection frames corresponding to each three-dimensional target are two-dimensional projection frames of the three-dimensional detection frames corresponding to each three-dimensional target on the two-dimensional image plane.
5. The method according to claim 3, characterized in that, The two-dimensional image detection features include the positions and sizes of two-dimensional detection frames of each two-dimensional target in the multiple two-dimensional targets; The performing feature extraction on the two-dimensional visual image to obtain two-dimensional image detection features includes: The two-dimensional visual image is subjected to feature extraction through a two-dimensional feature encoding module to obtain a first target feature; the two-dimensional feature encoding module includes a backbone network, a feature pyramid network, and a path aggregation network. The backbone network is used to perform feature extraction on the two-dimensional visual image to obtain a first feature; the feature pyramid network is used to perform semantic enhancement processing on the first feature to obtain a second feature; the path aggregation network is used to perform semantic supplementation on the second feature to obtain the first target feature; The first target feature is subjected to object detection through an object detection head to obtain a two-dimensional detection box corresponding to each two-dimensional object; according to the two-dimensional detection box corresponding to each two-dimensional object, the two-dimensional image detection feature is determined.
6. The method according to claim 3, wherein The performing fusion processing on the spliced target feature to obtain the fusion feature includes: Feature extraction is performed on the spliced target feature through grouped convolution with a preset size to obtain a second target feature; the number of groups of the grouped convolution is equal to half of the number of feature channels of the spliced target feature; Feature fusion processing is performed on the second target feature through a convolutional network with a preset depth to obtain the fusion feature.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: During the visual and point cloud joint annotation process, the three-dimensional detection boxes corresponding to the three-dimensional objects with the occlusion attribute being a specified attribute are removed, and the remaining three-dimensional objects are matched and fused with the two-dimensional detection boxes corresponding to each two-dimensional object in the multiple two-dimensional objects to obtain the visual and point cloud joint annotation data corresponding to each two-dimensional object.
8. A monocular vision occlusion property recognition device, characterized in that, including: An acquisition module, configured to acquire a two-dimensional visual image of a vehicle environment through a monocular vision sensor, and acquire three-dimensional point cloud data of the vehicle environment through an in-vehicle radar; The two-dimensional visual image includes visual information of multiple two-dimensional objects; The three-dimensional point cloud data includes distance information of multiple three-dimensional objects relative to the in-vehicle radar; A fusion module, configured to convert the three-dimensional point cloud data into a two-dimensional distance image, and perform fusion processing on the two-dimensional distance image and the two-dimensional visual image to obtain a fusion feature; An attribute classification module, configured to classify the fusion feature to obtain the occlusion attribute of each three-dimensional object in the multiple three-dimensional objects in the two-dimensional visual image; the occlusion attribute is used to remove the three-dimensional detection boxes corresponding to the three-dimensional objects with the occlusion attribute being a specified attribute during the visual and point cloud joint annotation process.
9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.