Environment perception method and device, storage medium and electronic equipment
By acquiring multiple environmental perspective images using a surround-view fisheye camera and extracting and transforming various types of features, combined with a multi-task perception model, the shortcomings of LiDAR and surround-view cameras in environmental perception are solved, achieving more accurate and comprehensive environmental perception.
Patent Information
- Application Number
- CN202410645664.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, the lack of rich texture and semantic information in lidar sensors leads to inaccurate perception results, while surround-view cameras cannot accurately perceive the environment under severe distortion conditions.
Multiple environmental perspective images are acquired by a surround-view fisheye camera, and various types of features are extracted, dimensionality is transformed, and feature transformation is performed. Feature recognition is then performed using a multi-task perception model to obtain drivable area segmentation, ground semantic segmentation, and target object detection results.
It improves the perception accuracy and comprehensiveness of images acquired by the surround-view fisheye camera, providing accurate and comprehensive environmental perception results.
Smart Images

Figure CN121010936A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an environmental sensing method, device, storage medium, and electronic device. Background Technology
[0002] With the development of autonomous driving technology, environmental perception systems have become a core component of autonomous vehicles and other mobile robots. Environmental perception involves using sensors and perception algorithms to acquire and understand information about the surrounding environment so that robots can accurately perceive, understand, and respond to various complex road and traffic situations.
[0003] During robot operation, environmental perception systems typically include various sensors such as cameras, lidar, ultrasonic sensors, and radar, used to detect obstacles, identify road markings, and measure distances and speeds. The data acquired by these sensors is processed and fused to provide the robot with rich information about its surroundings, including road conditions, obstacle locations, and pedestrian behavior.
[0004] To achieve efficient environmental perception, robotic systems also need to be equipped with advanced perception algorithms, such as object detection and tracking, semantic segmentation, and Simultaneous Localization and Mapping (SLAM). These algorithms can analyze and interpret the data acquired by sensors in real time, enabling the robot to make corresponding decisions and actions, ensuring safe and efficient navigation in complex and changing environments. Summary of the Invention
[0005] This application provides an environmental sensing method, device, computer storage medium, and electronic device, the technical solutions of which are as follows:
[0006] In a first aspect, embodiments of this application provide an environmental perception method applied to a drivable device, the method comprising:
[0007] Multiple environmental perspective images are acquired by using a panoramic fisheye camera, and various types of feature extraction processing is performed on the multiple environmental perspective images to obtain the image category feature data corresponding to the environmental perspective images.
[0008] The bird's-eye view feature data is obtained by performing dimensional transformation on the image type feature data;
[0009] The bird's-eye view feature data is subjected to feature transformation processing to obtain bird's-eye view hierarchical feature data;
[0010] The bird's-eye view hierarchical feature data is processed by a multi-task perception model to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. The multi-task perception model is trained based on the sample drivable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels corresponding to the sample bird's-eye view hierarchical feature data.
[0011] Secondly, embodiments of this application provide an environmental sensing device applied to a drivable device, the device comprising:
[0012] The feature extraction module is used to acquire multiple environmental perspective images through a panoramic fisheye camera, and perform various types of feature extraction processing on the multiple environmental perspective images to obtain the image type feature data corresponding to the environmental perspective images.
[0013] The feature transformation module is used to perform dimensional transformation on image type feature data to obtain bird's-eye view feature data.
[0014] The feature transformation module is used to perform feature transformation processing on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data.
[0015] The feature recognition module is used to perform feature recognition processing on the bird's-eye view hierarchical feature data using a multi-task perception model to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results; wherein, the multi-task perception model is obtained after training based on the sample drivable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels corresponding to the sample bird's-eye view hierarchical feature data.
[0016] Thirdly, embodiments of this application provide a computer storage medium having multiple instructions adapted for loading and executing the methods described above by a processor.
[0017] Fourthly, embodiments of this application provide an electronic device, which may include: a memory and a processor; wherein the memory stores a computer program adapted to be loaded by the memory and to execute the above-described method.
[0018] The beneficial effects of the technical solutions provided in this application include at least the following:
[0019] In this embodiment, multiple environmental perspective images are acquired using a surround-view fisheye camera. Then, various feature extraction processes are performed on these images to obtain image category feature data corresponding to each environmental perspective image. Next, the image category feature data undergoes dimensionality transformation to obtain bird's-eye view feature data. This bird's-eye view feature data is then subjected to feature transformation to obtain bird's-eye view hierarchical feature data. Finally, a multi-task perception model is used to perform feature recognition processing on the bird's-eye view hierarchical feature data to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. By employing the above technical solution, various feature extraction processes are performed on environmental perspective images acquired from multiple perspectives using a fisheye camera to obtain image category feature data containing multiple types of features. This makes the bird's-eye view feature data obtained from the image category feature data more accurate. Furthermore, the rich and accurate bird's-eye view feature data is transformed into bird's-eye view hierarchical feature data with multi-level features. Multi-task perception is then performed on this multi-level feature data to obtain accurate and comprehensive recognition results. Therefore, this embodiment improves the accuracy and comprehensiveness of the perception results of drivable devices based on images acquired by surround-view fisheye cameras. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a scene diagram illustrating an environmental perception method provided in an embodiment of this application;
[0022] Figure 2 This is a flowchart illustrating an environmental perception method provided in an embodiment of this application;
[0023] Figure 3 This is a flowchart illustrating another environmental perception method provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of the structure of an environmental sensing device provided in an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the structure of a feature conversion module provided in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To make the inventive objectives, features, and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0029] In related technologies, environmental perception plays a crucial role in robot navigation, providing robots with a comprehensive understanding and perception of their surroundings, and offering key support for safe and intelligent autonomous driving. Currently, there are two main approaches to environmental perception in robotics: one involves using specific LiDAR sensors to collect data and perform detection and segmentation tasks based on that data; the other uses surround-view camera sensors to collect data, performing detection and segmentation tasks on each camera's data, and then fusing the results from multiple cameras to obtain the perceived environment around the robot. For the first approach, while LiDAR sensors can easily and directly measure spatial distances with high reliability, task perception based solely on LiDAR sensor data lacks rich texture and semantic information, resulting in inaccurate perception results. The second approach does not consider camera distortion issues with surround-view cameras, leading to inaccurate perception results under severe distortion. Therefore, accurately perceiving the surrounding environment of a drivable device is a pressing technical problem that needs to be solved.
[0030] To address the aforementioned technical problems, the environmental perception method of this application will be further described below with reference to specific embodiments.
[0031] Please see Figure 1 This is a scene diagram illustrating an environmental perception method provided in an embodiment of this application.
[0032] like Figure 1 As shown, it illustrates a scene diagram illustrating the acquisition of environmental images from different perspectives by a drivable device from a top-down view. Figure 1 In this embodiment, the surround-view fisheye camera can be installed on the front, back, left, and right sides of the drivable device to capture environmental images from the front, back, left, and right sides of the drivable device, respectively. In other embodiments of this application, the number of surround-view fisheye cameras installed on the drivable device is not limited.
[0033] In this embodiment, the drivable device can be a mobile robot, such as a cleaning robot or a service robot. The drivable device can also be an autonomous vehicle or similar device.
[0034] In this embodiment, multiple environmental perspective images are acquired using a surround-view fisheye camera. Then, various feature extraction processes are performed on these images to obtain image category feature data corresponding to each environmental perspective image. Next, the image category feature data undergoes dimensionality transformation to obtain bird's-eye view feature data. This bird's-eye view feature data is then subjected to feature transformation to obtain bird's-eye view hierarchical feature data. Finally, a multi-task perception model is used to perform feature recognition processing on the bird's-eye view hierarchical feature data to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. By employing the above technical solution, various feature extraction processes are performed on environmental perspective images acquired from multiple perspectives using a fisheye camera to obtain image category feature data containing multiple types of features. This makes the bird's-eye view feature data obtained from the image category feature data more accurate. Furthermore, the rich and accurate bird's-eye view feature data is transformed into bird's-eye view hierarchical feature data with multi-level features. Multi-task perception is then performed on this multi-level feature data to obtain accurate and comprehensive recognition results. Therefore, this embodiment improves the accuracy and comprehensiveness of the perception results of drivable devices based on images acquired by surround-view fisheye cameras.
[0035] In the following method embodiments, each step is described in detail with electronic devices as the executing entity.
[0036] Please see Figure 2 This is a flowchart illustrating an environmental perception method provided in an embodiment of this application. Figure 2As shown, the method in this application embodiment may include the following steps:
[0037] S201 acquires multiple environmental perspective images through a panoramic fisheye camera, and performs various feature extraction processes on the multiple environmental perspective images to obtain image category feature data corresponding to the environmental perspective images.
[0038] As is easily understood, a surround-view fisheye camera refers to a fisheye camera arranged around an electronic device. A fisheye camera is a camera with a fisheye lens, which is a lens with an extremely short focal length and a field of view close to or equal to 180 degrees. Optionally, in embodiments of this application, one surround-view fisheye camera may be arranged on each of the four sides (front, back, left, and right) of the electronic device.
[0039] Environmental perspective images refer to environmental images captured by each panoramic fisheye camera from its own perspective.
[0040] Image category feature data refers to a variety of feature data, including feature data of image category, feature data of depth category, and feature data of semantic category. Image category feature data can refer to feature data collected through image feature extraction from environmental perspective images. Depth category feature data can refer to feature data obtained through depth prediction from environmental perspective images. Semantic category feature data can refer to feature data obtained through semantic recognition from environmental perspective images.
[0041] In some embodiments, environmental perspective images can be acquired by surround-view fisheye cameras arranged on different sides, each capturing an image within its field of view. For each environmental perspective image, image feature extraction, depth prediction, and semantic recognition can be performed simultaneously to obtain image category feature data corresponding to that environmental perspective image. Specifically, by simultaneously performing image feature extraction, depth prediction, and semantic recognition on the environmental perspective image, image category feature data, depth category feature data, and semantic category feature data can be obtained, respectively. Combining these feature data yields the image category feature data for that environmental perspective image. The image category feature data may include high-level image features, such as edge features, texture features, and corner features. The depth category feature data may include the depth value of each pixel in the image. The semantic category feature data may include the name or category of objects present in the image.
[0042] S202, perform dimensional transformation on the image type feature data to obtain bird's-eye view feature data.
[0043] In simple terms, bird's-eye view feature data refers to feature data from a bird's-eye view perspective.
[0044] In some embodiments, two-dimensional image category feature data can be converted into three-dimensional target image feature data using the fisheye camera parameters of a panoramic fisheye camera. Then, the three-dimensional target image feature data can be converted into feature data from a bird's-eye view, i.e., bird's-eye view feature data. Specifically, converting two-dimensional image category feature data into three-dimensional target image feature data using the fisheye camera parameters of a panoramic fisheye camera can involve converting the pixel coordinates of pixels in the environmental view image into three-dimensional coordinates in the world coordinate system using the fisheye camera parameters, and then using these three-dimensional coordinates to construct the three-dimensional target image feature data. The fisheye camera parameters may include parameters such as focal length, principal point coordinates, distortion parameters, and extrinsic parameters.
[0045] S203, perform feature transformation processing on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data.
[0046] Bird's-eye view layer feature data can be understood as feature data with more layers of image features under a bird's-eye view perspective, compared to bird's-eye view feature data.
[0047] Specifically, a deep convolutional neural network (DCNN) can be used to perform feature transformation on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data. The DCNN can be pre-trained, and the training process can be as follows: create an initial DCNN for the bird's-eye view feature transformation scenario; acquire sample bird's-eye view feature data, which can be feature data obtained by dimensional transformation of sample image type feature data, or feature data obtained by multi-class feature extraction of sample environmental view images captured by a surround-view fisheye camera; label the sample bird's-eye view feature data with corresponding sample bird's-eye view hierarchical feature data tags; input the sample bird's-eye view feature data into the initial DCNN for at least one round of training to obtain predicted bird's-eye view hierarchical feature data; calculate the loss value using a loss function based on the predicted bird's-eye view hierarchical feature data and the sample bird's-eye view hierarchical feature data tags; adjust the model parameters of the initial DCNN based on the loss value until the training termination condition is met to obtain the deep convolutional neural network.
[0048] Optionally, this application embodiment can use a Residual Network (ResNet) as the deep convolutional neural network. ResNet performs convolution, pooling, and other operations on the input bird's-eye view feature data layer by layer to extract features at different levels. ResNet can extract features at different levels by stacking multiple residual blocks. Earlier residual blocks are used to extract low-level features, while deeper residual blocks gradually extract higher-level features. For example, N is the total depth of ResNet, N / 8 represents 1 / 8 of the total network depth, and N / 4 represents 4 / 8 of the total network depth. From the perspective of layer level, the N / 4 layer is higher than the N / 8 layer. The shallower layers (N / 8 layers) extract features that are more biased towards low-level features such as edges, colors, and textures. These low-level features are used to identify basic elements and details in the image. The deeper layers (N / 4 layers) extract higher-level, more abstract features, which are used to understand the overall structure and semantic content of the image. Therefore, the bird's-eye view hierarchical feature data obtained through deep convolutional neural networks is a fusion of low-level and high-level features, which can provide rich image information for subsequent task perception tasks.
[0049] S204 uses a multi-task perception model to process the feature data of the bird's-eye view layer, and obtains the drivable area segmentation result, ground semantic segmentation result, and target object detection result.
[0050] In simple terms, drivable area segmentation refers to the segmentation of an environmental view image into drivable and non-drivable areas. Drivable areas are those where there are no obstacles on the ground and drivability is possible, while non-drivable areas are those where there are obstacles on the ground and drivability is impossible.
[0051] Ground semantic segmentation results refer to the segmentation results of dividing the ground in an environmental perspective image into elements such as lane lines, parking lines, and speed bumps.
[0052] The target object detection result refers to the detection results of the spatial position, bounding box size, and rotation angle of the target object in the environmental view image. The spatial position of the target object includes the pixel coordinates of its center point, the midpoint of the head side of the bounding box corresponding to the target object, and the midpoint of the tail side of the bounding box corresponding to the target object. The bounding box size refers to the length, width, and height of the 3D bounding box surrounding the target object. The rotation angle of the target object refers to its rotation angle in the world coordinate system, that is, its rotation angle relative to the electronic device.
[0053] Specifically, the bird's-eye view hierarchical feature data can be input into the multi-task perception model to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. The multi-task perception model in this embodiment can be pre-trained, and the training process of the multi-task perception model can include the following steps:
[0054] Model creation: Create an initial multi-task perception model for multi-task perception scenarios based on machine learning models;
[0055] Acquiring Sample Data: Acquiring a large amount of sample data, which is sample bird's-eye view layer feature data obtained by extracting and transforming features from a large number of sample environmental perspective images collected by a surround-view fisheye camera;
[0056] Labeling sample data: Based on the needs of multi-task perception scenarios, an expert service is introduced to manually label the sample data with corresponding sample labels. The sample labels include the sample drivable area segmentation result label, the sample ground semantic segmentation result label, and the sample target object detection result label for each sample data.
[0057] Model training process: Input sample data into the initial multi-task perception model for at least one round of model training to obtain prediction result data. The prediction result data includes predicted drivable area segmentation results, predicted ground semantic segmentation results, and predicted target object detection results. Based on the prediction result data (predicted drivable area segmentation results, predicted ground semantic segmentation results, and predicted target object detection results) and sample data labels (drivable area segmentation result labels, sample ground semantic segmentation result labels, and sample target object detection result labels), the model loss function is used to determine the model loss value. Based on the model loss value, the model parameters of the initial multi-task perception model are adjusted until the model training termination condition is met to obtain the multi-task perception model.
[0058] Optionally, the model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0059] It should be noted that the machine learning models involved in one or more embodiments of this application include, but are not limited to, fitting of one or more of the following machine learning models: Convolutional Neural Network (CNN) model, Deep Neural Network (DNN) model, Recurrent Neural Networks (RNN) model, embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.
[0060] In this embodiment, multiple environmental perspective images are acquired using a surround-view fisheye camera. Then, various feature extraction processes are performed on these images to obtain image category feature data corresponding to each environmental perspective image. Next, the image category feature data undergoes dimensionality transformation to obtain bird's-eye view feature data. This bird's-eye view feature data is then subjected to feature transformation to obtain bird's-eye view hierarchical feature data. Finally, a multi-task perception model is used to perform feature recognition processing on the bird's-eye view hierarchical feature data to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. By employing the above technical solution, various feature extraction processes are performed on environmental perspective images acquired from multiple perspectives using a fisheye camera to obtain image category feature data containing multiple types of features. This makes the bird's-eye view feature data obtained from the image category feature data more accurate. Furthermore, the rich and accurate bird's-eye view feature data is transformed into bird's-eye view hierarchical feature data with multi-level features. Multi-task perception is then performed on this multi-level feature data to obtain accurate and comprehensive recognition results. Therefore, this embodiment improves the accuracy and comprehensiveness of the perception results of drivable devices based on images acquired by surround-view fisheye cameras.
[0061] Please see Figure 3 This is a flowchart illustrating an environmental perception method provided in an embodiment of this application. Figure 3 As shown, the method in this application embodiment may include the following steps:
[0062] S301 acquires multiple environmental perspective images through a surround-view fisheye camera, and performs image feature extraction processing on the environmental perspective images to obtain the image feature data corresponding to the environmental perspective images.
[0063] For a detailed explanation of panoramic fisheye camera and environmental perspective images, please refer to [link / reference]. Figure 2 The description of S201 in the illustrated embodiment will not be repeated here.
[0064] Image feature data refers to feature data that includes features such as brightness, edges, texture, shape, and color of an image.
[0065] In some embodiments, environmental view images can be acquired by surround-view fisheye cameras arranged on different sides, each capturing an image within its field of view. For each environmental view image, a ResNet network and a feature pyramid are used to extract image features, yielding the corresponding image feature data. The ResNet network is a deep residual network that achieves layer-by-layer feature extraction and representation by stacking multiple residual modules. Specifically, the ResNet network is first used to extract features at multiple scales from the environmental view image, and then a feature pyramid is used to fuse the features extracted by the ResNet network at multiple scales to obtain the image feature data. For example, in an environmental view image with a pixel length of N, a ResNet network can be used to extract features at levels N / 2, N / 4, N / 8, and N / 16, among others. Further, a feature pyramid is used for low-level feature fusion to obtain fused N / 16-level features. These fused N / 16-level features are then converted to N / 8-level features. The converted N / 8-level features are then fused with the N / 8-level features. The fused N / 8-level features are then converted to N / 4-level features. The converted N / 4-level features are then fused with the N / 4-level features. Finally, the fused N / 4-level features are fused at higher levels, and so on, to obtain the image feature data of the environmental view image.
[0066] S302, perform depth prediction processing on the environmental perspective image to obtain the depth prediction data corresponding to the environmental perspective image.
[0067] In some embodiments, step S302 may specifically involve: obtaining the intrinsic parameters of the fisheye camera from the panoramic fisheye camera; and performing depth prediction processing on the environmental view image based on the fisheye camera intrinsic parameters to obtain the depth prediction data corresponding to the environmental view image.
[0068] Among them, the intrinsic parameters of a fisheye camera refer to parameters such as focal length, principal point coordinates, and distortion parameters.
[0069] Depth prediction data refers to the depth value of each pixel in an environmental view image.
[0070] Obtaining the intrinsic parameters of a panoramic fisheye camera can be understood as reading the intrinsic parameters from the camera parameter configuration file. In addition to storing the intrinsic parameters, the camera parameter configuration file can also store the extrinsic parameters, including parameters such as the extrinsic parameter matrix.
[0071] Depth prediction data for environmental perspective images is obtained by performing depth prediction processing based on fisheye camera intrinsic parameters. This can be understood as follows: the fisheye camera intrinsic parameters and the environmental perspective image are input into a depth prediction model to obtain the depth interval category label for each pixel in the environmental perspective image. The depth prediction data for the environmental perspective image is then obtained based on the depth interval category label for each pixel. The depth prediction model is trained based on the sample fisheye camera intrinsic parameters and the sample depth interval category labels for the sample environmental perspective images. The depth interval category label refers to the category label of the target depth interval into which the pixel's depth value falls among multiple defined depth intervals. The category label can be used to identify different depth intervals and can be determined based on the position of the depth interval and the range of depth values represented by the depth interval. Therefore, by predicting the depth interval category label for each pixel in the environmental perspective image through the depth prediction model, the depth value of each pixel can be determined.
[0072] Alternatively, the method for dividing the depth range can be to divide the area within 20 meters in front of the camera into multiple depth ranges using small grids of 0.2 meters each, with each depth range corresponding to a different range of depth values.
[0073] Optionally, the deep prediction model can be trained based on a deep learning model, which may include, but is not limited to, perceptrons, multilayer perceptrons, convolutional neural networks, recurrent neural networks, etc.
[0074] In this embodiment of the application, when predicting the depth value of each pixel, the prediction is not only based on the environmental perspective image, but also takes into account the camera's distortion parameters, so as to achieve a more accurate prediction of the depth value of each pixel.
[0075] S303, perform semantic segmentation on the environmental perspective image to obtain the semantic recognition data corresponding to the environmental perspective image.
[0076] In simple terms, semantic recognition data refers to semantic data that includes the names or categories of objects present in an image.
[0077] In some embodiments, semantic segmentation can be performed on each environmental perspective image using a semantic recognition model to obtain semantic recognition data corresponding to each environmental perspective image. The semantic recognition model can be trained based on the sample semantic recognition data labels corresponding to the sample environmental perspective images. The sample environmental perspective images can include sample environmental perspective images from different perspectives to enhance the semantic recognition capability of the semantic recognition model. The semantic recognition model can identify not only dynamic objects in the image but also static objects.
[0078] This application embodiment predicts the corresponding semantic results for each environmental perspective image captured by a surround-view fisheye camera individually, thereby increasing the semantic meaning of the image's feature data.
[0079] S304, determine the image type feature data corresponding to the environmental perspective image based on image feature data, depth prediction data, and semantic recognition data.
[0080] In some embodiments, for each environmental perspective image, the image feature data, depth prediction data, and semantic recognition data can be spliced or combined to obtain the image category feature data corresponding to that environmental perspective image.
[0081] The image type feature data obtained through the embodiments of this application contains rich feature data, and also provides more accurate feature data for subsequent perception tasks.
[0082] S305, retrieves the fisheye camera parameters of the panoramic fisheye camera.
[0083] In simple terms, fisheye camera parameters include intrinsic and extrinsic parameters. Intrinsic parameters include focal length, principal point coordinates, distortion parameters, etc. Extrinsic parameters may include an extrinsic parameter matrix, which contains the camera's position and orientation in three-dimensional coordinate space.
[0084] Specifically, the intrinsic and extrinsic parameters of the fisheye camera can be read from the camera parameter configuration file.
[0085] S306 performs coordinate system feature transformation on the image type feature data based on fisheye camera parameters to obtain the target image feature data corresponding to the environmental perspective image.
[0086] When performing step S406, the specific steps can be as follows:
[0087] A1: Obtain the target pixel corresponding to the target feature data in the image type feature data, and determine the pixel coordinates of the target pixel in the environmental view image;
[0088] For each image category feature data, it can contain feature data of multiple pixels. Each pixel's feature data can be a set of feature data, which may include image feature data, depth prediction data, and semantic recognition data. A set of feature data can be selected from the image category features as the target feature data, and the pixel corresponding to this target feature data is designated as the target pixel. The pixel coordinates of the target pixel are determined based on its position in the environmental view image. For example, if the target pixel is located at row 3, column 4 in the environmental view image, its pixel coordinates could be (3,4).
[0089] A2: Based on the intrinsic parameters of the fisheye camera, perform camera coordinate transformation on the pixel coordinates to obtain the first coordinates of the target pixel in the camera coordinate system;
[0090] When performing step A2, the specific steps can be:
[0091] a1: Obtain the focal length, principal point coordinates, and distortion parameters of the panoramic fisheye camera;
[0092] Focal length can include both horizontal and vertical focal lengths. The horizontal focal length can be represented as fx, and the vertical focal length can be represented as fy.
[0093] Principal point coordinates refer to the coordinate position of the principal point in the camera's optical system on the image plane. The principal point is the point on the image plane that intersects the camera's optical axis, usually represented as (cx, cy). cx represents the horizontal coordinate position of the principal point on the image plane, which determines the horizontal distance between the image center and the image edge. cy represents the vertical coordinate position of the principal point on the image plane, which determines the vertical distance between the image center and the image edge.
[0094] The distortion parameters can include four distortion parameters. These distortion parameters can be represented as k1, k2, k3, and k4.
[0095] a2: Based on the focal length, principal point coordinates, distortion parameters, and pixel coordinates, perform camera coordinate transformation to obtain the first coordinates of the target pixel in the camera coordinate system.
[0096] Specifically, calculating the first coordinate also requires obtaining the depth value corresponding to the target pixel. The depth value of the target pixel can be represented as d.
[0097] The first coordinate can be calculated using the first coordinate transformation formula.
[0098] The first coordinate transformation formula satisfies the following formula:
[0099]
[0100] The first coordinate can be represented by x, y, z in the above formula as (x, y, z);
[0101] u represents the horizontal pixel coordinate in the pixel coordinate system, cx represents the horizontal coordinate in the principal point coordinate system, and fx represents the focal length in the horizontal direction. v represents the vertical pixel coordinate in the pixel coordinate system, cy represents the vertical coordinate in the principal point coordinate system, fy represents the focal length in the vertical direction, and d represents the depth value of the target pixel.
[0102] The calculation of θ follows these steps: Step 1, assign θ the value 0; Step 2, substitute θ = 0 into the right side of the first calculation formula to obtain the first result; Step 3, substitute the first result into the right side of the first calculation formula to obtain the second result; Step 4, substitute the second result into the right side of the first calculation formula to obtain the third result; Step 5, substitute the third result into the right side of the first calculation formula to obtain the fourth result; Step 6, substitute the fourth result into the right side of the first calculation formula to obtain the fifth result, and assign the fifth result to θ. The first calculation formula satisfies the following formula:
[0103]
[0104] Where k1, k2, k3, and k4 represent the aforementioned distortion parameters.
[0105] A3: Based on the fisheye camera extrinsic parameters, the first coordinate is transformed into world coordinates to obtain the second coordinate of the target pixel in the world coordinate system;
[0106] The extrinsic parameters of a fisheye camera refer to the extrinsic parameter matrix, which includes a rotation matrix and a translation vector. The rotation matrix is a 3x3 matrix, and the translation vector is a 3x1 matrix. The extrinsic parameter matrix can be represented as [R|T], where R represents the rotation matrix and T represents the translation vector.
[0107] The first coordinate is transformed to obtain the first matrix. The first matrix is then multiplied on the left by the extrinsic parameter matrix to obtain the second coordinate of the target pixel in world coordinates. The second coordinate can be denoted as (xc, yc, zc), then [xc yc zc] = [R|T][xyc 1]. Here, [xyc 1] represents the first matrix.
[0108] In this embodiment of the application, when the first coordinates of the target pixel are upgraded to obtain the second three-dimensional coordinates, the distortion parameters of the camera are taken into account to achieve accurate mapping of the pixel coordinates to the three-dimensional coordinates in the three-dimensional space.
[0109] A4: Determine the target image feature data corresponding to the environmental perspective image based on the second coordinate and target feature data.
[0110] Specifically, since the target feature data is the feature data corresponding to the target pixel, there is a two-dimensional feature correspondence between the pixel coordinates of the target pixel in the camera coordinate system and the target feature data. When converting the pixel coordinates of the target pixel into three-dimensional second coordinates, the two-dimensional correspondence between the pixel coordinates and the target feature data can be converted into a three-dimensional feature correspondence between the two-dimensional second coordinates and the target feature data. For each target pixel, there is a corresponding three-dimensional feature correspondence. Thus, the three-dimensional feature correspondences corresponding to all target pixels can be obtained. The target image feature data can be obtained by associating the target feature data and the second coordinates with all the three-dimensional feature correspondences.
[0111] S307, Perform bird's-eye view feature transformation on the target image feature data to obtain bird's-eye view feature data.
[0112] It is easy to understand that the target image feature data is three-dimensional feature data.
[0113] When performing step S307, the specific steps can be as follows:
[0114] B1: Highly compressed feature data is obtained by performing high-dimensional feature compression on the target image feature data through an attention mechanism; B2: The orientation matrix and depth distance matrix are determined; B3: Bird's-eye view feature transformation is performed on the highly compressed feature data, orientation matrix, and depth distance matrix to obtain bird's-eye view feature data.
[0115] In step B1, an attention weight is determined for each target image feature data using an attention mechanism. The target image feature data is then weighted using these attention weights to obtain weighted image feature data. This weighted image feature data is then weighted and averaged at the height of the target image feature data to obtain height-compressed feature data. The height of the target image feature data is based on the pixel height in the corresponding environmental viewpoint data, which can be measured by a sensor. This weighting of the target image feature data allows feature data with higher attention weights to receive more attention. Since the drivable area segmentation results, ground semantic segmentation results, and target object detection results obtained through the method of this embodiment can be used for subsequent planning and control of the electronic device, and the subsequent planning and control of the electronic device in this embodiment is not sensitive to height information, the feature data can be compressed using height information to obtain two-dimensional feature data from a bird's-eye view.
[0116] In step B2, for each environmental view image, its image width is represented by w, and its image height by h. The depth values between 0.5 meters and 20 meters are evenly divided into 113 grids, the number of grids being d_num. Each grid represents a range of depth values, and the fixed point of each grid represents the reference point of that grid in terms of depth distance. The above fixed points form a reference point set, which contains the fixed points of all the above grids, and is represented by geom_sep. The coordinates of (w*h*d_num) in the world coordinate system are calculated through the above steps A1, A2, and A3. A matrix with d_num rows and L*L columns is used as the direction matrix. The direction matrix is initialized by setting all elements in the direction matrix to 0. Then, dir is used as the index, with the value of dir ranging from 0 to Nc*w, where Nc represents the number of fisheye cameras. The direction matrix is updated using the first update formula: Ray[i,geom_sep[dir]]+=1, where Ray represents the direction matrix. Use a matrix with Nc*w rows and L*L columns as the depth distance matrix, with L being 128. Initialize the depth distance matrix by setting all elements in the direction of the depth distance matrix to 0. Then, use dd as the index, with the value of dd ranging from 0 to d_num. Update the depth distance matrix using the second update formula, which is: Circle[dd,geom_sep[:,dd]]+=1, where Circle represents the depth distance matrix.
[0117] In step B3, the second calculation formula is used to calculate the highly compressed feature data, orientation matrix, and depth distance matrix to obtain the bird's-eye view feature data. The second calculation formula satisfies the following formula:
[0118]
[0119] Where F represents the bird's-eye view feature data, depth represents the depth measurement data obtained in step S302, Circle represents the depth distance matrix, Ray represents the direction matrix, f represents the height compressed feature data, and × represents matrix multiplication. This represents the Hadamard product operation.
[0120] S308 uses a feature extraction backbone network to perform hierarchical feature parsing on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data.
[0121] In simple terms, bird's-eye view hierarchical feature data refers to feature data with higher-level features from a bird's-eye view perspective, relative to bird's-eye view feature data.
[0122] In some embodiments, bird's-eye view feature data is input into a feature extraction backbone network for hierarchical feature parsing to obtain bird's-eye view hierarchical feature data. The feature extraction backbone network can be a network trained on a ResNet network, and can be trained using the labels of the sample bird's-eye view hierarchical feature data corresponding to the sample bird's-eye view feature data. The feature extraction backbone network can perform convolution, pooling, and other operations on the bird's-eye view feature data layer by layer to extract higher-level feature data from more bird's-eye view perspectives. The training process of the feature extraction backbone network can be as follows: First, create an initial feature extraction backbone network based on the ResNet network for bird's-eye view feature transformation scenarios. This initial network can be used to obtain sample bird's-eye view feature data, which can be obtained by dimensional transformation of sample image category feature data or by multi-category feature extraction of sample environmental view images captured by a surround-view fisheye camera. Second, label the sample bird's-eye view feature data with corresponding sample bird's-eye view hierarchical feature data labels. Third, input the sample bird's-eye view feature data into the initial feature extraction backbone network for at least one round of training to obtain predicted bird's-eye view hierarchical feature data. Fourth, calculate the loss value using a loss function based on the predicted bird's-eye view hierarchical feature data and the sample bird's-eye view hierarchical feature data labels. Fifth, adjust the model parameters of the initial feature extraction backbone network based on the loss value until the training termination condition is met, thus obtaining the feature extraction backbone network.
[0123] By extracting higher-level feature data from a bird's-eye view, we can subsequently perceive more accurate and comprehensive results in drivable area segmentation, ground semantic segmentation, and target object detection.
[0124] S309 uses a multi-task perception model to process the feature data of the bird's-eye view layer, and obtains the drivable area segmentation result, ground semantic segmentation result, and target object detection result.
[0125] Driving area segmentation results refer to the segmentation results of dividing the environmental view image into driving areas and non-driving areas. Driving areas are areas on the ground where there are no obstacles and driving is possible, while non-driving areas are areas on the ground where there are obstacles and driving is impossible.
[0126] Ground semantic segmentation results refer to the segmentation results of dividing the ground in an environmental perspective image into elements such as lane lines, parking lines, and speed bumps.
[0127] The target object detection result refers to the detection results of the spatial position, bounding box size, and rotation angle of the target object in the environmental view image. The spatial position of the target object includes the pixel coordinates of its center point, the midpoint of the head side of the bounding box corresponding to the target object, and the midpoint of the tail side of the bounding box corresponding to the target object. The bounding box size refers to the length, width, and height of the 3D bounding box surrounding the target object. The rotation angle of the target object refers to its rotation angle in the world coordinate system, that is, its rotation angle relative to the electronic device.
[0128] The multi-task perception model is trained based on the sample drivable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels corresponding to the sample bird's-eye view hierarchical feature data. The training process of the multi-task perception model in this embodiment may include the following steps:
[0129] Model creation: Create an initial multi-task perception model for multi-task perception scenarios based on machine learning models;
[0130] Acquiring Sample Data: Acquiring a large amount of sample data, which is sample bird's-eye view layer feature data obtained by extracting and transforming features from a large number of sample environmental perspective images collected by a surround-view fisheye camera;
[0131] Labeling sample data: Based on the needs of multi-task perception scenarios, an expert service is introduced to manually label the sample data with corresponding sample labels. The sample labels include the sample drivable area segmentation result label, the sample ground semantic segmentation result label, and the sample target object detection result label for each sample data.
[0132] Model training process: Input sample data into the initial multi-task perception model for at least one round of model training to obtain prediction result data. The prediction result data includes predicted drivable area segmentation results, predicted ground semantic segmentation results, and predicted target object detection results. Based on the prediction result data (predicted drivable area segmentation results, predicted ground semantic segmentation results, and predicted target object detection results) and sample data labels (drivable area segmentation result labels, sample ground semantic segmentation result labels, and sample target object detection result labels), the model loss function is used to determine the model loss value. Based on the model loss value, the model parameters of the initial multi-task perception model are adjusted until the model training termination condition is met to obtain the multi-task perception model.
[0133] Optionally, the model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.
[0134] It should be noted that the machine learning models involved in one or more embodiments of this application include, but are not limited to, fitting of one or more of the following machine learning models: Convolutional Neural Network (CNN) model, Deep Neural Network (DNN) model, Recurrent Neural Networks (RNN) model, embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.
[0135] In this embodiment, multiple environmental perspective images are acquired using a surround-view fisheye camera. Image feature extraction is performed on these images to obtain corresponding image feature data. Depth prediction is then performed to obtain corresponding depth prediction data. Semantic segmentation is performed to obtain corresponding semantic recognition data. Based on the image feature data, depth prediction data, and semantic recognition data, image category feature data corresponding to the environmental perspective images is determined. Therefore, the image category feature data in this embodiment contains rich feature data and provides more accurate feature data for subsequent perception tasks. Furthermore, this embodiment also performs coordinate system feature transformation on the image category feature data based on fisheye camera parameters to obtain target image feature data corresponding to the environmental perspective images. Bird's-eye view feature transformation is then performed on the target image feature data to obtain bird's-eye view feature data. Feature transformation is then performed on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data. Finally, based on the bird's-eye view hierarchical feature data, drivable area segmentation results, ground semantic segmentation results, and target object detection results are identified. Before recognizing feature data from a bird's-eye view, this embodiment first extracts higher-level feature data based on the bird's-eye view feature data. This ensures that the feature data from the bird's-eye view input to the multi-task perception model has richer and more accurate feature data, resulting in more accurate recognition results. Furthermore, this embodiment simultaneously performs multi-task perception, achieving precise and comprehensive recognition results.
[0136] The following will combine Figure 4 This application provides a detailed description of the environmental sensing device provided in its embodiments. It should be noted that... Figure 4 The environmental sensing device shown is used to perform the functions described in this application. Figure 2 and Figure 3The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figure 2 and Figure 3 The example shown.
[0137] Please see Figure 4 This diagram illustrates the structure of an environmental sensing device according to an embodiment of this application. The environmental sensing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the environmental sensing device 1 includes a feature extraction module 11, a feature conversion module 12, a feature transformation module 13, and a feature recognition module 14, specifically used for:
[0138] Feature extraction module 11 is used to acquire multiple environmental perspective images through a panoramic fisheye camera, and perform multiple types of feature extraction processing on the multiple environmental perspective images to obtain image type feature data corresponding to the environmental perspective images;
[0139] Feature transformation module 12 is used to perform dimensional transformation processing on image type feature data to obtain bird's-eye view feature data;
[0140] Feature transformation module 13 is used to perform feature transformation processing on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data;
[0141] The feature recognition module 14 is used to perform feature recognition processing on the bird's-eye view hierarchical feature data using a multi-task perception model to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results; wherein, the multi-task perception model is obtained after training based on the sample drivable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels corresponding to the sample bird's-eye view hierarchical feature data.
[0142] Optional, please see Figure 5 The schematic diagram of the feature transformation module 12 shown above includes a parameter acquisition unit 121, a coordinate transformation unit 122, and a feature transformation unit 123, which are specifically used for:
[0143] The parameter acquisition unit 121 is used to acquire the fisheye camera parameters of the panoramic fisheye camera;
[0144] The coordinate transformation unit 122 is used to perform coordinate system feature transformation processing on the image type feature data based on the fisheye camera parameters to obtain the target image feature data corresponding to the environmental perspective image.
[0145] The feature conversion unit 123 is used to perform bird's-eye view feature conversion processing on the target image feature data to obtain bird's-eye view feature data.
[0146] Optionally, the coordinate transformation unit 122 includes:
[0147] The first conversion unit is used to obtain the target pixel corresponding to the target feature data in the image type feature data and determine the pixel coordinates of the target pixel in the environmental view image.
[0148] The second transformation unit is used to perform camera coordinate transformation on the pixel coordinates based on the fisheye camera intrinsic parameters to obtain the first coordinates of the target pixel in the camera coordinate system.
[0149] The third transformation unit is used to perform world coordinate transformation on the first coordinate based on the fisheye camera extrinsic parameters to obtain the second coordinate of the target pixel in the world coordinate system.
[0150] The fourth transformation unit is used to determine the target image feature data corresponding to the environmental perspective image based on the second coordinates and target feature data.
[0151] Optional, the second conversion unit is specifically used for:
[0152] Obtain the focal length, principal point coordinates, and distortion parameters of the panoramic fisheye camera;
[0153] Camera coordinate transformation is performed based on focal length, principal point coordinates, distortion parameters, and pixel coordinates to obtain the first coordinates of the target pixel in the camera coordinate system.
[0154] Optional, feature transformation unit, specifically used for:
[0155] Highly compressed feature data is obtained by performing high-dimensional feature compression on the target image feature data through an attention mechanism.
[0156] Determine the orientation matrix and depth-distance matrix;
[0157] Bird's-eye view feature data is obtained by performing bird's-eye view feature transformation on highly compressed feature data, orientation matrix, and depth distance matrix.
[0158] Optional, the feature extraction module includes:
[0159] The first extraction unit is used to perform image feature extraction processing on the environmental perspective image to obtain the image feature data corresponding to the environmental perspective image;
[0160] The second extraction unit is used to perform depth prediction processing on the environmental perspective image to obtain the depth prediction data corresponding to the environmental perspective image.
[0161] The third extraction unit is used to perform semantic segmentation processing on the environmental perspective image to obtain the semantic recognition data corresponding to the environmental perspective image.
[0162] The fourth extraction unit is used to determine the image type feature data corresponding to the environmental perspective image based on image feature data, depth prediction data, and semantic recognition data.
[0163] Optional, the second extraction unit is specifically used for:
[0164] Obtain the intrinsic parameters of the fisheye camera from the panoramic fisheye camera;
[0165] Depth prediction data corresponding to environmental perspective images is obtained by performing depth prediction processing on environmental perspective images based on fisheye camera intrinsic parameters.
[0166] Optional, feature transformation module, specifically used for:
[0167] The feature extraction backbone network is used to perform hierarchical feature parsing on the bird's-eye view feature data to obtain the bird's-eye view hierarchical feature data.
[0168] Please refer to Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device in this embodiment may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 can be connected via the bus 150.
[0169] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device via various interfaces and lines, and performs various functions and processes data of electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of the following: central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately through a communication chip.
[0170] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets.
[0171] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In this embodiment, the input device 130 can be a temperature sensor for acquiring the operating temperature of the electronic device. The output device 140 can be a speaker for outputting audio signals.
[0172] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the terminal. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WIFI) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0173] In the embodiments of this application, the executing entity for each step can be the electronic device described above. Optionally, the executing entity for each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this embodiment of the application does not limit this.
[0174] exist Figure 6 In the electronic device, the processor 110 can be used to call the program of the environmental perception method stored in the memory 120, and the processor 110 can be used to call the program of the environmental perception method stored in the memory 120, and specifically perform the following operations:
[0175] Multiple environmental perspective images are acquired by using a panoramic fisheye camera, and various types of feature extraction processing is performed on the multiple environmental perspective images to obtain the image category feature data corresponding to the environmental perspective images.
[0176] The bird's-eye view feature data is obtained by performing dimensional transformation on the image type feature data;
[0177] The bird's-eye view feature data is subjected to feature transformation processing to obtain bird's-eye view hierarchical feature data;
[0178] The bird's-eye view hierarchical feature data is processed by a multi-task perception model to obtain drivable area segmentation results, ground semantic segmentation results, and target object detection results. The multi-task perception model is trained based on the sample drivable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels corresponding to the sample bird's-eye view hierarchical feature data.
[0179] In one embodiment, when the processor 110 performs the step of performing dimensional transformation processing on the image type feature data to obtain bird's-eye view feature data, it specifically performs the following operations:
[0180] Obtain the fisheye camera parameters of the panoramic fisheye camera;
[0181] Based on the fisheye camera parameters, coordinate system feature transformation is performed on the image type feature data to obtain the target image feature data corresponding to the environmental perspective image.
[0182] Bird's-eye view feature data is obtained by performing bird's-eye view feature transformation on the target image feature data.
[0183] In one embodiment, when the processor 110 performs the step of performing coordinate system feature transformation processing on image type feature data based on fisheye camera parameters to obtain target image feature data corresponding to the environmental perspective image, it specifically performs the following operations:
[0184] Obtain the target pixel corresponding to the target feature data in the image type feature data, and determine the pixel coordinates of the target pixel in the environmental view image;
[0185] Based on the intrinsic parameters of the fisheye camera, the pixel coordinates are transformed by camera coordinates to obtain the first coordinates of the target pixel in the camera coordinate system.
[0186] Based on the fisheye camera extrinsic parameters, the first coordinate is transformed into world coordinates to obtain the second coordinate of the target pixel in the world coordinate system.
[0187] The target image feature data corresponding to the environmental perspective image is determined based on the second coordinate and target feature data.
[0188] In one embodiment, when the processor 110 performs camera coordinate transformation processing on the pixel coordinates based on the fisheye camera intrinsic parameters to obtain the first coordinates of the target pixel in the camera coordinate system, it specifically performs the following operations:
[0189] Obtain the focal length, principal point coordinates, and distortion parameters of the panoramic fisheye camera;
[0190] Camera coordinate transformation is performed based on focal length, principal point coordinates, distortion parameters, and pixel coordinates to obtain the first coordinates of the target pixel in the camera coordinate system.
[0191] In one embodiment, when the processor 110 performs the step of performing bird's-eye view feature transformation processing on the target image feature data to obtain bird's-eye view feature data, it specifically performs the following operations:
[0192] Highly compressed feature data is obtained by performing high-dimensional feature compression on the target image feature data through an attention mechanism.
[0193] Determine the orientation matrix and depth-distance matrix;
[0194] Bird's-eye view feature data is obtained by performing bird's-eye view feature transformation on highly compressed feature data, orientation matrix, and depth distance matrix.
[0195] In one embodiment, when the processor 110 performs multi-class feature extraction processing on multiple environmental view images to obtain image category feature data corresponding to the environmental view images, it specifically performs the following operations:
[0196] Image feature extraction is performed on the environmental perspective image to obtain the image feature data corresponding to the environmental perspective image;
[0197] Depth prediction data corresponding to the environmental view image is obtained by performing depth prediction processing on the environmental view image.
[0198] Semantic segmentation is performed on the environmental perspective image to obtain the semantic recognition data corresponding to the environmental perspective image;
[0199] Image type feature data corresponding to environmental perspective images are determined based on image feature data, depth prediction data, and semantic recognition data.
[0200] In one embodiment, when the processor 110 performs the step of performing depth prediction processing on the environmental view image to obtain depth prediction data corresponding to the environmental view image, it specifically performs the following operations:
[0201] Obtain the intrinsic parameters of the fisheye camera from the panoramic fisheye camera;
[0202] Depth prediction data corresponding to environmental perspective images is obtained by performing depth prediction processing on environmental perspective images based on fisheye camera intrinsic parameters.
[0203] In one embodiment, when the processor 110 performs the step of performing feature transformation processing on the bird's-eye view feature data to obtain bird's-eye view hierarchical feature data, it specifically performs the following operations:
[0204] The feature extraction backbone network is used to perform hierarchical feature parsing on the bird's-eye view feature data to obtain the bird's-eye view hierarchical feature data.
[0205] This application also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor to implement the environment awareness methods as described in the above embodiments.
[0206] This application also provides a computer program product that stores at least one instruction, which is loaded and executed by the processor to implement the environmental perception methods of the above embodiments.
[0207] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0208] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An environmental perception method, characterized in that, The method is applied to a movable device, and comprises: A plurality of environment perspective images are collected by a surround-view fisheye camera, and a plurality of category feature extraction processes are performed on the plurality of environment perspective images to obtain image category feature data corresponding to the environment perspective images; Dimension conversion processing is performed on the image category feature data to obtain bird's-eye view perspective feature data; Feature transformation processing is performed on the bird's-eye view perspective feature data to obtain bird's-eye view perspective hierarchical feature data; Feature recognition processing is performed on the bird's-eye view perspective hierarchical feature data by using a multi-task perception model to obtain a movable area segmentation result, a ground semantic segmentation result, and a target object detection result; wherein the multi-task perception model is obtained by training based on sample bird's-eye view perspective hierarchical feature data corresponding to sample movable area segmentation result labels, sample ground semantic segmentation result labels, and target object detection result labels.
2. The method of claim 1, wherein, The dimension conversion processing on the image category feature data to obtain the bird's-eye view perspective feature data comprises: An fisheye camera parameter of the surround-view fisheye camera is obtained; Coordinate system feature conversion processing is performed on the image category feature data based on the fisheye camera parameter to obtain target image feature data corresponding to the environment perspective images; Bird's-eye view perspective feature conversion processing is performed on the target image feature data to obtain the bird's-eye view perspective feature data.
3. The method of claim 2, wherein, The fisheye camera parameter comprises an fisheye camera intrinsic parameter and an fisheye camera extrinsic parameter, and the coordinate system feature conversion processing on the image category feature data based on the fisheye camera parameter to obtain the target image feature data corresponding to the environment perspective images comprises: A target pixel in the image category feature data is obtained, and a pixel coordinate of the target pixel in the environment perspective image is determined; Camera coordinate conversion processing is performed on the pixel coordinate based on the fisheye camera intrinsic parameter to obtain a first coordinate of the target pixel in a camera coordinate system; World coordinate conversion processing is performed on the first coordinate based on the fisheye camera extrinsic parameter to obtain a second coordinate of the target pixel in a world coordinate system; The target image feature data corresponding to the environment perspective image is determined based on the second coordinate and the target feature data.
4. The method of claim 3, wherein, The camera coordinate conversion processing on the pixel coordinate based on the fisheye camera intrinsic parameter to obtain the first coordinate of the target pixel in the camera coordinate system comprises: A focal length, a principal point coordinate, and a distortion parameter of the surround-view fisheye camera are obtained; Camera coordinate conversion processing is performed based on the focal length, the principal point coordinate, the distortion parameter, and the pixel coordinate to obtain the first coordinate of the target pixel in the camera coordinate system.
5. The method of claim 2, wherein, The bird's-eye view perspective feature conversion processing on the target image feature data to obtain the bird's-eye view perspective feature data comprises: High-dimensional feature compression processing is performed on the target image feature data by using an attention mechanism to obtain high-compression feature data; A direction matrix and a depth distance matrix are determined; Bird's-eye view perspective feature conversion processing is performed on the high-compression feature data, the direction matrix, and the depth distance matrix to obtain the bird's-eye view perspective feature data.
6. The method of claim 1, wherein, The multi-category feature extraction processing on the multiple environment perspective images obtains image category feature data corresponding to the environment perspective images, including: The image feature extraction processing on the environment perspective images obtains image feature data corresponding to the environment perspective images; The depth prediction processing on the environment perspective images obtains depth prediction data corresponding to the environment perspective images; The semantic segmentation processing on the environment perspective images obtains semantic recognition data corresponding to the environment perspective images; The image category feature data corresponding to the environment perspective images is determined based on the image feature data, the depth prediction data and the semantic recognition data.
7. The method of claim 6, wherein, The depth prediction processing on the environment perspective images obtains depth prediction data corresponding to the environment perspective images, including: The fisheye camera intrinsic parameters of the surround-view fisheye camera are acquired; The depth prediction processing on the environment perspective images based on the fisheye camera intrinsic parameters obtains depth prediction data corresponding to the environment perspective images.
8. The method of claim 1, wherein, The feature transformation processing on the bird's-eye view perspective feature data obtains bird's-eye view perspective hierarchical feature data, including: The hierarchical feature analysis processing on the bird's-eye view perspective feature data by using a feature extraction backbone network obtains bird's-eye view perspective hierarchical feature data.
9. An environmental perception apparatus, characterized in that, The device is applied to a drivable equipment, and the device includes: A feature extraction module is configured to collect multiple environment perspective images by using a surround-view fisheye camera, and to obtain image category feature data corresponding to the environment perspective images by performing multi-category feature extraction processing on the multiple environment perspective images. A feature conversion module is configured to obtain bird's-eye view perspective feature data by performing dimension conversion processing on the image category feature data. A feature transformation module is configured to obtain bird's-eye view perspective hierarchical feature data by performing feature transformation processing on the bird's-eye view perspective feature data. A feature recognition module is configured to obtain drivable area segmentation results, ground semantic segmentation results and target object detection results by performing feature recognition processing on the bird's-eye view perspective hierarchical feature data by using a multi-task perception model.
10. A computer storage medium, characterized in that The computer storage medium stores instructions, and the instructions are adapted to be loaded and executed by the processor to implement the method in any one of claims 1-8.
11. An electronic device, comprising: The device includes: A processor and a memory, wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the method in any one of claims 1-8.