Image Processing Method, Apparatus and Electronic Device Based on Perspective Projection

By combining camera parameters and perspective projection method without camera parameters, the main view image is processed, and the problem of inaccurate projection in autonomous driving is solved, more accurate bird's-eye view semantic segmentation is achieved, and segmentation accuracy in autonomous driving scenes is improved.

CN114708332BActive Publication Date: 2025-07-08BEIJING PHIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210225391.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-07-08
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

The existing perspective projection algorithm based on the camera model is large in the projection results of the camera model-based autonomous driving, which leads to a large deviation in the projection results of the aerial view, resulting in a decrease in the semantic segmentation accuracy of the bird's eye view, and the model generalization ability without camera parameters is poor.

Method used

Combined with the perspective projection method using camera parameters and without camera parameters, features are acquired through the main view image encoding network, and feature processing and fusion are performed using the first perspective projection module and the second perspective projection module to obtain the bird's-eye view spatial features, and decode them through the bird's-eye view image decoding network to output the object segmented image.

Benefits of technology

The accuracy of semantic segmentation of bird's aerial view is improved, especially in scenes where the ground is undulating and the height of objects is inconsistent, and the segmentation accuracy of dynamic objects and small objects is improved, providing more accurate and robust technical support for autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708332B_ABST
    Figure CN114708332B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an image processing method, apparatus, and electronic device based on perspective projection. The method includes: obtaining camera parameters of a target camera and a main view image captured by the target camera; encoding the main view image through a main view image encoding network to obtain N first main view spatial features; processing the N first main view spatial features through a first perspective projection module to obtain a first bird's-eye view spatial feature, and processing a target main view spatial feature among the N first main view spatial features through a second perspective projection module to obtain a second bird's-eye view spatial feature; performing feature fusion processing on the first bird's-eye view spatial feature and the second bird's-eye view spatial feature to obtain a third bird's-eye view spatial feature; and decoding the third bird's-eye view spatial feature through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space. The above solution improves the semantic segmentation accuracy of the bird's-eye view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image processing method, apparatus, and electronic device based on perspective projection. Background Art

[0002] Currently, the existing perspective projection (FV2BEV, Frontal View to Bird’s-Eye View) algorithm for autonomous driving from a front view to a bird's-eye view based on a camera model has a very fatal flaw: the camera model strongly depends on the flat ground plane assumption, which is often not satisfied in autonomous driving scenarios. For example, in the following two driving scenarios: (a) the road type of uphill and downhill; (b) the road edge with large height fluctuations. The projection results of the algorithm using camera parameters at these positions will have a large deviation from the true annotation, resulting in a decrease in the semantic segmentation accuracy in the bird's-eye view (BEV) space and causing great difficulties for subsequent tasks such as road layout estimation and motion path planning. And the model without using camera parameters is prone to problems such as an overly large solution space to be searched and poor generalization ability. Therefore, the above two major types of solutions are prone to problems such as projection position deviation and inaccurate segmentation when dealing with the FV2BEV task. Summary of the Invention

[0003] The present invention provides an image processing method, apparatus, and electronic device based on perspective projection to solve, to a certain extent, the problem that the existing technology is prone to inaccurate projection and segmentation when dealing with the FV2BEV task.

[0004] In the first aspect of the implementation of the present invention, an image processing method based on perspective projection is provided, including:

[0005] Obtain the camera parameters of the target camera and the front view image captured by the target camera;

[0006] Encode the front view image through a front view image encoding network to obtain N first front view space features, where the N first front view space features are front view space features of N resolutions;

[0007] Process the N first front view space features through a first perspective projection module to obtain a first bird's-eye view space feature, and process the target front view space feature among the N first front view space features through a second perspective projection module to obtain a second bird's-eye view space feature;

[0008] Perform feature fusion processing on the first bird's-eye view space feature and the second bird's-eye view space feature to obtain a third bird's-eye view space feature;

[0009] Decode the third bird's-eye view spatial feature through a bird's-eye view image decoding network to obtain an object segmentation image of the bird's-eye view space;

[0010] Among them, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0011] Optionally, the process of obtaining the first bird's-eye view spatial feature by processing the N first front view spatial features through the first perspective projection module includes:

[0012] Process the N first front view spatial features through the first perspective projection module to obtain N first bird's-eye view sub-features;

[0013] Perform feature stitching on the N first bird's-eye view sub-features to obtain the first bird's-eye view spatial feature.

[0014] Optionally, the process of obtaining N first bird's-eye view sub-features by processing the N first front view spatial features through the first perspective projection module includes:

[0015] Process the N first front view spatial features through the first perspective projection module to obtain a first depth information feature;

[0016] Project the first depth information feature onto N depth ranges to obtain N first bird's-eye view sub-features in the N depth ranges.

[0017] Optionally, the process of obtaining a first depth information feature by processing the N first front view spatial features through the first perspective projection module includes:

[0018] Process the N first front view spatial features through a first convolutional network to obtain N two-dimensional spatial features;

[0019] Perform compression processing on the N two-dimensional spatial features in the height dimension to obtain N second front view spatial features;

[0020] Process the N second front view spatial features through a second convolutional network for feature transformation to obtain N one-dimensional spatial features;

[0021] Perform resampling on the N one-dimensional spatial features to obtain a first depth information feature.

[0022] Optionally, the process of obtaining a second bird's-eye view spatial feature by processing the target front view spatial feature among the N first front view spatial features through the second perspective projection module includes:

[0023] The target main view space feature among the N first main view space features is processed by the second perspective projection module to obtain N second bird's-eye view sub-features;

[0024] The N second bird's-eye view sub-features are subjected to feature stitching processing to obtain the second bird's-eye view space feature.

[0025] Optionally, the processing of the target main view space feature among the N first main view space features by the second perspective projection module to obtain N second bird's-eye view sub-features includes:

[0026] The target main view space feature among the N first main view space features is subjected to feature transformation processing by a first multi-layer perceptron network to obtain a fourth bird's-eye view space feature;

[0027] The fourth bird's-eye view space feature is subjected to feature transformation processing by a second multi-layer perceptron network to obtain a third main view space feature;

[0028] The target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature are input into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features.

[0029] Optionally, the inputting of the target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features includes:

[0030] The target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature are subjected to feature processing by the second perspective projection module to obtain a second depth information feature;

[0031] The second depth information feature is projected onto N depth ranges to obtain N second bird's-eye view sub-features in the N depth ranges.

[0032] In the second aspect of the implementation of the present invention, an image processing apparatus based on perspective projection is provided, including:

[0033] A first acquisition module, configured to acquire camera parameters of a target camera and a main view image captured by the target camera;

[0034] A first processing module, configured to encode the main view image through a main view image encoding network to obtain N first main view space features, where the N first main view space features are main view space features of N resolutions;

[0035] A second processing module, configured to process the N first front view space features through a first perspective projection module to obtain first bird's-eye view space features, and process a target front view space feature among the N first front view space features through a second perspective projection module to obtain second bird's-eye view space features;

[0036] A third processing module, configured to perform feature fusion processing on the first bird's-eye view space features and the second bird's-eye view space features to obtain third bird's-eye view space features;

[0037] A fourth processing module, configured to decode the third bird's-eye view space features through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space;

[0038] Wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0039] In a third aspect of the embodiments of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0040] The memory is used to store a computer program;

[0041] The processor is configured to implement the steps in the above-mentioned image processing method based on perspective projection when executing the program stored in the memory.

[0042] In a fourth aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned image processing method based on perspective projection.

[0043] In a fifth aspect of the embodiments of the present invention, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute the above-mentioned image processing method based on perspective projection.

[0044] For the prior art, one of the above technical solutions has the following advantages:

[0045] In the above embodiments of the present invention, by obtaining the camera parameters of the target camera and the main view image captured by the target camera, and encoding the main view image through the main view image encoding network to obtain the first main view spatial features of N resolutions, processing the N first main view spatial features through the first perspective projection module to obtain the first bird's-eye view spatial features, and processing the target main view spatial features among the N first main view spatial features through the second perspective projection module to obtain the second bird's-eye view spatial features; and performing feature fusion processing on the first bird's-eye view spatial features and the second bird's-eye view spatial features to obtain the third bird's-eye view spatial features; decoding the third bird's-eye view spatial features through the bird's-eye view image decoding network to obtain the object segmentation image in the bird's-eye view space; wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0046] In the above solution, first, the main view image and its corresponding camera parameters are obtained through the target camera, then the main view image is encoded, and the encoded first main view spatial features are respectively provided to the first perspective projection module using camera parameters and the second perspective projection module not using camera parameters to perform effective information extraction and fusion of two branches to obtain the third bird's-eye view spatial features, and the fused third bird's-eye view spatial features are sent to the bird's-eye view image decoding network, and finally the object segmentation image in the bird's-eye view space is output. By combining the perspective projection methods using camera parameters and not using camera parameters, more accurate and robust bird's-eye view semantic segmentation is achieved, and the bird's-eye view semantic segmentation accuracy in scenarios such as ground undulation and inconsistent object and ground heights is improved.

[0047] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Brief Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments.

[0049] Figure 1 It is one of the flow diagrams of the image processing method based on perspective projection provided by the embodiments of the present invention;

[0050] Figure 2 It is the second flow diagram of the image processing method based on perspective projection provided by the embodiments of the present invention;

[0051] Figure 3It is the third flowchart diagram of the image processing method based on perspective projection provided by the embodiment of the present invention;

[0052] Figure 4 It is the block diagram of the image processing device based on perspective projection provided by the embodiment of the present invention;

[0053] Figure 5 It is the block diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0055] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0056] The problem to be solved by FV2BEV is to obtain the segmentation mask map of different objects in the corresponding area from the input front view image. Currently, the existing FV2BEV algorithms can be divided into two major categories: perspective projection schemes with camera parameters and without camera parameters.

[0057] Among them, a typical representative of the FV2BEV algorithm using camera parameters is: the Pyramid Occupancy Network (PON) algorithm. The technical solution of the PON algorithm is as follows: Input the frontal view (FV) image in the autonomous driving scenario, first use the backbone network to extract the features of this image, where the backbone network uses the Residual Network (ResNet), and then output 5 feature maps with different scale sizes; Send these 5 feature maps with different scale sizes into the multi-scale dense mapping network, project them into the BEV space respectively, and generate 5 feature maps with different depth ranges (the dimension of the feature map is B×C×H×W, where B is the batch size, C is the number of channels, H is the height of the feature map, and W is the width of the feature map), and splice them together along the channel C direction to form the BEV space feature; Immediately send the BEV space feature into the decoding network to output the pixel-level classification result in the BEV space.

[0058] However, in the actual autonomous driving environment, it is easy to have situations where the road surface is uneven and the heights of objects distributed on the road are different. Therefore, it is unreliable to rely solely on the model depending on camera parameters for semantic segmentation in the BEV space. The existing technical solutions strongly rely on the assumption that the ground plane is absolutely flat, and this assumption is very harsh and even unreasonable for the autonomous driving scenario. Because in the actual driving scenario, uphill and downhill road conditions are often encountered, and at this time the ground is uneven relative to the camera; in addition, objects such as the road edge have a certain height relative to the ground, and these objects are also uneven relative to the camera. Once the assumption that the ground plane is absolutely flat does not hold, technologies such as PON will have large segmentation errors, and the mean Intersection of Union (mIoU) index of semantic segmentation will decrease accordingly. Especially for categories that account for a small proportion in the BEV space itself, the inaccurate projection from the FV space to the BEV space will greatly reduce the segmentation effect of such objects. For example, for pedestrians and reflective traffic cones, the areas occupied by these two types of objects in the BEV space are very small, and once misclassified, it is easy to cause a significant drop in the overall index.

[0059] A typical representative of the FV2BEV algorithm that does not adopt camera parameters is the Projecting Your View Attentively (PYVA); this type of algorithm completely abandons the prior knowledge of the pinhole imaging camera model, which will cause redundant calculations to a certain extent; at the same time, searching for reasonable solutions in an overly large solution space, and quite a part of the solutions are obviously inconsistent with the prior knowledge. If the training scenario is not rich enough, it is often easy to lead to overfitting problems. When the trained model is transferred to a new scenario for the semantic segmentation task of the autonomous driving bird's-eye view, the generalization performance is often poor and the semantic segmentation is inaccurate.

[0060] Therefore, the embodiments of the present invention provide an image processing method, device and electronic device based on perspective projection, which combines the perspective projection methods with and without camera parameters to achieve more accurate and robust BEV semantic segmentation, and improve the BEV semantic segmentation accuracy in scenarios such as ground undulation and inconsistent object and ground heights.

[0061] Specifically, as Figure 1 and Figure 2 shown, an image processing method based on perspective projection provided by the embodiments of the present invention may specifically include the following steps:

[0062] Step 101, obtain the camera parameters of the target camera and the main view image captured by the target camera.

[0063] Specifically, by capturing a front view image with the target camera, that is, obtaining the main view image, and obtaining the camera parameters of the target camera. The camera parameters include: camera intrinsics and camera extrinsics. Among them, the camera intrinsics include the focal lengths and offsets in the horizontal and vertical axes, and the camera extrinsics include the rotation angle and translation amount between the camera coordinate system and the world coordinate system.

[0064] Among them, the target camera can be a vehicle-mounted surround-view camera or other types of cameras, which are not specifically limited here.

[0065] Step 102, encode the main view image through the main view image encoding network to obtain N first main view spatial features, where the N first main view spatial features are main view spatial features of N resolutions, and N can be a positive integer.

[0066] Specifically, after the main view image is encoded through the main view image encoding network, N first main view spatial features of different resolutions are obtained. If N is a positive integer greater than 1, the N first main view spatial features are N different resolution main view spatial features.

[0067] Among them, the main view image encoding network can be a Residual Network (ResNet) or a Swin Transformer. Through this main view image encoding network, N feature maps with different resolutions in the main view space (i.e., the first main view space features) are encoded. The N first main view space features with different resolutions respectively contain semantic features at different levels. For example: as Figure 2 shown, if N is equal to 5, then after the main view image is encoded by the main view image encoding network, 5 first main view space features with different resolutions are obtained, namely the first main view space feature 1, the first main view space feature 2, the first main view space feature 3, the first main view space feature 4, and the first main view space feature 5.

[0068] Step 103: Process the N first main view space features through the first perspective projection module to obtain the first bird's-eye view space features, and process the target main view space features among the N first main view space features through the second perspective projection module to obtain the second bird's-eye view space features;

[0069] Among them, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0070] Specifically, input the N first main view space features into the first perspective projection module for processing to output the first bird's-eye view space features. And input the N first main view space features into the second perspective projection module for processing to output the second bird's-eye view space features. That is, the N first main view space features are used as the input features of the subsequent two branches. The two branches refer to the first perspective projection module branch using camera parameters and the second perspective projection module branch not using camera parameters.

[0071] Step 104: Perform feature fusion processing on the first bird's-eye view space features and the second bird's-eye view space features to obtain the third bird's-eye view space features.

[0072] Specifically, the first perspective projection module branch that uses camera parameters and the second perspective projection module branch that does not use camera parameters are combined, which is equivalent to a two-stream model. That is, the first bird's-eye view spatial feature obtained from the first perspective projection module branch and the second bird's-eye view spatial feature obtained from the second perspective projection module branch are subjected to feature fusion processing through pixel-by-pixel superposition to obtain the third bird's-eye view spatial feature after fusion. The perspective projection algorithm that uses camera parameters is more accurate for perspective transformation of static objects with flat surfaces such as the ground and crosswalks, but less accurate for objects with uneven surfaces; while the perspective projection algorithm that does not use camera parameters is more accurate for perspective transformation of dynamic objects such as buses and motorcycles, but is prone to overfitting. By combining the advantages of these two branches and fusing the features of the two branches, the model can not only master powerful camera prior knowledge, narrow the search range of feasible solutions, and prevent overfitting, but also maintain strong learning ability for complex scenarios such as uneven roads and uneven object surfaces.

[0073] Among them, the second perspective projection module branch that does not use camera parameters can use a multi-layer perceptron (MLP) or an attention network, etc.

[0074] Step 105: Decode the third bird's-eye view spatial feature through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space.

[0075] Specifically, the third bird's-eye view spatial feature after fusion undergoes post-processing by the bird's-eye view image decoding network to obtain the final object segmentation image in the bird's-eye view space. A bird's-eye view image decoding network is constructed using stacked ResNet modules and a classifier to process the third bird's-eye view spatial feature after perspective view fusion transformation, thereby obtaining a semantic probability map in the bird's-eye view space, that is, an object segmentation image, and obtaining pixel-by-pixel class information.

[0076] In the above embodiments of the present invention, the main view image and its corresponding camera parameters are obtained through a target camera, and then the main view image is encoded. The encoded first main view spatial feature is respectively provided to the first perspective projection module that uses camera parameters and the second perspective projection module that does not use camera parameters to perform effective information extraction and fusion of the two branches to obtain the third bird's-eye view spatial feature, and the third bird's-eye view spatial feature after fusion is sent to the bird's-eye view image decoding network, and finally an object segmentation image in the bird's-eye view space is output. By combining the perspective projection methods with and without using camera parameters, more accurate and robust bird's-eye view semantic segmentation is achieved, and the accuracy of bird's-eye view semantic segmentation in scenarios such as ground undulation and inconsistent object-ground height is improved.

[0077] As an optional embodiment, step 103 processes the N first front view space features through a first perspective projection module to obtain first bird's-eye view space features, which may specifically include:

[0078] Process the N first front view space features through the first perspective projection module to obtain N first bird's-eye view sub-features;

[0079] Perform feature stitching processing on the N first bird's-eye view sub-features to obtain the first bird's-eye view space features.

[0080] Specifically, input the N first front view space features into the first perspective projection module for processing, and correspondingly output N first bird's-eye view sub-features. Stitch these N first bird's-eye view sub-features to obtain the first bird's-eye view space features.

[0081] For example: as Figure 2 shown, if the value of N is 5, then the 5 first front view space features are input into the first perspective projection module, and the corresponding 5 first bird's-eye view sub-features are output. The 5 sub-features correspond to the features of the bird's-eye view space in different depth ranges. These 5 sub-features are stitched together and fused to form the final first bird's-eye view space features.

[0082] As an optional embodiment, the step of processing the N first front view space features through the first perspective projection module to obtain N first bird's-eye view sub-features may specifically include:

[0083] Process the N first front view space features through the first perspective projection module to obtain a first depth information feature;

[0084] Project the first depth information feature onto N depth ranges to obtain N first bird's-eye view sub-features in N depth ranges.

[0085] Specifically, process the N first front view space features through the first perspective projection module to obtain a feature with depth information, that is, the first depth information feature (Depth Context Feature, DCF). If N is a positive integer greater than 1, then project the first depth information feature onto N different depth ranges to obtain N first bird's-eye view sub-features in N different depth ranges.

[0086] As an optional embodiment, the step of processing the N first front view space features through the first perspective projection module to obtain a first depth information feature may specifically include:

[0087] Process the N first front view space features through a first convolutional network to obtain N two-dimensional space features;

[0088] Perform compression processing on the N two-dimensional spatial features in the height dimension to obtain N second front view spatial features;

[0089] Perform feature transformation processing on the N second front view spatial features through a second convolutional network to obtain N one-dimensional spatial features;

[0090] Perform resampling processing on the N one-dimensional spatial features to obtain a first depth information feature.

[0091] Specifically, as Figure 3 shown, the front view image 31 passes through the front view image encoding network to obtain preliminary features, that is, N front view spatial features 32 with different resolutions. Then, perform feature processing on the N first front view spatial features through a first convolutional network, and perform compression processing on the N two-dimensional spatial features obtained after feature processing in the height dimension, that is, obtain N second front view spatial features 33. Further, perform feature transformation processing on the above N second front view spatial features through a second convolutional network, and perform resampling processing on the N one-dimensional spatial features after feature transformation processing to obtain features with depth information, that is, a first depth information feature 34. Project the first depth information feature into different depth ranges to obtain first bird's-eye view sub-features in different depth ranges, and perform feature stitching processing on the first bird's-eye view sub-features in different depth ranges to obtain a first bird's-eye view spatial feature 35. Perform feature fusion processing on the first bird's-eye view spatial feature and the second bird's-eye view spatial feature obtained through the second perspective projection module branch that does not adopt camera parameters to obtain a third bird's-eye view spatial feature, and perform decoding processing on the third bird's-eye view spatial feature through the bird's-eye view image decoding network to obtain an object segmentation image 36 in the bird's-eye view space.

[0092] Among them, the first convolutional network is a two-dimensional convolutional network, and the second convolutional network is a one-dimensional convolutional network.

[0093] As an optional embodiment, step 103 processes the target front view spatial feature among the N first front view spatial features through a second perspective projection module to obtain a second bird's-eye view spatial feature, which specifically includes:

[0094] Process the target front view spatial feature among the N first front view spatial features through the second perspective projection module to obtain N second bird's-eye view sub-features;

[0095] Perform feature stitching processing on the N second bird's-eye view sub-features to obtain a second bird's-eye view spatial feature.

[0096] Specifically, the target front view space feature is one of the N first front view space features. The target front view space feature is input into the second perspective projection module for processing, and N second bird's-eye view sub-features are output. The N second bird's-eye view sub-features are feature-stitched to obtain the second bird's-eye view space feature.

[0097] As an optional embodiment, the step of processing the target front view space feature among the N first front view space features through the second perspective projection module to obtain N second bird's-eye view sub-features may specifically include:

[0098] Performing feature transformation processing on the target front view space feature among the N first front view space features through a first multi-layer perceptron network to obtain a fourth bird's-eye view space feature;

[0099] Performing feature transformation processing on the fourth bird's-eye view space feature through a second multi-layer perceptron network to obtain a third front view space feature;

[0100] Inputting the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features.

[0101] Specifically, performing feature transformation on the target front view space feature through a first multi-layer perceptron (MLP) network can obtain the features of the bird's-eye view space, that is, the fourth bird's-eye view space feature. Then, performing feature transformation on the above fourth bird's-eye view space feature through a second multi-layer perceptron (MLP) network can obtain the new features of the new front view space, that is, the third front view space feature. Inputting the above target front view space feature, fourth bird's-eye view space feature, and third front view space feature into the second perspective projection module can output N second bird's-eye view sub-features.

[0102] For example: As Figure 2 shown, if N is set to 5, then the target front view space feature is input into the second perspective projection module, and 5 second bird's-eye view sub-features are output. The 5 sub-features correspond to the features of the bird's-eye view space in different depth ranges. These 5 sub-features are stitched together and fused to form the final second bird's-eye view space feature.

[0103] As an optional embodiment, the step of inputting the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features may specifically include:

[0104] The second perspective projection module performs feature processing on the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature to obtain a second depth information feature;

[0105] The second depth information feature is projected onto N depth ranges to obtain second bird's-eye view sub-features of the N depth ranges.

[0106] Specifically, the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature are input into the second perspective projection module for feature processing to obtain a feature with depth information, that is, the second depth information feature. If N is a positive integer greater than 1, then the second depth information feature is projected onto N different depth ranges, and second bird's-eye view sub-features of N different depth ranges can be obtained.

[0107] It should be noted that in the above embodiments, hardware and software are combined. The hardware can provide a front view image through an in-vehicle surround camera and does not rely on specific hardware. The algorithm model is trained on a 4-card NVIDIA GeForce GTX 3090 graphics card. The batch size of a single card is batch size = 12, and the total batch size is batch size = 48. The optimizer used can be the AdamW optimizer.

[0108] In summary, the embodiments of the present invention effectively fuse the advantages of models with and without camera parameters, enabling the model to not only master powerful camera prior knowledge, narrow the search range of feasible solutions, and prevent overfitting, but also maintain a powerful learning ability for complex scenarios such as road unevenness and height differences on the surface of objects, improving the accuracy of perspective projection and semantic segmentation of FV2BEV. It can achieve the highest segmentation accuracy among existing methods on three extremely challenging autonomous driving datasets, namely nuScenes, Argoverse, and KITTI. In particular, it improves the segmentation accuracy of dynamic objects and small objects in autonomous driving scenarios, providing strong technical support for tasks such as 3D detection, trajectory prediction, and motion planning in autonomous driving scenarios.

[0109] As Figure 4 shown, an image processing device 400 based on perspective projection provided by an embodiment of the present invention includes:

[0110] A first acquisition module 401, configured to acquire camera parameters of a target camera and a front view image captured by the target camera;

[0111] The first processing module 402 is configured to encode the main view image through a main view image encoding network to obtain N first main view spatial features, where the N first main view spatial features are main view spatial features of N resolutions;

[0112] The second processing module 403 is configured to process the N first main view spatial features through a first perspective projection module to obtain a first bird's-eye view spatial feature, and process the target main view spatial feature among the N first main view spatial features through a second perspective projection module to obtain a second bird's-eye view spatial feature;

[0113] The third processing module 404 is configured to perform feature fusion processing on the first bird's-eye view spatial feature and the second bird's-eye view spatial feature to obtain a third bird's-eye view spatial feature;

[0114] The fourth processing module 405 is configured to decode the third bird's-eye view spatial feature through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space;

[0115] Wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0116] Optionally, when the second processing module 403 processes the N first main view spatial features through the first perspective projection module to obtain a first bird's-eye view spatial feature, it is specifically configured to:

[0117] Process the N first main view spatial features through the first perspective projection module to obtain N first bird's-eye view sub-features;

[0118] Perform feature stitching processing on the N first bird's-eye view sub-features to obtain the first bird's-eye view spatial feature.

[0119] Optionally, when the second processing module 403 processes the N first main view spatial features through the first perspective projection module to obtain N first bird's-eye view sub-features, it is specifically configured to:

[0120] Process the N first main view spatial features through the first perspective projection module to obtain a first depth information feature;

[0121] Project the first depth information feature onto N depth ranges to obtain N first bird's-eye view sub-features in N depth ranges.

[0122] Optionally, when the second processing module 403 processes the N first main view spatial features through the first perspective projection module to obtain a first depth information feature, it is specifically configured to:

[0123] Feature process the N first front view space features through a first convolutional network to obtain N two-dimensional space features;

[0124] Perform compression processing on the N two-dimensional space features in the height dimension to obtain N second front view space features;

[0125] Feature transform the N second front view space features through a second convolutional network to obtain N one-dimensional space features;

[0126] Perform resampling processing on the N one-dimensional space features to obtain a first depth information feature.

[0127] Optionally, when the second processing module 403 processes the target front view space feature among the N first front view space features through the second perspective projection module to obtain a second bird's-eye view space feature, it specifically is used for:

[0128] Process the target front view space feature among the N first front view space features through the second perspective projection module to obtain N second bird's-eye view sub-features;

[0129] Perform feature stitching processing on the N second bird's-eye view sub-features to obtain a second bird's-eye view space feature.

[0130] Optionally, when the second processing module 403 processes the target front view space feature among the N first front view space features through the second perspective projection module to obtain N second bird's-eye view sub-features, it specifically is used for:

[0131] Feature transform the target front view space feature among the N first front view space features through a first multi-layer perceptron network to obtain a fourth bird's-eye view space feature;

[0132] Feature transform the fourth bird's-eye view space feature through a second multi-layer perceptron network to obtain a third front view space feature;

[0133] Input the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features.

[0134] Optionally, when the second processing module 403 inputs the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature into the second perspective projection module for processing to obtain N second bird's-eye view sub-features, it specifically is used for:

[0135] The second perspective projection module performs feature processing on the target front view space feature, the fourth bird's-eye view space feature, and the third front view space feature to obtain a second depth information feature;

[0136] The second depth information feature is projected onto N depth ranges to obtain second bird's-eye view sub-features for the N depth ranges.

[0137] It should be noted that the embodiment of the image processing device based on perspective projection is a device corresponding to the above-mentioned image processing method based on perspective projection. All implementation manners of the above method embodiment are applicable to this device embodiment and can achieve the same technical effects, which will not be elaborated here.

[0138] In summary, the embodiment of the present invention effectively integrates the advantages of models with and without camera parameters, enabling the model to not only master powerful camera prior knowledge, narrow the search range of feasible solutions, and prevent overfitting, but also maintain a strong learning ability for complex scenarios such as road unevenness and uneven object surfaces, improving the accuracy of perspective projection and semantic segmentation of FV2BEV. It can achieve the highest segmentation accuracy among existing methods on autonomous driving datasets such as nuScenes, Argoverse, and KITTI. In particular, it improves the segmentation accuracy of dynamic objects and small objects in autonomous driving scenarios, providing strong technical support for tasks such as 3D detection, trajectory prediction, and motion planning in autonomous driving scenarios.

[0139] The embodiment of the present invention also provides an electronic device. As Figure 5 shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 complete communication with each other through the communication bus 504.

[0140] The memory 503 is used to store a computer program.

[0141] When the processor 501 is used to execute the program stored on the memory 503, the following steps are implemented:

[0142] Obtain the camera parameters of the target camera and the front view image captured by the target camera;

[0143] Encode the front view image through a front view image encoding network to obtain N first front view space features, where the N first front view space features are front view space features with N resolutions;

[0144] Process the N first front view space features through a first perspective projection module to obtain first bird's-eye view space features, and process the target front view space features among the N first front view space features through a second perspective projection module to obtain second bird's-eye view space features;

[0145] Perform feature fusion processing on the first bird's-eye view space features and the second bird's-eye view space features to obtain third bird's-eye view space features;

[0146] Decode the third bird's-eye view space features through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space;

[0147] Wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

[0148] Optionally, when the processor 501 processes the N first front view space features through the first perspective projection module to obtain first bird's-eye view space features, it is specifically used for:

[0149] Process the N first front view space features through the first perspective projection module to obtain N first bird's-eye view sub-features;

[0150] Perform feature stitching processing on the N first bird's-eye view sub-features to obtain the first bird's-eye view space features.

[0151] Optionally, when the processor 501 processes the N first front view space features through the first perspective projection module to obtain N first bird's-eye view sub-features, it is specifically used for:

[0152] Process the N first front view space features through the first perspective projection module to obtain first depth information features;

[0153] Project the first depth information features into N depth ranges to obtain first bird's-eye view sub-features in N depth ranges.

[0154] Optionally, when the processor 501 processes the N first front view space features through the first perspective projection module to obtain first depth information features, it is specifically used for:

[0155] Process the N first front view space features through a first convolutional network to obtain N two-dimensional space features;

[0156] Perform compression processing on the N two-dimensional space features in the height dimension to obtain N second front view space features;

[0157] Performing feature transformation processing on the N second principal view space features through a second convolutional network to obtain N one-dimensional space features;

[0158] Performing resampling processing on the N one-dimensional space features to obtain a first depth information feature.

[0159] Optionally, when the processor 501 processes the target principal view space feature among the N first principal view space features through the second perspective projection module to obtain a second bird's-eye view space feature, it is specifically used for:

[0160] Processing the target principal view space feature among the N first principal view space features through the second perspective projection module to obtain N second bird's-eye view sub-features;

[0161] Performing feature stitching processing on the N second bird's-eye view sub-features to obtain the second bird's-eye view space feature.

[0162] Optionally, when the processor 501 processes the target principal view space feature among the N first principal view space features through the second perspective projection module to obtain N second bird's-eye view sub-features, it is specifically used for:

[0163] Performing feature transformation processing on the target principal view space feature among the N first principal view space features through a first multi-layer perceptron network to obtain a fourth bird's-eye view space feature;

[0164] Performing feature transformation processing on the fourth bird's-eye view space feature through a second multi-layer perceptron network to obtain a third principal view space feature;

[0165] Inputting the target principal view space feature, the fourth bird's-eye view space feature, and the third principal view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features.

[0166] Optionally, when the processor 501 inputs the target principal view space feature, the fourth bird's-eye view space feature, and the third principal view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features, it is specifically used for:

[0167] Performing feature processing on the target principal view space feature, the fourth bird's-eye view space feature, and the third principal view space feature through the second perspective projection module to obtain a second depth information feature;

[0168] Projecting the second depth information feature onto N depth ranges to obtain second bird's-eye view sub-features of the N depth ranges.

[0169] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0170] The communication interface is used for communication between the above-mentioned terminal and other devices.

[0171] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0172] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0173] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the image processing method based on perspective projection described in the above-mentioned embodiment.

[0174] Those of ordinary skill in the art can understand that all or part of the steps in implementing the method of the above-mentioned embodiment can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes the steps of the image processing method based on perspective projection. The storage medium, such as: ROM / RAM, magnetic disk, optical disc, etc.

[0175] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the image processing method based on perspective projection described in the above embodiment.

[0176] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

[0177] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise", or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0178] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, reference can be made to the corresponding description in the method embodiment.

[0179] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. An image processing method based on perspective projection, characterized in that, Including: Obtaining the camera parameters of a target camera and a main view image captured by the target camera; Encoding the main view image through a main view image encoding network to obtain N first main view spatial features, where the N first main view spatial features are main view spatial features of N resolutions; Processing the N first main view spatial features through a first perspective projection module to obtain a first bird's-eye view spatial feature, and processing a target main view spatial feature among the N first main view spatial features through a second perspective projection module to obtain a second bird's-eye view spatial feature; the target main view spatial feature is one of the N first main view spatial features; Performing feature fusion processing on the first bird's-eye view spatial feature and the second bird's-eye view spatial feature through pixel-by-pixel superposition to obtain a third bird's-eye view spatial feature; Decoding the third bird's-eye view spatial feature through a bird's-eye view image decoding network to obtain an object segmentation image in the bird's-eye view space; Wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

2. The method according to claim 1, wherein The step of processing the N first main view spatial features through the first perspective projection module to obtain a first bird's-eye view spatial feature includes: Processing the N first main view spatial features through the first perspective projection module to obtain N first bird's-eye view sub-features; Performing feature splicing processing on the N first bird's-eye view sub-features to obtain the first bird's-eye view spatial feature.

3. The method according to claim 2, wherein The step of processing the N first main view spatial features through the first perspective projection module to obtain N first bird's-eye view sub-features includes: Performing feature processing on the N first main view spatial features through the first perspective projection module to obtain a first depth information feature; Projecting the first depth information feature onto N depth ranges to obtain first bird's-eye view sub-features of N depth ranges.

4. The method according to claim 3, characterized in that, The step of performing feature processing on the N first main view spatial features through the first perspective projection module to obtain a first depth information feature includes: Performing feature processing on the N first main view spatial features through a first convolutional network to obtain N two-dimensional spatial features; Performing compression processing on the N two-dimensional spatial features in the height dimension to obtain N second main view spatial features; Performing feature transformation processing on the N second main view spatial features through a second convolutional network to obtain N one-dimensional spatial features; Performing resampling processing on the N one-dimensional spatial features to obtain a first depth information feature.

5. The method according to claim 1, characterized in that, The step of processing the target main view spatial feature among the N first main view spatial features through the second perspective projection module to obtain a second bird's-eye view spatial feature includes: Processing the target main view spatial feature among the N first main view spatial features through the second perspective projection module to obtain N second bird's-eye view sub-features; Performing feature splicing processing on the N second bird's-eye view sub-features to obtain the second bird's-eye view spatial feature.

6. The method according to claim 5, characterized in that, Processing the target main view space feature among the N first main view space features through the second perspective projection module to obtain N second bird's-eye view sub-features, including: Performing feature transformation processing on the target main view space feature among the N first main view space features through a first multi-layer perceptron network to obtain a fourth bird's-eye view space feature; Performing feature transformation processing on the fourth bird's-eye view space feature through a second multi-layer perceptron network to obtain a third main view space feature; Inputting the target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features.

7. The method according to claim 6, wherein The step of inputting the target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature into the second perspective projection module for processing to obtain the N second bird's-eye view sub-features includes: Performing feature processing on the target main view space feature, the fourth bird's-eye view space feature, and the third main view space feature through the second perspective projection module to obtain a second depth information feature; Projecting the second depth information feature onto N depth ranges to obtain second bird's-eye view sub-features of the N depth ranges.

8. An image processing device based on perspective projection, characterized in that, Including: A first acquisition module, configured to acquire camera parameters of a target camera and a main view image captured by the target camera; A first processing module, configured to encode the main view image through a main view image encoding network to obtain N first main view space features, where the N first main view space features are main view space features of N resolutions; A second processing module, configured to process the N first main view space features through a first perspective projection module to obtain a first bird's-eye view space feature, and process the target main view space feature among the N first main view space features through a second perspective projection module to obtain a second bird's-eye view space feature; the target main view space feature is one of the N first main view space features; A third processing module, configured to perform feature fusion processing on the first bird's-eye view space feature and the second bird's-eye view space feature through pixel-by-pixel superposition to obtain a third bird's-eye view space feature; A fourth processing module, configured to decode the third bird's-eye view space feature through a bird's-eye view image decoding network to obtain an object segmentation image of the bird's-eye view space; Wherein, the first perspective projection module is a perspective projection module using camera parameters, and the second perspective projection module is a perspective projection module not using camera parameters.

9. An electronic device, characterized in that, Including: A processor, a communication interface, a memory, and a communication bus; wherein, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing a computer program; The processor, when executing the program stored on the memory, implements the steps in the image processing method based on perspective projection according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the perspective-projection-based image processing method according to any one of claims 1 to 7.