Method and device for converting 2D (two-dimensional) spatial features into 3D (three-dimensional) spatial features

By constructing a lookup table for surround-viewing multi-camera images and 3D spatial reference point coordinates, and only calculating multi-scale deformable attention operation of 3D points covered by the camera, the problem of low operating efficiency of 2D to 3D spatial feature conversion in the prior art is solved, and more efficient feature representation and utilization of computing resources are achieved.

CN120147658APending Publication Date: 2025-06-13SHENZHEN DEEPROUTE AI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311705563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the process of converting 2D spatial features to 3D spatial features, there is a problem of low computing power platform operation efficiency.

Method used

By acquiring the multi-camera image of the surround view, the correspondence between the coordinates of the 2D pixel spatial reference point and the coordinates of the 3D spatial reference point under each surround view camera is calculated, a lookup table is constructed, and only the multi-scale deformable attention operation of the points covered by the camera range of the 3D space is calculated.

Benefits of technology

The calculation process of multi-scale deformable attention parameters of irrelevant points is reduced, and the operation efficiency of the perceptual model in low-computing platform is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147658A_ABST
    Figure CN120147658A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for converting 2D spatial features into 3D spatial features, and the method comprises the steps: obtaining a look-around multi-camera image, and carrying out the processing of the look-around multi-camera image, and obtaining the multi-camera multi-scale features of a 2D pixel space; obtaining a 3D space reference point coordinate, and calculating to obtain a 2D pixel space reference point coordinate under each look-around camera according to the 3D space reference point coordinate; determining 2D pixel space reference point coordinates meeting conditions and corresponding 3D space reference point coordinates under each look-around camera, and constructing a lookup table of the 2D pixel space reference point coordinates and the 3D space reference point coordinates; according to the lookup table, multi-scale deformable attention operation of points, covered by the camera range, in the 3D space is calculated, and feature representation in the final aerial view space is obtained. According to the method, the calculation process of the multi-scale deformable attention parameters of the irrelevant points can be reduced, the feature representation in the final aerial view space is obtained, and the operation efficiency of the perceptual model on a low-calculation-power platform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a method and device for converting 2D spatial features to 3D spatial features. Background Art

[0002] In the autonomous driving scenario, multi-camera perception is required. Placing the perception model in the BEV (Bird's Eye View) space can efficiently unify the features of different cameras in the BEV space for expression, facilitating the fusion of multi-camera features and the fusion of different modality features. Among them, a key step is the conversion of 2D spatial features to 3D spatial features.

[0003] Currently, in the process of converting 2D spatial features to 3D spatial features, the commonly used conversion methods are mainly divided into two types: the method based on forward projection and the method based on reverse sampling. Since the method based on forward projection has problems such as large video memory occupation and sparse feature representation, most of the current perception models use the method based on reverse sampling. And the method based on reverse sampling needs to perform MSDA (Multi-Scale Deformable Attention) processing on each feature point in the BEV space. Therefore, this type of method also has the problem of low operation efficiency on low-computing-power platforms.

[0004] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that, aiming at the defects of the existing technology, the present invention provides a method and device for converting 2D spatial features to 3D spatial features to solve the problem of low operation efficiency of the existing perception model on low-computing-power platforms.

[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0007] In a first aspect, the present invention provides a method for converting 2D spatial features to 3D spatial features, including:

[0008] Obtain panoramic multi-camera images, and process the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space;

[0009] Obtain 3D space reference point coordinates, and calculate 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates;

[0010] Determine the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates under each panoramic camera, and construct a lookup table of 2D pixel space reference point coordinates and 3D space reference point coordinates;

[0011] According to the lookup table, calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range to obtain the feature representation in the final bird's-eye view space.

[0012] In one implementation, the obtaining the panoramic multi-camera images and processing the panoramic multi-camera images to obtain the multi-camera multi-scale features in the 2D pixel space includes:

[0013] Obtain N panoramic multi-camera images;

[0014] Input the N panoramic multi-camera images into a perception model, and extract the multi-camera multi-scale features in the 2D pixel space through a backbone network and a feature pyramid.

[0015] In one implementation, before the obtaining the 3D space reference point coordinates, it includes:

[0016] Generate the 3D space reference point coordinates based on the predefined representation range of the bird's-eye view features in the 3D space and the representation granularity size of each bird's-eye view feature pixel point.

[0017] In one implementation, the calculating the 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates includes:

[0018] Right-multiply the 3D space reference point coordinates by the external parameters of the panoramic multi-camera to obtain the point coordinates in the panoramic camera coordinate system corresponding to the 3D space reference point coordinates;

[0019] Right-multiply the point coordinates in the panoramic camera coordinate system by the internal parameters of the panoramic multi-camera to obtain the 2D pixel space reference point coordinates under each panoramic camera.

[0020] In one implementation, the determining the 2D pixel space reference point coordinates that meet the conditions under each panoramic camera and the corresponding 3D space reference point coordinates, and constructing a lookup table of the 2D pixel space reference point coordinates and the 3D space reference point coordinates includes:

[0021] Filter the 2D pixel space reference point coordinates that meet the conditions under each panoramic camera;

[0022] Look up the 3D space reference point coordinates corresponding to the bird's-eye view space according to the 2D pixel space reference point coordinates that meet the conditions;

[0023] Construct the lookup table according to the relationship between the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates.

[0024] In one implementation, the filtering the 2D pixel space reference point coordinates that meet the conditions under each panoramic camera includes:

[0025] Determine whether the values of the coordinates of each 2D pixel space reference point exceed the length and width of the image, and determine whether the values of the coordinates of each 2D pixel space reference point are less than 0;

[0026] If it does not exceed the length and width of the image and the value is not less than 0, then determine that the corresponding 2D pixel space reference point coordinate is the 2D pixel space reference point coordinate that meets the conditions.

[0027] In one implementation manner, the calculating the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the look-up table to obtain the feature representation in the final bird's-eye view space includes:

[0028] Calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the key-value correspondence relationship in the look-up table to obtain the feature representation in the final bird's-eye view space;

[0029] Output the obtained feature representation in the final bird's-eye view space.

[0030] In a second aspect, the present invention provides a conversion device for 2D space features to 3D space features, including:

[0031] An image acquisition module, configured to acquire a panoramic multi-camera image and process the panoramic multi-camera image to obtain multi-camera multi-scale features in the 2D pixel space;

[0032] A reference point coordinate module, configured to acquire 3D space reference point coordinates and calculate 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates;

[0033] A look-up table module, configured to determine the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates under each panoramic camera, and construct a look-up table of 2D pixel space reference point coordinates and 3D space reference point coordinates;

[0034] A feature output module, configured to calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the look-up table to obtain the feature representation in the final bird's-eye view space.

[0035] In a third aspect, the present invention provides a terminal, including: a processor and a memory, the memory stores a conversion program for 2D space features to 3D space features, and when the conversion program for 2D space features to 3D space features is executed by the processor, it is used to implement the operations of the 2D space feature to 3D space feature conversion method described in the first aspect.

[0036] Fourthly, the present invention further provides a medium, which is a computer-readable storage medium. The medium stores a conversion program from 2D spatial features to 3D spatial features. When the conversion program from 2D spatial features to 3D spatial features is executed by a processor, it is used to implement the operations of the method for converting 2D spatial features to 3D spatial features as described in the first aspect.

[0037] The present invention adopts the above technical solutions and has the following effects:

[0038] The present invention considers the effective representation range of cameras with different perspectives in the bird's-eye view space, calculates the one-to-one correspondence between the 2D pixel coordinate system of each camera and the points in the 3D ego-vehicle coordinate system, and constructs a lookup table; then, according to the lookup table, only calculates the multi-scale deformable attention operations of the points covered by the camera range in the 3D space, thereby completing the feature representation of the perception model in the bird's-eye view space and realizing the conversion operation of features from 2D to 3D space. The present invention can reduce the calculation process of the multi-scale deformable attention parameters of irrelevant points, obtain the final feature representation in the bird's-eye view space, and improve the operation efficiency of the perception model on low-computing-power platforms. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the structures shown in these drawings.

[0040] Figure 1 is a flowchart of the method for converting 2D spatial features to 3D spatial features in an implementation manner of the present invention.

[0041] Figure 2 is a processing schematic diagram of the conversion algorithm from 2D spatial features to 3D spatial features in an implementation manner of the present invention.

[0042] Figure 3 is a functional schematic diagram of a terminal in an implementation manner of the present invention.

[0043] The implementation, functional features, and advantages of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments

[0044] To make the purpose, technical solutions, and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] Exemplary Method

[0046] Currently, in the process of converting 2D spatial features to 3D spatial features, the commonly used conversion methods are mainly divided into two types: the method based on forward projection and the method based on reverse sampling. Since the method based on forward projection has problems such as large video memory occupation and sparse feature representation, most current perception models use the method based on reverse sampling. And the method based on reverse sampling needs to perform MSDA (Multi-Scale Deformable Attention) processing on each feature point in the BEV space. Therefore, this type of method also has the problem of low operating efficiency on low-computing-power platforms.

[0047] In view of the above technical problems, an embodiment of the present invention provides a method for converting 2D spatial features to 3D spatial features. This method considers the effective representation range of different perspective cameras in the bird's-eye view space, calculates the one-to-one correspondence between the 2D pixel coordinate system of each camera and the points in the 3D ego-vehicle coordinate system, and constructs a lookup table; then, according to the lookup table, only calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range, so as to complete the feature representation of the perception model in the bird's-eye view space and realize the conversion operation of features from 2D to 3D space. Therefore, the embodiment of the present invention can reduce the calculation process of the multi-scale deformable attention parameters of irrelevant points, obtain the final feature representation in the bird's-eye view space, and improve the operating efficiency of the perception model on low-computing-power platforms.

[0048] As Figure 1 shown, an embodiment of the present invention provides a method for converting 2D spatial features to 3D spatial features, including the following steps:

[0049] Step S100, obtain panoramic multi-camera images, and process the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space.

[0050] In this embodiment, in the autonomous driving scenario, multi-camera perception is required. By placing the perception model in the bird's-eye view space, the features of different cameras can be efficiently unified and expressed in the bird's-eye view space, which is convenient for the feature fusion of multi-cameras and the fusion of different modality features. Among them, a key step is the conversion of 2D pixel space features to 3D spatial features.

[0051] In this embodiment, in order to improve the conversion efficiency of 2D pixel space features to 3D spatial features, a lookup table of coordinate points between 2D and 3D spaces is constructed. Then, according to the lookup table, only calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range, so as to complete the feature representation of the perception model in the bird's-eye view space and efficiently realize the conversion operation of features from 2D to 3D space.

[0052] Specifically, in one implementation of this embodiment, step S100 includes the following steps:

[0053] Step S101, obtain N panoramic multi-camera images;

[0054] Step S102, input the N panoramic multi-camera images into a perception model, and obtain the multi-camera multi-scale features in the 2D pixel space through extraction by a backbone network and a feature pyramid.

[0055] In this embodiment, by obtaining N panoramic multi-camera images, inputting the N panoramic multi-camera images into a perception model, and obtaining the multi-camera multi-scale features in the 2D pixel space through extraction by a backbone network and a feature pyramid, the shape is B*N*C*H*W. Wherein, B represents the batch number, N represents the number of panoramic cameras, C represents the number of feature channels, and H and W respectively represent the length and width of the multi-camera multi-scale features in the 2D pixel space.

[0056] In this embodiment, by obtaining the multi-camera multi-scale features in the 2D pixel space, the 2D pixel space reference point coordinates under each panoramic camera can be calculated according to the defined 3D space reference point coordinates and the internal and external parameters of each panoramic camera, so as to consider the effective representation range of cameras with different perspectives in the bird's-eye view space.

[0057] As Figure 1 shown, in one implementation of the embodiment of the present invention, the method for converting 2D space features to 3D space features further includes the following steps:

[0058] Step S200, obtain 3D space reference point coordinates, and calculate the 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates.

[0059] In this embodiment, for the multi-camera multi-scale features in the 2D pixel space obtained in the above step S100, the 3D space reference point coordinates can be set first, and then the 2D pixel space reference point coordinates under each panoramic camera can be calculated according to these 3D space reference point coordinates.

[0060] Specifically, in one implementation of this embodiment, before step S200 includes the following steps:

[0061] Step S201a, generate the 3D space reference point coordinates based on the pre-defined representation range of the 3D space bird's-eye view features and the representation granularity size of each bird's-eye view feature pixel point.

[0062] In this embodiment, a series of reference point coordinates in the bird's-eye view space (i.e., the 3D space reference point coordinates) are generated according to the predefined representation range of the 3D space bird's-eye view and the representation granularity of each pixel on each bird's-eye view feature. Among them, the bird's-eye view space range and the representation granularity of each pixel on the bird's-eye view feature can be set in the visual bird's-eye view model. Generally, one bird's-eye view point represents 0.5 meters in the real world.

[0063] In this embodiment, there are X * Y * Z 3D space reference point coordinates in total, where X represents the length of the bird's-eye view feature, Y represents the width of the bird's-eye view feature, and Z represents the height of the bird's-eye view feature; each 3D space reference point coordinate corresponds to the central coordinate of each pixel of the bird's-eye view feature.

[0064] As Figure 2 shown, multiplying these 3D space reference point coordinates on the right by the external parameters of the surround multi-camera can obtain the point coordinates in the camera coordinate system corresponding to these 3D space reference point coordinates. Then, multiplying the point coordinates in the camera coordinate system on the right by the internal parameters of the surround multi-camera can obtain the point coordinates in the 2D pixel space.

[0065] Specifically, in one implementation manner of this embodiment, step S200 includes the following steps:

[0066] Step S201, multiplying the 3D space reference point coordinates on the right by the external parameters of the surround multi-camera to obtain the point coordinates in the surround camera coordinate system corresponding to the 3D space reference point coordinates;

[0067] Step S202, multiplying the point coordinates in the surround camera coordinate system on the right by the internal parameters of the surround multi-camera to obtain the 2D pixel space reference point coordinates for each surround camera.

[0068] In this embodiment, multiplying the obtained X * Y * Z 3D space reference point coordinates on the right by the external camera parameters of the surround multi-camera to obtain the corresponding reference points in the surround camera coordinate system.

[0069] As an example, if there are N surround-view cameras, the extrinsic parameters of the surround-view multi-camera are in the shape of N*4*4. Through the reshape operation, it becomes in the shape of N*1*4*4, and then, through the repeat operation, it becomes in the shape of N*(X*Y*Z)*4*4, denoted as A. The shape corresponding to X*Y*Z 3D space reference points is X*Y*Z*4. Through the reshape operation, it becomes in the shape of 1*(X*Y*Z)*4*1, and then through the repeat operation, it becomes in the shape of N*(X*Y*Z)*4*1, denoted as B. Finally, the reference point C corresponding to the surround-view camera coordinate system is calculated by the function torch.bmm(A,B), and its shape is N*(X*Y*Z)*4*1). The reference point C is the point coordinate in the surround-view camera coordinate system corresponding to the 3D space reference point coordinate.

[0070] Further, based on the reference point corresponding to the surround-view camera coordinate system, multiplying the point coordinate in the surround-view camera coordinate system by the intrinsic parameters of the surround-view multi-camera on the right can obtain the 2D pixel space reference point coordinates under each surround-view camera.

[0071] As an example, if there are N surround-view cameras, the intrinsic parameters of the surround-view multi-camera are in the shape of N*3*4. Through the reshape operation, it becomes in the shape of N*1*3*4, and then through the repeat operation, it becomes in the shape of N*(X*Y*Z)*3*4, denoted as D. Combining with the reference point C obtained in the above process, using the function torch.bmm(D,C), the 2D pixel space reference point coordinates E under each surround-view camera are calculated, and its shape is N*(X*Y*Z)*3*1; for the 2D pixel space reference point coordinates E, by discarding the last dimension, it becomes N*(X*Y*Z)*3), and the 2D pixel space reference point coordinates under each surround-view camera are obtained.

[0072] In this embodiment, by calculating the 2D pixel space reference point coordinates under each surround-view camera, the effective representation range of different perspective cameras in the bird's-eye view space can be determined, and a lookup table is constructed according to the one-to-one correspondence between the 2D pixel coordinate system under each surround-view camera and the points in the 3D ego-vehicle coordinate system.

[0073] As Figure 1 shown, in an implementation manner of the embodiment of the present invention, the method for converting 2D space features to 3D space features further includes the following steps:

[0074] Step S300, determine the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates under each surround-view camera, and construct a lookup table of the 2D pixel space reference point coordinates and the 3D space reference point coordinates.

[0075] In this embodiment, the 2D pixel space reference point coordinates obtained for each surround-view camera are screened to obtain the valid 2D space reference point coordinates in the 2D pixel space reference point coordinates, thereby removing the invalid 2D space reference point coordinates and the invalid 3D reference points corresponding to the bird's-eye view space. This can avoid the calculation of invalid 3D reference points during the conversion process from 2D pixel space features to 3D space features and reduce the overhead of computing resources.

[0076] Specifically, in an implementation manner of this embodiment, step S300 includes the following steps:

[0077] Step S301, screening the 2D pixel space reference point coordinates that meet the conditions for each surround-view camera.

[0078] Specifically, the screening of the 2D pixel space reference point coordinates that meet the conditions for each surround-view camera includes the following steps:

[0079] Step S301a, determining whether the values of the 2D pixel space reference point coordinates exceed the length and width of the image, and determining whether the values of the 2D pixel space reference point coordinates are less than 0;

[0080] Step S301b, if the values do not exceed the length and width of the image and are not less than 0, determining that the corresponding 2D pixel space reference point coordinates are the 2D pixel space reference point coordinates that meet the conditions.

[0081] In this embodiment, for the 2D pixel space reference point coordinates calculated in the above step S200 for each surround-view camera, the valid 2D space reference point coordinates are first screened and determined; that is, the valid 2D space reference point coordinates for each surround-view camera are determined, and the invalid 3D space reference points corresponding to the invalid 2D space reference point coordinates in the bird's-eye view space are determined.

[0082] As an example, for the determination condition of the valid 2D space reference point coordinates, it can be first determined whether the value of the 2D space reference point coordinates exceeds the length and width of the image, and then it is determined whether the value of the 2D space reference point coordinates is less than 0; if either of the two conditions is met, it means that the corresponding 2D space reference point is an invalid point; if neither of the two conditions is met, it means that the corresponding 2D space reference point is a valid point; among them, the determined valid 2D space reference point coordinates are the coordinates required for establishing the lookup table.

[0083] Before multi-camera fusion, the shape of the coordinate points in the 3D space is N*(X*Y*Z)*4, where the N*(X*Y*Z) part and the N*(X*Y*Z) part in the 2D reference point coordinates N*(X*Y*Z)*3 are in one-to-one correspondence. Therefore, if the invalid 2D space reference point coordinates are found, the corresponding 3D space coordinate points at the corresponding positions are also found.

[0084] Specifically, in one implementation of this embodiment, step S300 further includes the following steps:

[0085] Step S302, find the 3D space reference point coordinates corresponding to the 2D pixel space reference point coordinates that meet the conditions in the bird's-eye view space;

[0086] Step S303, construct the lookup table according to the relationship between the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates.

[0087] In this embodiment, based on the valid 2D pixel space reference point coordinates that meet the conditions under each surround-view camera, first find the 3D space reference point coordinates corresponding to these 2D pixel space reference point coordinates in the bird's-eye view space, then establish the corresponding relationship between the two, and construct the lookup table using the relationship between these 2D pixel space reference point coordinates and the corresponding 3D space reference point coordinates; where the key-value correspondence in the lookup table is: 3D valid space reference point coordinates -> 2D pixel space reference point coordinates.

[0088] In this embodiment, by constructing the lookup table, the calculation of invalid 3D reference points during the conversion from 2D pixel space features to 3D space features can be avoided, reducing the overhead of computing resources.

[0089] As Figure 1 shown, in one implementation of the embodiment of the present invention, the method for converting 2D space features to 3D space features further includes the following steps:

[0090] Step S400, calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the lookup table, and obtain the feature representation in the final bird's-eye view space.

[0091] In this embodiment, based on the constructed lookup table, during the conversion from 2D pixel space features to 3D space features, only calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range to obtain the feature representation in the final bird's-eye view space, which greatly reduces the amount of calculation and the overhead of computing resources.

[0092] Specifically, in one implementation of this embodiment, step S400 includes the following steps:

[0093] Step S401, calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the key-value correspondence in the lookup table, and obtain the feature representation in the final bird's-eye view space;

[0094] Step S402, output the obtained feature representation in the final bird's-eye view space.

[0095] In this embodiment, for the 3D coordinate points in the conversion process from 2D pixel space features to 3D space features, by looking up the key-value in the lookup table, the corresponding 2D space reference point coordinates are found, and then only the multi-scale deformable attention operation of the points in the 3D space covered by the camera range is calculated to obtain the feature representation in the final bird's-eye view space.

[0096] Since the key in the lookup table corresponds to the index of the points in the 3D space covered by the camera in all 3D space reference points, the indices are directly taken one by one in parallel from the lookup table, and then based on the key-value correspondence, the corresponding 2D space reference point coordinates are obtained, and then only the multi-scale deformable attention operation of the points in the 3D space covered by the camera range is calculated.

[0097] In this embodiment, only the multi-scale deformable attention operation of the points in the 3D space covered by the camera range is calculated, which can reduce the calculation of the multi-scale deformable attention of irrelevant points and obtain the feature representation in the final bird's-eye view space. When starting the CUDA (a computing platform launched by graphics card manufacturer NVIDIA) kernel function, the total thread resources applied can be reduced from the original B*N*X*Y*Z threads to the number of pixels of the BEV within the effective camera range. In actual business, the thread resource overhead can be reduced by more than 30%.

[0098] In summary, in this embodiment, a large amount of invalid multi-scale deformable attention calculations of 3D points in the bird's-eye view space are reduced, a large amount of thread computing resources are released, the efficiency of the 2D feature to 3D feature conversion is effectively improved, and it is more easily deployed to low-computing power platforms.

[0099] This embodiment achieves the following technical effects through the above technical solutions:

[0100] This embodiment considers the effective representation range of different perspective cameras in the bird's-eye view space, constructs a lookup table by calculating the one-to-one correspondence between the 2D pixel coordinate system of each camera and the points in the 3D ego vehicle coordinate system; and then according to the lookup table, only calculates the multi-scale deformable attention operation of the points in the 3D space covered by the camera range, thereby completing the feature representation of the perception model in the bird's-eye view space and realizing the conversion operation of the feature from 2D to 3D space. This embodiment can reduce the calculation process of the multi-scale deformable attention parameters of irrelevant points, obtain the feature representation in the final bird's-eye view space, and improve the operation efficiency of the perception model on low-computing power platforms.

[0101] Exemplary device

[0102] Based on the above embodiment, the present invention further provides a conversion device for 2D space features to 3D space features, including:

[0103] An image acquisition module, configured to acquire panoramic multi-camera images, and process the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space;

[0104] A reference point coordinate module, configured to acquire 3D space reference point coordinates, and calculate 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates;

[0105] A look-up table module, configured to determine 2D pixel space reference point coordinates that meet the conditions and corresponding 3D space reference point coordinates under each panoramic camera, and construct a look-up table of 2D pixel space reference point coordinates and 3D space reference point coordinates;

[0106] A feature output module, configured to calculate the multi-scale deformable attention operation of points in the 3D space covered by the camera range according to the look-up table, and obtain the feature representation in the final bird's-eye view space.

[0107] Through the above technical solutions, the present embodiment achieves the following technical effects:

[0108] This embodiment considers the effective representation range of cameras with different perspectives in the bird's-eye view space. By calculating the one-to-one correspondence between the 2D pixel coordinate system of each camera and the points in the 3D ego-vehicle coordinate system, a look-up table is constructed; then, according to the look-up table, only the multi-scale deformable attention operation of points in the 3D space covered by the camera range is calculated, thereby completing the feature representation of the perception model in the bird's-eye view space and realizing the conversion operation of features from 2D to 3D space. This embodiment can reduce the calculation process of the multi-scale deformable attention parameters of irrelevant points, obtain the feature representation in the final bird's-eye view space, and improve the operation efficiency of the perception model on low-computing-power platforms.

[0109] Based on the above embodiments, the present invention further provides a terminal, and its principle block diagram can be as Figure 3 shown.

[0110] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected through a system bus; wherein, the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect external devices; the display screen is used to display corresponding information; the communication module is used to communicate with a cloud server or other devices.

[0111] When the computer program is executed by the processor, it is used to implement the operations of the method for converting 2D space features to 3D space features.

[0112] Those skilled in the art can understand that Figure 3The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0113] In one embodiment, a terminal is provided, which includes: a processor and a memory. The memory stores a conversion program from 2D space features to 3D space features. When the conversion program from 2D space features to 3D space features is executed by the processor, it is used to implement the operations of the above method for converting 2D space features to 3D space features.

[0114] In one embodiment, a storage medium is provided, which stores a conversion program from 2D space features to 3D space features. When the conversion program from 2D space features to 3D space features is executed by the processor, it is used to implement the operations of the above method for converting 2D space features to 3D space features.

[0115] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and volatile memories.

[0116] In summary, the present invention provides a method and device for converting 2D space features to 3D space features. The method includes: obtaining panoramic multi-camera images, processing the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space; obtaining 3D space reference point coordinates, and calculating 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates; determining the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates under each panoramic camera, and constructing a lookup table of 2D pixel space reference point coordinates and 3D space reference point coordinates; according to the lookup table, calculating the multi-scale deformable attention operation of the points covered by the camera range in the 3D space to obtain the feature representation in the final bird's-eye view space. The present invention can reduce the calculation process of the multi-scale deformable attention parameters of irrelevant points, obtain the feature representation in the final bird's-eye view space, and improve the operation efficiency of the perception model on a low-computing-power platform.

[0117] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for converting 2D spatial features to 3D spatial features, characterized in that, it includes: Obtain panoramic multi-camera images, and process the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space; Obtain 3D spatial reference point coordinates, and calculate the 2D pixel space reference point coordinates under each panoramic camera according to the 3D spatial reference point coordinates; Determine the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D spatial reference point coordinates under each panoramic camera, and construct a lookup table of 2D pixel space reference point coordinates and 3D spatial reference point coordinates; According to the lookup table, calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range to obtain the feature representation in the final bird's-eye view space.

2. The method for converting 2D spatial features to 3D spatial features according to claim 1, characterized in that, The obtaining of panoramic multi-camera images and processing the panoramic multi-camera images to obtain multi-camera multi-scale features in the 2D pixel space includes: Obtain N panoramic multi-camera images; Input the N panoramic multi-camera images into a perception model, and extract the multi-camera multi-scale features in the 2D pixel space through a backbone network and a feature pyramid.

3. The method for converting 2D spatial features to 3D spatial features according to claim 1, characterized in that, Before the obtaining of the 3D spatial reference point coordinates, it includes: Generate the 3D spatial reference point coordinates based on the pre-defined representation range of the 3D spatial bird's-eye view features and the representation granularity size of each bird's-eye view feature pixel point.

4. The method for converting 2D spatial features to 3D spatial features according to claim 1, characterized in that, The calculating of the 2D pixel space reference point coordinates under each panoramic camera according to the 3D spatial reference point coordinates includes: Right-multiply the 3D spatial reference point coordinates by the external parameters of the panoramic multi-camera to obtain the point coordinates in the panoramic camera coordinate system corresponding to the 3D spatial reference point coordinates; Right-multiply the point coordinates in the panoramic camera coordinate system by the internal parameters of the panoramic multi-camera to obtain the 2D pixel space reference point coordinates under each panoramic camera.

5. The method for converting 2D spatial features to 3D spatial features according to claim 1, characterized in that, The determining of the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D spatial reference point coordinates under each panoramic camera, and constructing a lookup table of 2D pixel space reference point coordinates and 3D spatial reference point coordinates includes: Screen the 2D pixel space reference point coordinates that meet the conditions under each panoramic camera; Find the 3D spatial reference point coordinates corresponding to the bird's-eye view space according to the 2D pixel space reference point coordinates that meet the conditions; Construct the lookup table according to the relationship between the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D spatial reference point coordinates.

6. The method for converting 2D spatial features to 3D spatial features according to claim 5, characterized in that, The screening of the 2D pixel space reference point coordinates that meet the conditions under each panoramic camera includes: Determine whether the numerical values of the coordinates of each 2D pixel space reference point exceed the length and width of the image, and determine whether the numerical values of the coordinates of each 2D pixel space reference point are less than 0; If it does not exceed the length and width of the image and the value is not less than 0, then determine that the corresponding 2D pixel space reference point coordinate is the 2D pixel space reference point coordinate that meets the conditions.

7. The method for converting 2D space features to 3D space features according to claim 1, characterized in that, The calculating of the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the look-up table to obtain the feature representation in the final bird's-eye view space includes: Calculating the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the key-value correspondence in the look-up table to obtain the feature representation in the final bird's-eye view space; Output the obtained feature representation in the final bird's-eye view space.

8. A device for converting 2D space features to 3D space features, characterized in that, comprising: An image acquisition module, configured to acquire a panoramic multi-camera image and process the panoramic multi-camera image to obtain multi-camera multi-scale features in the 2D pixel space; A reference point coordinate module, configured to acquire 3D space reference point coordinates and calculate 2D pixel space reference point coordinates under each panoramic camera according to the 3D space reference point coordinates; A look-up table module, configured to determine the 2D pixel space reference point coordinates that meet the conditions and the corresponding 3D space reference point coordinates under each panoramic camera, and construct a look-up table of 2D pixel space reference point coordinates and 3D space reference point coordinates; A feature output module, configured to calculate the multi-scale deformable attention operation of the points in the 3D space covered by the camera range according to the look-up table to obtain the feature representation in the final bird's-eye view space.

9. A terminal, characterized in that, comprising: A processor and a memory, the memory stores a conversion program for 2D space features to 3D space features, and when the conversion program for 2D space features to 3D space features is executed by the processor, it is used to implement the operations of the method for converting 2D space features to 3D space features according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a conversion program for 2D space features to 3D space features, and when the conversion program for 2D space features to 3D space features is executed by a processor, it is used to implement the operations of the method for converting 2D space features to 3D space features according to any one of claims 1-7.