Image generation method and electronic device
By projecting a 3D scene on the side of the pyramid to generate environment maps and optimizing distortion models, the problems of insufficient scene coverage and difficulty in obtaining data in the existing technology are solved, and high-quality and efficient simulation image generation is achieved, which improves the recognition performance and training and verification efficiency of the perceptual model.
Patent Information
- Application Number
- PCT/CN2025/072723
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2025-01-16
- Publication Date
- 2025-08-14
AI Technical Summary
When generating simulated images of distorted sensor data, the prior art has problems such as insufficient scene coverage and difficulty in obtaining data, resulting in insufficient recognition performance of the perceptual model and low training and testing verification efficiency.
By projecting the 3D scene of the target device's driving environment on multiple sides of the pyramid, multiple environmental maps are generated, and combined with the Gaussian-Newtonian method to optimize the distortion model, the coordinates of pixel points in the simulation image are determined under the camera coordinates, and the interpolation algorithm is used to fuse the pixel values of multiple environmental maps to generate high-quality simulated images.
The simulation images generated at the same efficiency are higher in quality, or the simulation images generated at the same quality are more efficient, which improves the recognition performance and training and verification efficiency of the perceptual model.
Smart Images

Figure CN2025072723_14082025_PF_FP_ABST
Abstract
Description
Image generation method and electronic device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on February 8, 2024, with application number 202410177218.8 and application name “Image Generation Method and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing, and in particular to an image generation method and electronic device. Background Art
[0003] Sensors are essential hardware for smart vehicles; examples include fisheye cameras and distorted pinhole cameras. Currently, vehicles are equipped with these cameras on the front, rear, left, and right sides of the vehicle. While driving or parking, the onboard system uses a perception model to identify the distances of obstacles, pedestrians, other vehicles, animals, and so on, using the distorted images captured by these cameras. Based on this distance, the system then controls the vehicle (e.g., for automated parking and autonomous driving) or provides user notifications. Therefore, the recognition performance of the perception model is crucial.
[0004] Typically, a large amount of distortion sensor data (i.e., distorted images) and ground truth data (such as the location information of obstacles and pedestrians) is required to train and test the perception model to improve its recognition performance. To address the problems of insufficient scene coverage and difficulty in obtaining distortion sensor data, simulated images (i.e., distorted images obtained through simulation) can be used instead of distorted images captured by the distortion camera to train and test the perception model. This simulation method can significantly reduce the cost of data generation and annotation, and can also improve the efficiency of model training and testing. Summary of the Invention
[0005] In view of this, the present application provides an image generation method and electronic device. Compared to the prior art, this method generates a simulated image of a target distorted camera with higher quality while maintaining the same efficiency; or, while maintaining the same quality, generates a simulated image with higher efficiency.
[0006] In a first aspect, an embodiment of the present application provides an image generation method, which includes: first, obtaining a three-dimensional (3D) scene of the driving environment of a target device; then, projecting the 3D scene onto N sides of a pyramid to obtain N environment maps; wherein N is an integer greater than or equal to 3; then, determining the pixel point corresponding to each pixel point in the M environment maps of the simulated image of the target distorted camera; wherein M is a positive integer less than or equal to N, and the target device includes a target distorted camera; thereafter, determining the pixel value of each pixel point in the simulated image based on the pixel value of the corresponding pixel point in the M environment maps of each pixel point in the simulated image.
[0007] When N is equal to 3, the present application projects the 3D scene of the target device's driving environment onto the three sides of a pyramid to obtain three environment maps. These three environment maps can support the generation of simulated images of a distortion camera with a maximum field of view (FOV) of about 240°; and the actual maximum field of view of a distortion camera is usually about 220°. That is to say, the present application projects the 3D scene onto the three sides of a triangular pyramid to obtain three environment maps, which can support the generation of simulated images of all (or most) distortion cameras. The prior art, on the other hand, projects the 3D scene of the target device's driving environment onto the five or six faces of a cube to obtain five or six environment maps; these five or six environment maps can support the generation of simulated images of a distortion camera with a maximum field of view of about 270°. It can be seen that the prior art generates five or six environment maps with redundancy in field of view.
[0008] When N is equal to 3, and each pair of adjacent side surfaces of the three pairs of triangular pyramids are orthogonal, each pixel in the simulated image of the target distortion camera in the present application has a corresponding pixel in one environment map; each pixel in the simulated image of the target distortion camera in the prior art also has a corresponding pixel in one environment map; that is, the pixel value of a pixel in the simulated image generated by the present application and the prior art is determined based on the pixel value of one pixel; it can be seen that the quality of the simulated image generated by the present application is the same as the quality of the simulated image generated by the prior art. In summary, when N is equal to 3, and each pair of adjacent side surfaces of the three pairs of triangular pyramids are orthogonal, the quality of the simulated image generated by the present application is the same as the quality of the simulated image generated by the prior art, but the number of environment maps required to be generated by the present application is less, so the efficiency of generating simulated images by the present application is higher.
[0009] When N is equal to 3 and at least one of the three pairs of adjacent side surfaces of the triangular pyramid intersects, some of the pixels in the simulated image of the target distortion camera in the present application have corresponding pixels in one environment map, and some of the other pixels have corresponding pixels in two environment maps. In other words, the pixel values of some of the pixels in the simulated image of the present application are determined based on the pixel values of two pixels, wherein the more pixels in the environment map used to determine the pixel values of the pixels of the simulated image, the higher the probability that the pixel values of the pixels in the multiple pixels of the environment map are close to the true pixel values of the pixels in the simulated image. In this way, the pixel values of the pixels in the determined simulated image are more likely to be more accurate; it can be seen that the quality of the simulated image generated by the present application is higher than the quality of the simulated image generated by the prior art. In summary, when N is equal to 3 and at least one of the three pairs of adjacent side surfaces of the triangular pyramid intersects, the quality of the simulated image generated by the present application is not only higher than the quality of the simulated image generated by the prior art, but also the efficiency of generating simulated images by the present application is higher.
[0010] When N is equal to 4, the present application projects the 3D scene of the target device's driving environment onto the four sides of a pyramid to obtain four environment maps. These four environment maps can support the generation of simulated images of a distortion camera with a maximum field of view angle slightly greater than 240°; and the maximum actual field of view angle of a distortion camera is usually around 220°, that is, the four environment maps obtained by projecting the 3D scene onto the four sides of a quadrangular pyramid in the present application can support the generation of simulated images of all (or most) distortion cameras. In addition, in the present application, some pixels in the simulated image of the target distortion camera have corresponding pixels in one environment map, and some other pixels have corresponding pixels in two environment maps; that is, the pixel values of some pixels in the simulated image of the present application are determined based on the pixel values of two pixels; it can be seen that the quality of the simulated image generated by the present application is higher than that of the simulated image generated by the prior art.
[0011] In summary, when N is equal to 4, the quality of the simulated image generated by the present application is not only higher than that of the simulated image generated by the prior art, but the efficiency of generating the simulated image by the present application is also higher.
[0012] When N is equal to 5 or 6, the present application projects the 3D scene of the target device's driving environment onto 5 or 6 faces of a pyramid to obtain 5 or 6 environment maps; these 5 or 6 environment maps can support the generation of simulated images of a distortion camera with a maximum field of view of about 260°; and the actual maximum field of view of a distortion camera is usually about 220°, that is, the 5 or 6 environment maps obtained by projecting the 3D scene onto 5 or 6 faces of a pyramid in the present application can support the generation of simulated images of all (or most) distortion cameras. In addition, in the simulated image of the target distortion camera in the present application, some pixels have corresponding pixels in one environment map, and some other pixels have corresponding pixels in multiple environment maps; that is, the pixel values of some pixels in the simulated image of the present application are determined based on the pixel values of multiple pixels (including 2 or more pixels); it can be seen that the quality of the simulated image generated by the present application is higher than that of the simulated image generated by the prior art.
[0013] In summary, when N is equal to 5 or 6, the efficiency of the simulated images generated by the present application is the same as that of the prior art, but the quality of the simulated images generated by the present application is higher.
[0014] When N is greater than 6, although this application needs to generate more environment maps and requires higher computing power, the quality of the generated simulation image is also higher, which is something that cannot be achieved by the existing technology of projecting the 3D scene of the target device's driving environment onto 5 or 6 faces of a cube.
[0015] For example, the target device may include a vehicle, a delivery robot, a food delivery robot, an airport guide robot, a shopping mall service robot, etc., and this application does not limit this. This application takes the target device as an example of a vehicle for explanation.
[0016] For example, when the target device is a vehicle, the 3D scene of the vehicle driving environment may include: a 3D scene of a parking lot, a 3D scene of a city road, a 3D scene of a suburban road, a 3D scene of a mountain road, a 3D scene of a highway, etc., and this application does not impose any restrictions on this.
[0017] Exemplarily, when the target device is a food delivery robot, the 3D scene of the food delivery robot's driving environment may include: a 3D scene of an apartment, a 3D scene of an office building, etc.
[0018] Exemplarily, when the target device is an airport guidance robot, the 3D scene of the airport guidance robot's driving environment may include: a 3D scene of an airport waiting room, etc.
[0019] Exemplarily, when the target device is a shopping mall service robot, the 3D scene of the shopping mall service robot's driving environment may include: a 3D scene of the shopping mall, etc.
[0020] It should be noted that since the environment map generated based on the 3D scene must be rectangular, this application expands each side of the pyramid into a rectangle. The number of sides of the pyramid can be represented by N, and the pyramid can include N pairs of adjacent sides, and the angle between each pair of adjacent sides can be represented by α.
[0021] In the embodiment of the present application, N may be an integer greater than or equal to 3; and the present application does not limit the size of the side angle α between each pair of adjacent side surfaces.
[0022] For example, if N=3, the angle α between side 1 and side 2 is 90°, the angle α between side 2 and side 3 is 90°, and the angle α between side 3 and side 1 is 90°. The triangular pyramid in this case can also be called an orthogonal triangular prism. For another example, if N=3, the angle α between side 1 and side 2 is 60°, the angle α between side 2 and side 3 is 60°, and the angle α between side 3 and side 1 is 60°. For another example, if N=4, the angle α between side 1 and side 2 is 90°, the angle α between side 2 and side 3 is 60°, the angle α between side 3 and side 4 is 60°, and the angle α between side 4 and side 1 is 60°. And so on. The examples are not listed here one by one.
[0023] Exemplarily, a plurality of distortion cameras may be arranged on the target device, and the target distortion camera may be any one of the plurality of distortion cameras arranged on the target device.
[0024] It should be understood that the image generation method of the present application can generate a distorted image for a target distorted camera. In other words, the image generation method of the present application can simulate a distorted image of a target distorted camera, that is, a simulated image of a target distorted camera. In other words, the simulated image of a target distorted camera can refer to a distorted image of the target distorted camera obtained through simulation.
[0025] For example, the size of the simulated image of the target distorted camera is the same as the size of the distorted image captured by the target distorted camera. Pixels in the simulated image of the target distorted camera correspond one-to-one to pixels in the distorted image captured by the target distorted camera.
[0026] For example, the environment map is also called a material map in some scenes.
[0027] Exemplarily, the pixel values of each pixel of the "simulated image" in "determining the pixel points corresponding to each pixel point in the simulated image of the target distortion camera in the M environment maps" may be preset values. For example, the simulated image is a blank image. "Determining the pixel value of each pixel in the simulated image based on the pixel values of the corresponding pixel points in the M environment maps for each pixel point in the simulated image" may be implemented by determining a target value based on the pixel values of the corresponding pixel points in the M environment maps for each pixel point in the simulated image; and updating the pixel values of all pixels in the simulated image from the preset values to the target values. In other words, the pixel values of each pixel in the resulting simulated image are the target values.
[0028] Illustratively, the field of view of the target distortion camera of the first aspect is less than or equal to 260°.
[0029] For example, the target device includes a target distortion camera. This can be understood as the target distortion camera being currently installed on the target device, or being installed on the target device at some point in the future. In other words, the target distortion camera does not necessarily have to be permanently installed on the target device, but can be installed (or mounted) on the target device when needed. The target distortion camera can be installed in the front, back, left, or right of the target device, and this application does not impose any restrictions on this.
[0030] Exemplarily, the target device includes a target distortion camera, which can also be understood as the target distortion camera being a part of the target device.
[0031] According to the first aspect, the pixel value of each pixel point in the simulated image is determined based on the pixel value of the corresponding pixel point in M environment maps of each pixel point in the simulated image, including: when M is equal to 1, the pixel value of the corresponding pixel point of the first pixel point in an environment map is determined as the pixel value of the first pixel point; wherein the first pixel point is any pixel point in the simulated image.
[0032] That is, for any pixel point in the simulated image, when the pixel point has a corresponding pixel point in an environment map, the pixel value of the corresponding pixel point in the environment map can be determined as the pixel value of the pixel point.
[0033] According to the first aspect, or any implementation method of the first aspect above, the pixel value of each pixel point in the simulated image is determined based on the pixel value of the corresponding pixel point in M environment maps of each pixel point in the simulated image, including: when M is greater than 1, the pixel value of the corresponding pixel point of the second pixel point in the M environment maps is fused to obtain the pixel value of the second pixel point; wherein the second pixel point is any pixel point in the simulated image.
[0034] That is to say, for any pixel point in the simulated image, when the pixel point has corresponding pixel points in multiple environment maps, the pixel values of the corresponding pixel points in multiple environment maps can be fused to obtain the pixel value of the pixel point.
[0035] Among them, the more pixel points in the environment map used to determine the pixel values of the pixel points of the simulated image, the higher the probability that the pixel points whose pixel values are close to the actual pixel values of the pixel points in the simulated image appear in these multiple pixel points of the environment map. In this way, the pixel values of the pixel points in the determined simulated image are more likely to be more accurate; in this way, the quality of the simulated image can be improved.
[0036] According to the first aspect, or any implementation method of the first aspect above, the pixel values of the corresponding pixel points of the second pixel point in M environment maps are fused to obtain the pixel value of the second pixel point, including: interpolating the pixel values of the corresponding pixel points of the second pixel point in the M environment maps to obtain the pixel value of the second pixel point.
[0037] It should be understood that the present application does not limit the interpolation algorithm used, for example, it can be the nearest neighbor interpolation method.
[0038] According to the first aspect, or any implementation method of the first aspect above, determining the pixel point corresponding to each pixel point in the simulated image of the target distortion camera in M environment maps includes: obtaining the coordinates of each pixel point in the simulated image in the camera coordinate system; determining the pixel point corresponding to each pixel point in the simulated image in the M environment maps based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; wherein the third pixel point is any pixel point in the simulated image, and the first environment map is any environment map among the N environment maps.
[0039] According to the first aspect, or any implementation method of the first aspect above, the pixel point corresponding to each pixel point in the simulated image in the M environment maps is determined according to the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system, including: determining the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map according to the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; when determining that the pixel point corresponding to the third pixel point in the projection plane of the first environment map belongs to the first environment map according to the size of the first environment map and the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map, the pixel point corresponding to the third pixel point in the projection plane of the first environment map is determined as the pixel point corresponding to the third pixel point in the first environment map.
[0040] In this way, it is possible to determine the pixel point corresponding to each pixel point in the simulated image of the target fisheye camera in the M environment maps.
[0041] Exemplarily, each of the N environment maps has a projection plane.
[0042] According to the first aspect, or any implementation of the first aspect above, obtaining the coordinates of each pixel point in the simulated image in the camera coordinate system includes: obtaining a distortion model of the target distortion camera; solving the distortion model using the Gauss-Newton method to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0043] Compared with the prior art that uses a reverse ray tracing solution to determine the coordinates of pixel points in the simulated image under camera coordinates, this application uses the Gauss-Newton method to optimize the distortion model. The time required to determine the coordinates of pixel points in the simulated image under camera coordinates is shorter, and there is no dependency between each pixel point, so parallel calculations can be performed.
[0044] In addition, some existing technologies use a custom inverse distortion projection model to determine the coordinates of pixels in the simulated image in camera coordinates. Since the distortion parameters of the custom distortion camera have certain errors during the custom inverse distortion projection model process, the accuracy of the coordinates of pixels in the simulated image in camera coordinates determined by these existing technologies is not high. However, the present application determines the coordinates of pixels in the simulated image in camera coordinates by optimizing the distortion model of the distortion camera. This eliminates the need for custom distortion parameters of the distortion camera. As can be seen, the present application determines the coordinates of pixels in the simulated image in camera coordinates with greater accuracy.
[0045] According to the first aspect, or any implementation of the first aspect above, a back projection lookup table is obtained; wherein the back projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image and the coordinates in the camera coordinate system; according to the position of each pixel point in the simulated image in the simulated image, the back projection lookup table is searched to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0046] The back-projection lookup table can be pre-established, so that the efficiency of determining the coordinates of each pixel in the simulated image in the camera coordinate system can be improved by looking up the table, thereby improving the efficiency of generating the simulated image of the target distortion camera.
[0047] According to the first aspect, or any implementation of the first aspect above, the target distortion camera includes a target fisheye camera.
[0048] It should be understood that the target distortion camera in the first aspect or any implementation of the first aspect may also include other distortion cameras having similar features to the target fisheye camera, and this application does not limit this.
[0049] According to the first aspect, or any implementation of the first aspect above, any pair of adjacent side surfaces among the N side surfaces of the pyramid are orthogonal.
[0050] According to the first aspect, or any implementation of the first aspect above, at least one pair of adjacent side surfaces among the N side surfaces intersect.
[0051] According to the first aspect, or any implementation of the first aspect, the target device includes a vehicle.
[0052] In a second aspect, an embodiment of the present application provides an image generating device, the device comprising:
[0053] A first scene acquisition module is used to acquire a 3D scene of the target device's driving environment;
[0054] A first projection module is configured to project the 3D scene onto N sides of the pyramid to obtain N environment maps, where N is an integer greater than or equal to 3;
[0055] a first pixel point determination module, configured to determine a corresponding pixel point in M environment maps for each pixel point in the simulated image of the target distorted camera; wherein M is a positive integer less than or equal to N, and the target device includes the target distorted camera;
[0056] The first pixel value determination module is used to determine the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the M environment maps.
[0057] According to the second aspect, the first pixel value determination module is specifically used to determine the pixel value of the corresponding pixel point of the first pixel point in an environment map as the pixel value of the first pixel point when M is equal to 1; wherein the first pixel point is any pixel point in the simulated image.
[0058] According to the second aspect, or any implementation method of the above second aspect, the first pixel value determination module is specifically used to, when M is greater than 1, fuse the pixel values of the corresponding pixel points of the second pixel point in M environment maps to obtain the pixel value of the second pixel point; wherein the second pixel point is any pixel point in the simulated image.
[0059] According to the second aspect, or any implementation method of the second aspect above, the first pixel point determination module is used to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system; according to the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system, determine the pixel point corresponding to each pixel point in the simulated image in the M environment maps; wherein the third pixel point is any pixel point in the simulated image, and the first environment map is any environment map among the N environment maps.
[0060] According to the second aspect, or any implementation method of the above second aspect, the first pixel point determination module is specifically used to determine the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; when it is determined that the pixel point corresponding to the third pixel point in the projection plane of the first environment map belongs to the first environment map based on the size of the first environment map and the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map, the pixel point corresponding to the third pixel point in the projection plane of the first environment map is determined as the pixel point corresponding to the third pixel point in the first environment map.
[0061] According to the second aspect, or any implementation of the second aspect above, the first pixel point determination module is specifically used to obtain a distortion model of the target distortion camera; the Gauss-Newton method is used to solve the distortion model to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0062] According to the second aspect, or any implementation method of the above second aspect, the first pixel point determination module is specifically used to obtain a back projection lookup table; wherein the back projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image and the coordinates in the camera coordinate system; according to the position of each pixel point in the simulated image in the simulated image, the back projection lookup table is searched to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0063] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0064] In a third aspect, an embodiment of the present application provides an image generation method, which includes: first, obtaining a 3D scene of the driving environment of the target device; then, projecting the 3D scene onto a plane to obtain an environment map; and obtaining the coordinates of each pixel point in the simulated image of the target distortion camera in the camera coordinate system, wherein the coordinates of each pixel point in the simulated image in the camera coordinate system are obtained by solving the distortion model of the target distortion camera using the Gauss-Newton method, and the target device includes a target distortion camera; then, according to the coordinates of each pixel point in the simulated image in the camera coordinate system, determining the corresponding pixel point of each pixel point in the simulated image of the target distortion camera in the environment map; thereafter, determining the pixel value of each pixel point in the simulated image according to the pixel value of the corresponding pixel point in the environment map of each pixel point in the simulated image.
[0065] Compared with the prior art that uses a reverse ray tracing solution to determine the coordinates of pixel points in the simulated image under camera coordinates, this application uses the Gauss-Newton method to optimize the distortion model. The time required to determine the coordinates of pixel points in the simulated image under camera coordinates is shorter, and there is no dependency between each pixel point, so parallel calculations can be performed.
[0066] In addition, some existing technologies use a custom inverse distortion projection model to determine the coordinates of pixels in the simulated image in camera coordinates. Since the distortion parameters of the custom distortion camera have certain errors during the custom inverse distortion projection model process, the accuracy of the coordinates of pixels in the simulated image in camera coordinates determined by these existing technologies is not high. However, the present application determines the coordinates of pixels in the simulated image in camera coordinates by optimizing the distortion model of the distortion camera. This eliminates the need for custom distortion parameters of the distortion camera. As can be seen, the present application determines the coordinates of pixels in the simulated image in camera coordinates with greater accuracy.
[0067] Exemplarily, the environment map of the third aspect may also be referred to as a planar map.
[0068] For example, the pose of the camera can be determined based on the pose of the distorted pinhole camera on the target device; and the size of the environment map (height is H, width is W) can be determined; then, the 3D scene is rendered according to the pose of the camera and the size of the environment map to obtain the environment map.
[0069] Illustratively, the field of view of the target distortion camera of the third aspect is less than 180°.
[0070] According to the third aspect, the pixel value of each pixel in the simulated image is determined based on the pixel value of the corresponding pixel in the environment map of each pixel in the simulated image, including: determining the pixel value of the corresponding pixel in the environment map of each pixel in the simulated image as the pixel value of each pixel in the simulated image.
[0071] According to the third aspect, or any implementation of the third aspect, the target distortion camera is a target distortion pinhole camera.
[0072] It should be understood that the target distortion camera in the third aspect or any implementation of the third aspect may also include other distortion cameras having similar features to the distortion pinhole camera, and this application does not limit this.
[0073] In a fourth aspect, an embodiment of the present application provides an image generating device, the image generating device comprising:
[0074] The second scene acquisition module is used to acquire the 3D scene of the target device's driving environment;
[0075] The second projection module is used to project the 3D scene onto a plane to obtain an environment map;
[0076] a coordinate acquisition module, configured to acquire the coordinates of each pixel in a simulated image of a target distortion camera in a camera coordinate system, wherein the coordinates of each pixel in the simulated image in the camera coordinate system are obtained by solving a distortion model of the target distortion camera using a Gauss-Newton method, wherein the target device includes the target distortion camera;
[0077] A second pixel point determination module is used to determine the corresponding pixel point in the environment map for each pixel point in the simulated image of the target distortion camera according to the coordinates of each pixel point in the simulated image in the camera coordinate system;
[0078] The second pixel value determination module is used to determine the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the environment map of each pixel in the simulated image.
[0079] The fourth aspect and any implementation of the fourth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0080] In a fifth aspect, an embodiment of the present application provides an electronic device comprising: a memory and a processor, wherein the memory is coupled to the processor; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method in the first aspect or any possible implementation of the first aspect.
[0081] Exemplarily, the electronic device may be a server or a terminal device, and this application does not impose any restrictions on this.
[0082] The fifth aspect and any implementation of the fifth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0083] In the sixth aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, enables the electronic device to execute the method in the third aspect or any possible implementation of the third aspect.
[0084] Exemplarily, the electronic device may be a server or a terminal device, and this application does not impose any restrictions on this.
[0085] The sixth aspect and any implementation of the sixth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the sixth aspect and any implementation of the sixth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0086] In the seventh aspect, an embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method in the first aspect or any possible implementation of the first aspect are executed.
[0087] The seventh aspect and any implementation of the seventh aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the seventh aspect and any implementation of the seventh aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0088] In an eighth aspect, an embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method in the third aspect or any possible implementation of the third aspect are executed.
[0089] The eighth aspect and any implementation of the eighth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the eighth aspect and any implementation of the eighth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0090] In the ninth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, it enables the computer or the processor to execute the method in the first aspect or any possible implementation of the first aspect.
[0091] The ninth aspect and any implementation of the ninth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the ninth aspect and any implementation of the ninth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect, and will not be repeated here.
[0092] In the tenth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, it enables the computer or the processor to execute the method in the third aspect or any possible implementation of the third aspect.
[0093] The tenth aspect and any implementation of the tenth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the tenth aspect and any implementation of the tenth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here.
[0094] In the eleventh aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer or the processor executes the method in the first aspect or any possible implementation of the first aspect.
[0095] The eleventh aspect and any implementation of the eleventh aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the eleventh aspect and any implementation of the eleventh aspect can be referred to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, and will not be repeated here.
[0096] In the twelfth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer or the processor executes the method in the third aspect or any possible implementation of the third aspect.
[0097] The twelfth aspect and any implementation of the twelfth aspect correspond to the third aspect and any implementation of the third aspect, respectively. The technical effects corresponding to the twelfth aspect and any implementation of the twelfth aspect can be referred to the technical effects corresponding to the third aspect and any implementation of the third aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] FIG1A is a schematic diagram illustrating an exemplary application scenario;
[0099] FIG1B is a schematic diagram illustrating an exemplary application scenario;
[0100] FIG2A is a schematic diagram illustrating an exemplary image generation process 200;
[0101] FIG2B is a schematic diagram of an exemplary pyramid;
[0102] FIG2C is a schematic diagram illustrating an exemplary image generation process;
[0103] FIG3A is a schematic diagram illustrating an exemplary image generation process 300;
[0104] FIG3B is a schematic diagram showing the geometric relationship of an orthogonal triangular prism;
[0105] FIG3C is a schematic diagram illustrating an exemplary imaging principle of a fisheye camera;
[0106] FIG4 is a schematic diagram illustrating an exemplary image generation process 400;
[0107] FIG5 is a schematic diagram of an exemplary image generating apparatus 500;
[0108] FIG6 is a schematic diagram of an exemplary image generating apparatus 600;
[0109] FIG7 is a schematic structural diagram of an exemplary device. DETAILED DESCRIPTION
[0110] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0111] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0112] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first target object" and "second target object" are used to distinguish different objects, rather than to describe a specific order of objects.
[0113] In the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0114] In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more. For example, "multiple processing units" means two or more processing units; "multiple systems" means two or more systems.
[0115] Figure 1A is a schematic diagram of an exemplary application scenario. Figure 1A shows a training scenario 100 of a perception model.
[0116] 1A , illustratively, in a perception model training scenario 100, a simulated image 102 from a distorted camera deployed on a target device can be generated based on a 3D scene 101 of the target device's driving environment. Furthermore, real data 103 (e.g., location information and 3D models of objects (such as obstacles, pedestrians, and other vehicles) contained in 3D scene 101) can be generated based on the 3D scene 101 of the target device's driving environment. Subsequently, a perception model 104 can be trained using the simulated image 102 from the distorted camera and the real data 103.
[0117] It should be noted that there can be multiple 3D scenes 101; the target device can be set at multiple locations within the current 3D scene 101. When the target device is at its current location within the current 3D scene, simulated images 102 are generated for the multiple distortion cameras deployed on the target device. In other words, multiple simulated images 102 can be generated when the target device is at its current location within the current 3D scene. The following describes the process of training the perception model using a single simulated image 102 as an example.
[0118] For example, a simulated image 102 can be input into the perception model 104, which then performs forward computation to determine the predicted distance between each object in the simulated image 102 and the target device. Subsequently, a reference distance between each object in the simulated image 102 and the target device can be determined based on the position information of the objects in the 3D scene 101, the 3D model, and the position information of the target device in the 3D scene. Next, with the goal of minimizing the reference distance and predicted distance between each object in the simulated image 102 and the target device, backpropagation is performed on the perception model 104 to adjust its model parameters.
[0119] After the training of the perception model 104 is completed, the trained perception model 104 can be installed on the target device. In this way, while the target device is driving, the trained perception model 104 can be used to identify the distance between obstacles, pedestrians, other vehicles, animals, etc. and the vehicle based on the distorted images captured by the distortion camera deployed on the target device. The target device can then control its position, motion status, etc. based on this distance, or prompt the user.
[0120] For example, the target device may include a vehicle, a delivery robot, a food delivery robot, an airport guide robot, a shopping mall service robot, etc., and this application does not limit this. This application takes the target device as an example of a vehicle for explanation.
[0121] Figure 1B is a schematic diagram of an exemplary application scenario. Figure 1B shows an autonomous driving algorithm test scenario 110.
[0122] 1B , illustratively, in an autonomous driving algorithm test scenario 110, simulated images 112 from multiple (e.g., six) distorted cameras mounted on the vehicle (i.e., simulated images from multi-view distorted cameras) can be generated based on a 3D scene 111 of the vehicle's driving environment. Furthermore, real data 113 (e.g., location information and 3D models of objects (e.g., obstacles, pedestrians, other vehicles, etc.) contained in 3D scene 111) can be generated based on the 3D scene 111 of the vehicle's driving environment. Subsequently, the autonomous driving algorithm 114 can be tested using the simulated images 112 from the multi-view distorted cameras, the real data 113, and the trained perception model 104 of FIG. 1A .
[0123] It should be noted that there can be multiple 3D scenes 111; the vehicle can be set at multiple positions in the current 3D scene 111, and when the vehicle is at the current position of the current 3D scene and the vehicle is in the current posture, a simulation image 112 of the multi-view distortion camera is generated.
[0124] For example, a multi-view distorted camera simulated image 112 generated when the vehicle is at its current position and posture in the current 3D scene is input into the trained perception model 104. The trained perception model 104 processes the simulated image 112 and outputs the predicted distance between each object and the vehicle. The predicted distance between each object and the vehicle in the simulated image 112 is then input into the autonomous driving algorithm. The autonomous driving algorithm can determine predictive control information (e.g., steering wheel angle, vehicle speed, gear position, etc.) based on the predicted distance between each object and the vehicle in the simulated image 112, the vehicle's current position, and the vehicle's current posture. The algorithm then calculates the vehicle's position, posture, and motion state after controlling the vehicle based on the predictive control information (e.g., adjusting the steering wheel angle, vehicle speed, gear position, etc.). The algorithm then determines whether the vehicle's position, posture, and motion state after controlling the vehicle meet safety standards, thereby testing whether the autonomous driving algorithm meets safety requirements. If these safety requirements are not met, the autonomous driving algorithm can be adjusted.
[0125] It should be noted that the above operations of generating simulated images 102 (112), generating real data 103 (113), training the perception model 104, and testing the autonomous driving algorithm 114 can be performed by a server or by a terminal device, and this application does not impose any restrictions on this.
[0126] It should be noted that the function of generating the simulated image 102 (112) can be implemented by software (e.g., image simulation software) or by hardware, and this application does not impose any restrictions on this. The image generation method involved in this application can be used to generate the simulated image 102 (112). In some scenarios, the image generation method involved in this application can also be called an image simulation method.
[0127] The image generation method involved in this application is described below.
[0128] FIG. 2A is a schematic diagram illustrating an exemplary image generation process 200 .
[0129] S201: Acquire a 3D scene of a target device's driving environment.
[0130] Exemplarily, a 3D scene may refer to a virtual three-dimensional environment generated by an electronic device.
[0131] For example, when the target device is a vehicle, the 3D scene of the vehicle driving environment may include: a 3D scene of a parking lot, a 3D scene of a city road, a 3D scene of a suburban road, a 3D scene of a mountain road, a 3D scene of a highway, etc., and this application does not impose any restrictions on this.
[0132] Exemplarily, when the target device is a food delivery robot, the 3D scene of the food delivery robot's driving environment may include: a 3D scene of an apartment, a 3D scene of an office building, etc.
[0133] Exemplarily, when the target device is an airport guidance robot, the 3D scene of the airport guidance robot's driving environment may include: a 3D scene of an airport waiting room, etc.
[0134] Exemplarily, when the target device is a shopping mall service robot, the 3D scene of the shopping mall service robot's driving environment may include: a 3D scene of the shopping mall, etc.
[0135] S202 , projecting the 3D scene onto N side faces of a pyramid to obtain N environment maps; wherein N is an integer greater than or equal to 3.
[0136] It should be noted that since the environment map generated based on the 3D scene must be rectangular, this application expands each side of the pyramid into a rectangle. The number of sides of the pyramid can be represented by N, and the pyramid can include N pairs of adjacent sides, and the angle between each pair of adjacent sides can be represented by α.
[0137] In the embodiment of the present application, N may be an integer greater than or equal to 3; and the present application does not limit the size of the side angle α between each pair of adjacent side surfaces.
[0138] For example, if N=3, the angle α between side 1 and side 2 is 90°, the angle α between side 2 and side 3 is 90°, and the angle α between side 3 and side 1 is 90°. The triangular pyramid in this case can also be called an orthogonal triangular prism. For another example, if N=3, the angle α between side 1 and side 2 is 60°, the angle α between side 2 and side 3 is 60°, and the angle α between side 3 and side 1 is 60°. For another example, if N=4, the angle α between side 1 and side 2 is 90°, the angle α between side 2 and side 3 is 60°, the angle α between side 3 and side 4 is 60°, and the angle α between side 4 and side 1 is 60°. And so on. The examples are not listed here one by one.
[0139] Exemplarily, at least one pair of adjacent side surfaces of the pyramid are orthogonal; that is, the side angle α between at least one pair of adjacent side surfaces of the pyramid is 90°.
[0140] Exemplarily, edges of at least one pair of adjacent side surfaces among the N side surfaces of the pyramid intersect.
[0141] Exemplarily, at least one pair of adjacent side surfaces among the N side surfaces of the pyramid intersect.
[0142] Exemplarily, the vertex of the pyramid is on the optical axis, and the N side faces of the pyramid wrap around the imaging center point (i.e., the camera center point of the 3D scene). Exemplarily, the 3D scene can be projected onto the N side faces of the pyramid by rendering the 3D scene N times, thereby obtaining N environment maps. Specifically, the 3D scene can be projected onto one side face of the pyramid by rendering the 3D scene once, thereby obtaining an environment map. In this way, each of the N side faces of the pyramid is an environment map. It should be understood that the N environment maps correspond one-to-one to the N side faces of the pyramid.
[0143] Specifically, the size of side n (n less than or equal to N) can be determined (height represented by H (H is a positive integer) and width represented by W (W is a positive integer)). Next, the camera is positioned at the camera center and facing side n of the pyramid. This determines the camera pose. Then, based on the camera pose and the size of side n, the 3D scene can be rendered to project the 3D scene onto side n of the pyramid, generating an environment map n. It should be understood that the environment map n and side n have the same size.
[0144] For example, the size of the side n can be determined according to the field of view of the target distortion camera; details will be described later.
[0145] Fig. 2B is a schematic diagram of an exemplary pyramid, wherein the N side angles α of N pairs of adjacent side surfaces of the pyramid are the same, and the height and width of the N side surfaces are the same.
[0146] FIG2B(1) shows a triangular pyramid in which the angle α between the three sides is 90°, and the height and width of the three sides are both 1. In this case, each pair of adjacent sides of the three pairs of adjacent sides of the triangular pyramid are orthogonal, and the edges of each pair of adjacent sides of the three pairs of adjacent sides intersect.
[0147] FIG2B(2) shows a triangular pyramid in which the angle α between the three sides is 90°, and the height and width of the three sides are both 1.5. In this case, each pair of adjacent sides of the triangular pyramid is orthogonal, and one pair of adjacent sides of the three pairs intersects.
[0148] FIG2B(3) shows a triangular pyramid, wherein the angle α among the three side faces is 60°, and the height and width of the three side faces are both 1. In this case, each pair of adjacent side faces of the three adjacent pairs of the triangular pyramid intersects.
[0149] FIG2B(4) shows a triangular pyramid in which the angle α between the three sides is 90°, and the height and width of the three sides are both 1. In this case, each pair of adjacent sides of the three pairs of the triangular pyramid are orthogonal, and the edges of each pair of adjacent sides intersect.
[0150] FIG2B(5) shows a quadrangular pyramid, wherein the angle α between the four sides is 90°, and the sum of the widths of the four sides is 1. In this case, each pair of adjacent sides of the four pairs of adjacent sides of the quadrangular pyramid is orthogonal, and each pair of adjacent sides of the four pairs of adjacent sides intersects.
[0151] FIG2B(6) shows an octagonal pyramid, wherein the angle α among the eight sides is 90°, and the sum of the widths of the eight sides is 1. In this case, each pair of adjacent sides among the eight pairs of adjacent sides of the octagonal pyramid is orthogonal, and each pair of adjacent sides among the eight pairs of adjacent sides intersects.
[0152] It should be noted that the sphere in the middle of each pyramid in FIG2B may be a schematic representation of a 3D scene.
[0153] S203: Determine the pixel point corresponding to each pixel point in the simulated image of the target distorted camera in the M environment maps; where M is a positive integer less than or equal to N, and the target distorted camera is any one of the multiple distorted cameras included in the target device.
[0154] For example, one of the multiple distortion cameras included in the target device can be determined as the target distortion camera. Next, the size of the distorted image obtained by shooting with the target distortion camera (that is, the resolution of the distortion camera) is determined; wherein, the height of the distorted image can be represented by h (h is a positive integer), and the width of the distorted image can be represented by w (w is a positive integer) (that is, the distorted image includes h*w pixels). Afterwards, a simulated image (for example, a blank simulated image) can be generated with a height of h, a width of w, and pixel values of h*w pixels all being preset values. In other words, the size of the simulated image is the same as the size of the distorted image.
[0155] For example, the target distortion camera may be a target fisheye camera. It should be understood that the target distortion camera may also include other distortion cameras having similar features to the target fisheye camera, and this application does not limit this.
[0156] For example, after obtaining N environment maps, the corresponding pixel in each of the N environment maps can be determined for each pixel in the simulated image. It should be noted that each pixel in the simulated image may only have a corresponding pixel in M of the N environment maps, where M is a positive integer less than or equal to N.
[0157] For example, when N is equal to 3, as in the case of FIG2B(1) and FIG2B(4), since the edges of each pair of adjacent side surfaces of the three pairs of adjacent side surfaces of the triangular pyramid intersect, each pixel point in the simulated image has a corresponding pixel point in only one environment map. In the case of FIG2B(2) and FIG2B(3), since at least one pair of adjacent side surfaces of the three pairs of adjacent side surfaces of the triangular pyramid intersects, some pixels points in the simulated image have corresponding pixel points in both environment maps.
[0158] Exemplarily, when N is greater than or equal to 3, since each pair of adjacent side surfaces of the N pairs of adjacent side surfaces of the pyramid intersects, some pixels in the simulation image have corresponding pixels in multiple environment maps.
[0159] S204 , determining the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the M environment maps.
[0160] For example, for each pixel in the simulated image, the pixel value of the corresponding pixel in the M environment maps can be determined. Then, a target value can be determined based on the pixel value of the corresponding pixel in the M environment maps. After that, the pixel values of all pixels in the simulated image can be updated from the preset values to the target values. In this way, after the pixel values of all pixels in the simulated image are updated from the preset values to the target values, the final simulated image of the target distorted camera can be obtained; that is, the distorted image of the target distorted camera obtained through simulation.
[0161] FIG2C is a schematic diagram illustrating an exemplary image generation process.
[0162] FIG2C(1) shows a schematic diagram of the image generation process of the prior art. In FIG2C(1), the 3D scene is projected onto the five sides of a cube to obtain five environment maps. Among them, it can be determined that the pixel point Px_1 in the simulated image of the fisheye camera has a corresponding pixel point Pt_1 in one environment map; the pixel value of the pixel point Pt_1 can be used as the pixel value of the pixel point Px_1 in the simulated image.
[0163] FIG2C(2) shows a schematic diagram of the image generation process of the present application. In FIG2C(2), the 3D scene is projected onto the three sides of a triangular pyramid to obtain three environment maps. Among them, it can be determined that the pixel point Px_2 in the simulated image of the fisheye camera has a corresponding pixel point Pt_2 in one environment map; the pixel value of the pixel point Pt_2 can be used as the pixel value of the pixel point Px_2 in the simulated image.
[0164] FIG2C(3) shows a schematic diagram of the image generation process of the present application. In FIG2C(3), the 3D scene is projected onto the three sides of a triangular pyramid to obtain three environment maps. Among them, it can be determined that the pixel point Px_3 in the simulated image of the fisheye camera has corresponding pixel points Pt_3 and Pt_4 in the two environment maps; the pixel value of the pixel point Px_3 in the simulated image can be determined based on the pixel value of the pixel point Pt_3 and the pixel value of Pt_4.
[0165] When N is equal to 3, the present application projects the 3D scene of the target device's driving environment onto the three sides of a pyramid to obtain three environment maps. These three environment maps can support the generation of simulated images of a distorted camera with a maximum field of view of about 240°; and the maximum actual field of view of a distorted camera is usually about 220°. That is to say, the present application projects the 3D scene onto the three sides of a triangular pyramid to obtain three environment maps, which can support the generation of simulated images of all (or most) distorted cameras. The prior art, on the other hand, projects the 3D scene of the target device's driving environment onto five or six faces of a cube to obtain five or six environment maps; these five or six environment maps can support the generation of simulated images of a distorted camera with a maximum field of view of about 270°. It can be seen that the prior art generates five or six environment maps with redundancy in field of view.
[0166] When N is equal to 3, and each pair of adjacent side surfaces of the three pairs of triangular pyramids are orthogonal, each pixel in the simulated image of the target distortion camera in the present application has a corresponding pixel in one environment map; each pixel in the simulated image of the target distortion camera in the prior art also has a corresponding pixel in one environment map; that is, the pixel value of a pixel in the simulated image generated by the present application and the prior art is determined based on the pixel value of one pixel; it can be seen that the quality of the simulated image generated by the present application is the same as the quality of the simulated image generated by the prior art. In summary, when N is equal to 3, and each pair of adjacent side surfaces of the three pairs of triangular pyramids are orthogonal, the quality of the simulated image generated by the present application is the same as the quality of the simulated image generated by the prior art, but the number of environment maps required to be generated by the present application is less, so the efficiency of generating simulated images by the present application is higher.
[0167] When N is equal to 3, and at least one of the three pairs of adjacent side surfaces of the triangular pyramid intersects, some of the pixels in the simulated image of the target distortion camera in the present application have corresponding pixels in one environment map, and some of the other pixels have corresponding pixels in two environment maps. In other words, the pixel values of some of the pixels in the simulated image of the present application are determined based on the pixel values of two pixels; wherein, the more pixels in the environment map used to determine the pixel values of the pixels of the simulated image, the higher the probability that the pixel values of the pixels in the multiple pixels of the environment map are close to the true pixel values of the pixels in the simulated image, so that the pixel values of the pixels in the determined simulated image are more likely to be more accurate; it can be seen that the quality of the simulated image generated by the present application is higher than the quality of the simulated image generated by the prior art. In summary, when N is equal to 3, and at least one of the three pairs of adjacent side surfaces of the triangular pyramid intersects, the quality of the simulated image generated by the present application is not only higher than the quality of the simulated image generated by the prior art, but also the efficiency of generating simulated images by the present application is higher.
[0168] When N is equal to 4, the present application projects the 3D scene of the target device's driving environment onto the four sides of a pyramid to obtain four environment maps. These four environment maps can support the generation of simulated images of a distortion camera with a maximum field of view angle slightly greater than 240°; and the maximum actual field of view angle of a distortion camera is usually around 220°, that is, the four environment maps obtained by projecting the 3D scene onto the four sides of a quadrangular pyramid in the present application can support the generation of simulated images of all (or most) distortion cameras. In addition, in the present application, some pixels in the simulated image of the target distortion camera have corresponding pixels in one environment map, and some other pixels have corresponding pixels in two environment maps; that is, the pixel values of some pixels in the simulated image of the present application are determined based on the pixel values of two pixels; it can be seen that the quality of the simulated image generated by the present application is higher than that of the simulated image generated by the prior art.
[0169] In summary, when N is equal to 4, the quality of the simulated image generated by the present application is not only higher than that of the simulated image generated by the prior art, but the efficiency of generating the simulated image by the present application is also higher.
[0170] When N is equal to 5 or 6, the present application projects the 3D scene of the target device's driving environment onto 5 or 6 faces of a pyramid to obtain 5 or 6 environment maps; these 5 or 6 environment maps can support the generation of simulated images of a distortion camera with a maximum field of view of about 260°; and the actual maximum field of view of a distortion camera is usually about 220°, that is, the 5 or 6 environment maps obtained by projecting the 3D scene onto 5 or 6 faces of a pyramid in the present application can support the generation of simulated images of all (or most) distortion cameras. In addition, in the present application, some pixels in the simulated image of the target distortion camera have corresponding pixels in one environment map, and some other pixels have corresponding pixels in multiple environment maps; that is, the pixel values of some pixels in the simulated image of the present application are determined based on the pixel values of multiple pixels; it can be seen that the quality of the simulated image generated by the present application is higher than that of the simulated image generated by the prior art.
[0171] In summary, when N is equal to 5 or 6, the efficiency of the simulated images generated by the present application is the same as that of the prior art, but the quality of the simulated images generated by the present application is higher.
[0172] When N is greater than 6, although this application needs to generate more environment maps and requires higher computing power, the quality of the generated simulation image is also higher, which is something that cannot be achieved by the existing technology of projecting the 3D scene of the target device's driving environment onto 5 or 6 faces of a cube.
[0173] The following describes the process of determining the pixel point corresponding to each pixel point in the M environment maps in the simulated image of the target distorted camera, and the process of determining the pixel value of each pixel point in the simulated image.
[0174] 3A is a schematic diagram illustrating an exemplary image generation process 300. In the image generation process 300, the target distortion camera is a target fisheye camera.
[0175] S301: Acquire a 3D scene of a target device's driving environment.
[0176] For example, S301 may refer to the description of S201 and will not be repeated here.
[0177] S302 : Project the 3D scene onto N sides of the pyramid to obtain N environment maps.
[0178] The following uses the target distortion camera as a target fisheye camera and N equals 3 as an example to illustrate the process of determining the size of each side.
[0179] Figure 3B is a schematic diagram illustrating the geometric relationship of an orthogonal triangular prism. The coordinate system in Figure 3B is the camera coordinate system (i.e., the polar coordinate system), which includes the camera center O, the x-axis, the y-axis, and the z-axis; the x-axis, the y-axis, and the z-axis are orthogonal to each other.
[0180] 3B , exemplarily, plane ABEC, plane ABFD and plane ACGD are three side faces of an orthogonal triangular prism, B′ is a point symmetrical about B on the z-axis (B′ is not on plane ACGD), and H0 is the projection point of point O on line segment AB.
[0181] For example, the distortion parameters of the target fisheye camera include {K1, K2, K3, K4}, and the intrinsic parameters of the target fisheye camera are {C x , C y , f x , f y}, and the resolution is h*w. Where ∠B′OB is the maximum FOV of the target fisheye camera, denoted as θ. The relationship between θ and the intrinsic parameters and resolution of the target fisheye camera can be referred to the following formula (1):
[0182] For example, according to the geometric relationship, the following formulas (2) and (3) can be obtained:
[0183] For example, according to formula (2) and formula (3), the following formula (4) is obtained:
[0184] For example, according to formula (4), the following formula (5) is derived:
[0185] Referring to formula (1) and formula (5), it can be seen that the sizes of the N side surfaces of the pyramid can be determined according to the field of view of the target fisheye camera.
[0186] It should be understood that when the number N of sides of the pyramid is other values, or the side angle is other values, the mathematical relationship between the size of the side and the field of view angle of the target fisheye camera can also be derived based on the geometric relationship between the N sides of the pyramid, which will not be repeated here.
[0187] Afterwards, N environment maps can be obtained by rendering the 3D scene N times and projecting the 3D scene onto N sides of the pyramid.
[0188] Illustratively, S203 may include the following descriptions of S303 to S304.
[0189] S303: Obtain the coordinates of each pixel in the simulation image in the camera coordinate system.
[0190] For example, the coordinates of each pixel point in the simulated image of the target fisheye camera in the distorted image coordinate system can be converted to the camera coordinate system, thereby obtaining the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0191] FIG3C is a schematic diagram illustrating an exemplary imaging principle of a fisheye camera.
[0192] 3C , exemplarily, the angle between the incident direction of light and the central axis of the optical center (the dotted line in FIG3C ) is θ, and the light refracted through the lens of the fisheye camera corresponds to the pixel point (u', v') on the distorted image (which can also be regarded as the pixel point (u', v') in the simulated image, where (u', v') is used to represent the position of the pixel point in the simulated image, assuming that the pixel point in the upper left corner of the simulated image (i.e., the 1st row and 1st column) is represented by the pixel point (0, 0), and the pixel point (u', v') can refer to the pixel point in the u'-1th row and v'-1th column in the simulated image), and the coordinates of the pixel point (u', v') in the distorted coordinate system are In this way, the distance from the pixel point (u', v') to the center axis of the optical center can be obtained.
[0193] Exemplarily, S303 may include the following: S3031 to S3033:
[0194] S3031, obtain the distortion model of the target fisheye camera.
[0195] For example, the distortion model of the fisheye camera (Kannala-Brandt model) is shown in the following formula (6): θ d =θ+k1θ 3 +k2θ 5 +k3θ 7 +k4θ 9 (6)
[0196] For example, according to formula (6), the following equation can be obtained, as shown in formula (7): θ+k1θ 3 +k2θ 5 +k3θ 7 +k4θ 9 -θ d =0 (7)
[0197] S3032: Solve the distortion model using the Gauss-Newton method to obtain the first component of the coordinate of each pixel in the camera coordinate system.
[0198] For example, the optimization algorithm can be used to solve formula (7) to obtain the θ of the pixel point (u', v') d The corresponding θ value.
[0199] For example, the Gauss-Newton method (also known as the Newton method iteration method) may be used to solve Formula (7).
[0200] Exemplarily, solving formula (7) can be regarded as finding the root of the function f(θ) to be found on the left side of the equal sign in formula (7).
[0201] For example, the formula of the Gauss-Newton method can be shown as formula (8):
[0202] The derivative f'(θ) of f(θ) can be expressed as follows:
[0203] In formula (8), n is the number of iterations, θ n is the value of θ at the nth iteration, f(θ n ) is the value of f(θ) under the value of θ at the nth iteration, f'(θ n ) is f(θ) at the value of θ at the nth iteration.
[0204] For example, the θ of pixel (u', v') d The corresponding θ value is the first component of the coordinate of the pixel point (u', v') in the camera coordinate system.
[0205] Compared with the prior art that uses a reverse ray tracing solution to determine the coordinates of pixel points in the simulated image under camera coordinates, this application uses the Gauss-Newton method to optimize the distortion model. The time required to determine the coordinates of pixel points in the simulated image under camera coordinates is shorter, and there is no dependency between each pixel point, so parallel calculations can be performed.
[0206] In addition, some existing technologies use a custom inverse distortion projection model to determine the coordinates of pixels in the simulated image in camera coordinates. Since the distortion parameters of the custom distortion camera have certain errors during the custom inverse distortion projection model process, the accuracy of the coordinates of pixels in the simulated image in camera coordinates determined by these existing technologies is not high. However, the present application determines the coordinates of pixels in the simulated image in camera coordinates by optimizing the distortion model of the distortion camera. This eliminates the need for custom distortion parameters of the distortion camera. As can be seen, the present application determines the coordinates of pixels in the simulated image in camera coordinates with greater accuracy.
[0207] S3033 : Determine a second component of the coordinates of each pixel point in the camera coordinate system according to the coordinates of each pixel point in the simulated image in the distorted image coordinate system.
[0208] For example, the second component of the coordinates of the pixel point (u', v') in the camera coordinate system can be determined based on xd and yd: It can be expressed as the following formula (10):
[0209] That is to say, the coordinates of the pixel point (u', v') in the camera coordinate system are
[0210] It should be understood that the above method can also be used to obtain the coordinates of other pixel points in the camera coordinate system.
[0211] It should be noted that another implementation of S303 may be as shown in S3034 to S3035 below:
[0212] S3034, obtaining a back-projection lookup table; wherein the back-projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image and the coordinates in the camera coordinate system.
[0213] Exemplarily, the back-projection lookup table may be established in the manner of S3031 to S3033 described above, which will not be described in detail here.
[0214] For example, the back-projection lookup table is represented by LUT.
[0215] S3035 , searching the inverse projection lookup table according to the position of each pixel point in the simulated image in the simulated image, and obtaining the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0216] S304 , determining the pixel point of each pixel in the simulated image in the M environment maps according to the projection matrices corresponding to the N projection planes of the N environment maps and the coordinates of each pixel in the simulated image of the target fisheye camera in the camera coordinate system.
[0217] For example, the process of determining the corresponding pixel point in any one of the N environment maps for any pixel point in the simulated image can be as follows S3041-S3042. For ease of explanation, any pixel point in the simulated image can be referred to as a third pixel point, and any one of the N environment maps can be referred to as a first environment map.
[0218] S3041, determining the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map according to the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system.
[0219] For example, the coordinates of the third pixel point (u', v') in the camera coordinate system can be first Convert to the Cartesian coordinate system and get the coordinates of the third pixel point (u', v') in the Cartesian coordinate system
[0220] Next, determine the projection matrix corresponding to the projection plane of the first environment map. The projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point (u', v') in the Cartesian coordinate system are used. Get the position of the third pixel point (u', v') corresponding to the pixel point in the projection plane of the first environment map. Assume that the position of the third pixel point (u', v') corresponding to the pixel point in the projection plane of the first environment map is the u-th pixel point in the projection plane of the first environment map. tex -1 row v tex -1 column, then the pixel corresponding to the third pixel point (u', v') in the projection plane of the first environment map can be (u tex ,v tex ) is expressed as shown in formula (11):
[0221] Among them, p camera is the coordinate of the third pixel (u', v') in the Cartesian coordinate system is the projection matrix of the projection plane of the i-th environment map (i.e. the third environment map).
[0222] S3042, when determining that the pixel point corresponding to the third pixel point in the projection plane of the first environment map belongs to the first environment map based on the size of the first environment map and the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map, the pixel point corresponding to the third pixel point in the projection plane of the first environment map is determined as the pixel point corresponding to the third pixel point in the first environment map.
[0223] For example, the first environment map may include H rows*W columns of pixels, so it can be determined that u tex Whether it belongs to [0,H-1], and v tex Whether it belongs to [0,W-1]. tex belongs to [0,H-1], and v tex When it belongs to [0, W-1], it can be determined that the third pixel point has a corresponding pixel point in the first environment map, and the pixel point is the uth pixel point in the first environment map. tex -1 row v tex Otherwise, it is determined that the third pixel does not have a corresponding pixel in the first environment map. In this case, another environment map in the N environment maps can be determined as the first environment map, and S3041 to S3042 are executed again.
[0224] S305 , determining the pixel value of each pixel in the simulated image of the target fisheye camera according to the pixel value of the corresponding pixel in the M environment maps.
[0225] In one possible approach, when M is 1, the pixel value of a corresponding pixel point in an environment map for a first pixel is determined as the pixel value of the first pixel point; wherein the first pixel point is any pixel point in the simulated image. In other words, for any pixel point in the simulated image, if the pixel point has a corresponding pixel point in an environment map, the pixel value of the corresponding pixel point in the environment map can be determined as the pixel value of the pixel point.
[0226] In one possible approach, when M is greater than 1, the pixel values of the second pixel corresponding to the M environment maps are fused to obtain the pixel value of the second pixel; wherein the second pixel is any pixel in the simulated image. In other words, for any pixel in the simulated image, if the pixel has a corresponding pixel in multiple environment maps, the pixel values of the corresponding pixels in the multiple environment maps can be fused to obtain the pixel value of the pixel.
[0227] For example, the pixel value of the second pixel can be obtained by interpolating the pixel values of the corresponding pixel points of the second pixel in the M environment maps. For example, the pixel values of the corresponding pixel points of the second pixel in the M environment maps can be interpolated using the nearest neighbor interpolation method; this application does not limit the interpolation method used.
[0228] 4 is a schematic diagram illustrating an exemplary image generation process 400. In the image generation process 400, the target distortion camera is a target distortion pinhole camera.
[0229] S401: Acquire a 3D scene of a target device's driving environment.
[0230] For example, S401 may refer to the description of S201 above, which will not be repeated here.
[0231] S402, projecting the 3D scene onto a plane to obtain an environment map.
[0232] For example, since the field of view of the distorted pinhole camera is less than 180°, an environment map (also called a planar map) can be generated based on the 3D scene. For example, the camera's pose can be determined based on the pose of the distorted pinhole camera on the target device; and the size of the environment map (height is H, width is W) can be determined. Then, the 3D scene is rendered based on the camera's pose and the size of the environment map to obtain the environment map. It should be noted that S402 performs a rendering to obtain an environment map.
[0233] For example, the process of determining the size of the environment map may be as follows:
[0234] For example, the distortion parameters of the target distortion pinhole camera include {K1, K2, P1, P2, K3}, and the intrinsic parameters of the target distortion pinhole camera are {C x , C y , f x , f y}, the resolution is h*w.
[0235] For example, the angle between the incident direction of light and the axis of the optical center is θ, and the pixel point corresponding to the distorted image after the light is refracted by the lens of the distorted pinhole camera is (u', v') (which can also be regarded as the pixel point (u', v') in the simulated image). The coordinates of the pixel point (u', v') in the distorted coordinate system are
[0236] For example, the coordinates (x, y) of the pixel point (u, v) in the environment map corresponding to the pixel point (u', v') in the simulation image = The distance from the pixel point (u,v) to the center axis of the optical center
[0237] For example, the height H and width W of the environment map can be obtained by solving the following formula (12).
[0238] Among them, formula (12) is the distortion model (Brown-Conrady model) of the distorted pinhole camera.
[0239] Exemplarily, the present application may use a binary search method to solve the inequality to obtain the height H and width W of the environment map.
[0240] Based on the above description, we can see that for coordinates satisfying The coordinates of all pixel points projected onto the projection plane of the environment map satisfy In other words, all pixels in the simulated image can be projected into the environment map.
[0241] S403 , obtaining the coordinates of each pixel point in the simulated image of the target distorted pinhole camera in the camera coordinate system.
[0242] For example, the coordinates of each pixel point in the simulated image of the target distorted pinhole camera in the distorted image coordinate system can be converted to the camera coordinate system, thereby obtaining the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0243] Exemplarily, S403 may include the following: S4031 to S4033:
[0244] S4031, obtaining a distortion model of a target distorted pinhole camera.
[0245] For example, according to the distortion model of the distorted pinhole camera (such as formula (12)), the equation shown in the following formula (13) is obtained:
[0246] S4032: Use the Gauss-Newton method to solve the distortion model and obtain the coordinates of each pixel in the camera coordinate system.
[0247] For example, an optimization algorithm may be used to solve formula (13), and the coordinates of the pixel point (u', v') in the camera coordinate system are {x, y}.
[0248] For example, the Gauss-Newton method can be used to solve formula (13), and the coordinates of the pixel point (u', v') in the camera coordinate system are obtained as {x, y}.
[0249] Exemplarily, solving formula (13) can be regarded as finding the root of the function f(θ) to be found on the left side of the equal sign in formula (13).
[0250] The formula of the Gauss-Newton method can be shown as formula (14):
[0251] Where J is the Jacobian matrix, as shown in formula (15):
[0252] In formula (14), n is the number of iterations, is the {x, y} value of the nth iteration, J -1 is the transpose of the Jacobian matrix.
[0253] For example, the {x, y} value of the nth iteration can be substituted into formula (15) to obtain the following formula (16):
[0254] Among them, in formula (16) Refer to the following formula (17), formula (16) Refer to the following formula (18), formula (16) Refer to the following formula (19), formula (16) Please refer to the following formula (20).
[0255] It should be understood that the above method can also be used to obtain the coordinates of other pixel points in the camera coordinate system.
[0256] Compared with the prior art that uses a reverse ray tracing solution to determine the coordinates of pixel points in the simulated image under camera coordinates, this application uses the Gauss-Newton method to optimize the distortion model. The time required to determine the coordinates of pixel points in the simulated image under camera coordinates is shorter, and there is no dependency between each pixel point, so parallel calculations can be performed.
[0257] In addition, some existing technologies use a custom inverse distortion projection model to determine the coordinates of pixels in the simulated image in camera coordinates. Since the distortion parameters of the custom distortion camera have certain errors during the custom inverse distortion projection model process, the accuracy of the coordinates of pixels in the simulated image in camera coordinates determined by these existing technologies is not high. However, the present application determines the coordinates of pixels in the simulated image in camera coordinates by optimizing the distortion model of the distortion camera. This eliminates the need for custom distortion parameters of the distortion camera. As can be seen, the present application determines the coordinates of pixels in the simulated image in camera coordinates with greater accuracy.
[0258] It should be noted that another implementation of S403 may be as shown in S4033 to S4034 below:
[0259] S4033, obtaining a back-projection lookup table; wherein the back-projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image and the coordinates in the camera coordinate system.
[0260] Exemplarily, the back-projection lookup table may be established according to the above-mentioned steps S4031 to S4032, which will not be described in detail here.
[0261] For example, the back-projection lookup table is represented by LUT, where LUT(u',v')={x,y}.
[0262] S4034: Search the back-projection lookup table according to the position of each pixel point in the simulated image to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0263] It should be noted that this application does not limit the execution order of S403, S401 and S402.
[0264] S404 , determining the pixel point of each pixel in the simulated image in the environment map according to the coordinates of each pixel in the simulated image of the target distorted pinhole camera in the camera coordinate system.
[0265] For example, the following formula (21) and formula (22) can be used to determine the pixel point (u', v') in the environment map corresponding to the pixel point (u, v):
[0266] S405 , determining a pixel value of each pixel in the simulated image of the target distorted pinhole camera according to the pixel value of a corresponding pixel in the environment map.
[0267] Exemplarily, the pixel value of each pixel in the simulated image corresponding to the pixel in the environment map may be determined as the pixel value of each pixel in the simulated image.
[0268] 5 is a schematic diagram of an exemplary image generating apparatus 500. The image generating apparatus 500 can be used to execute the method of the aforementioned embodiment. Therefore, the beneficial effects achieved by the apparatus 500 can refer to the beneficial effects of the corresponding method provided above, which will not be described in detail here.
[0269] Exemplarily, the image generating device 500 includes:
[0270] A first scene acquisition module 501 is used to acquire a 3D scene of the target device's driving environment;
[0271] A first projection module 502 is configured to project the 3D scene onto N sides of a pyramid to obtain N environment maps, where N is an integer greater than or equal to 3;
[0272] A first pixel point determination module 503 is configured to determine the corresponding pixel point in the M environment maps for each pixel point in the simulated image of the target distorted camera; wherein M is a positive integer less than or equal to N, and the target device includes the target distorted camera;
[0273] The first pixel value determination module 504 is configured to determine the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the M environment maps.
[0274] Exemplarily, the first pixel value determination module 504 is specifically used to determine the pixel value of the first pixel corresponding to the pixel in an environment map as the pixel value of the first pixel when M is equal to 1; wherein the first pixel is any pixel in the simulated image.
[0275] Exemplarily, the first pixel value determination module 504 is specifically used to fuse the pixel values of the second pixel point corresponding to the M environment maps to obtain the pixel value of the second pixel point when M is greater than 1; wherein the second pixel point is any pixel point in the simulated image.
[0276] Exemplarily, the first pixel point determination module 503 is used to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system; determine the pixel point corresponding to each pixel point in the simulated image in the M environment maps based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; wherein the third pixel point is any pixel point in the simulated image, and the first environment map is any environment map among the N environment maps.
[0277] Exemplarily, the first pixel point determination module 503 is specifically used to determine the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; when it is determined that the pixel point corresponding to the third pixel point in the projection plane of the first environment map belongs to the first environment map based on the size of the first environment map and the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map, the pixel point corresponding to the third pixel point in the projection plane of the first environment map is determined as the pixel point corresponding to the third pixel point in the first environment map.
[0278] Illustratively, the first pixel point determination module 503 is specifically configured to obtain a distortion model of a target distortion camera; and solve the distortion model using the Gauss-Newton method to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0279] Exemplarily, the first pixel point determination module 503 is specifically used to obtain a back projection lookup table; wherein the back projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image and the coordinates in the camera coordinate system; according to the position of each pixel point in the simulated image in the simulated image, the back projection lookup table is searched to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
[0280] 6 is a schematic diagram of an exemplary image generating apparatus 600. The image generating apparatus 600 can be used to execute the method of the aforementioned embodiment. Therefore, the beneficial effects achieved by the apparatus 600 can refer to the beneficial effects of the corresponding method provided above, which will not be repeated here.
[0281] Exemplarily, the image generating device 600 includes:
[0282] The second scene acquisition module 601 is used to acquire the 3D scene of the target device's driving environment;
[0283] A second projection module 602 is used to project the 3D scene onto a plane to obtain an environment map;
[0284] A coordinate acquisition module 603 is configured to acquire the coordinates of each pixel in a simulated image of a target distortion camera in a camera coordinate system, wherein the coordinates of each pixel in the simulated image in the camera coordinate system are obtained by solving a distortion model of the target distortion camera using a Gauss-Newton method, where the target device includes the target distortion camera;
[0285] A second pixel point determination module 604 is configured to determine a corresponding pixel point in the environment map for each pixel point in the simulated image of the target distortion camera according to the coordinates of each pixel point in the simulated image in the camera coordinate system;
[0286] The second pixel value determination module 605 is configured to determine the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the environment map.
[0287] In an example, FIG7 shows a schematic block diagram of an apparatus 700 according to an embodiment of the present application. The apparatus 700 may include: a processor 701 and a transceiver / transceiver pin 702 , and optionally, a memory 703 .
[0288] The various components of the device 700 are coupled together via a bus 704, wherein the bus 704 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the various buses are collectively referred to as bus 704 in the figure.
[0289] Optionally, the memory 703 may be used to store instructions in the aforementioned method embodiment. The processor 701 may be used to execute the instructions in the memory 703 and control the receiving pin to receive a signal and control the transmitting pin to send a signal.
[0290] The apparatus 700 may be the electronic device or a chip of the electronic device in the above method embodiment.
[0291] Exemplarily, the electronic device may be a server or a terminal device.
[0292] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0293] The present application also provides a chip including one or more interface circuits and one or more processors. The one or more processors receive or send data via the one or more interface circuits. When the one or more processors execute computer instructions, the steps of the above-mentioned related methods are implemented. The interface circuit is a transceiver / transceiver pin 702.
[0294] This embodiment further provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the method in the above-mentioned embodiment.
[0295] This embodiment further provides a computer program product, which includes computer instructions. When the computer instructions are executed by a computer or a processor, the computer executes the above-mentioned related steps to implement the method in the above-mentioned embodiment.
[0296] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the methods in the above-mentioned method embodiments.
[0297] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0298] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0299] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0300] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0301] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0302] Any content of each embodiment of this application, as well as any content of the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0303] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0304] The steps of the method or algorithm described in conjunction with the disclosure of the embodiments of the present application can be implemented in a hardware manner, or can be implemented by a processor executing a software instruction. The software instruction can be composed of corresponding software modules, and the software module can be stored in a random access memory (Random Access Memory, RAM), a flash memory, a read-only memory (Read Only Memory, ROM), an erasable programmable read-only memory (Erasable Programmable ROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM), a register, a hard disk, a mobile hard disk, a read-only compact disc (CD-ROM) or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.
[0305] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0306] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. An image generation method, characterized in that: The method includes: Obtain the three-dimensional 3D scene of the target device's driving environment; Projecting the 3D scene onto N sides of a pyramid to obtain N environment maps, wherein N is an integer greater than or equal to 3; Determine the pixel point corresponding to each pixel point in the simulated image of the target distorted camera in M environment maps; wherein M is a positive integer less than or equal to N, and the target device includes the target distorted camera; The pixel value of each pixel in the simulated image is determined according to the pixel value of the corresponding pixel in the M environment maps.
2. The method according to claim 1, characterized in that Determining the pixel value of each pixel in the simulated image according to the pixel value of a corresponding pixel in the M environment maps includes: When M is equal to 1, the pixel value of the first pixel corresponding to the pixel in an environment map is determined as the pixel value of the first pixel; wherein the first pixel is any pixel in the simulation image.
3. The method according to claim 1 or 2, characterized in that Determining the pixel value of each pixel in the simulated image according to the pixel value of a corresponding pixel in the M environment maps includes: When M is greater than 1, the pixel values of the second pixel corresponding to the pixel points in the M environment maps are fused to obtain the pixel value of the second pixel; wherein the second pixel is any pixel in the simulated image.
4. The method according to claim 3, characterized in that The fusing pixel values of corresponding pixels of the second pixel in the M environment maps to obtain the pixel value of the second pixel includes: Interpolation is performed according to pixel values of corresponding pixels of the second pixel in the M environment maps to obtain a pixel value of the second pixel.
5. The method according to any one of claims 1 to 4, characterized in that The step of determining the pixel point corresponding to each pixel point in the simulated image of the target distortion camera in the M environment maps includes: Obtaining the coordinates of each pixel in the simulated image in the camera coordinate system; According to the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system, determine the pixel point corresponding to each pixel point in the simulated image in the M environment maps; wherein the third pixel point is any pixel point in the simulated image, and the first environment map is any environment map among the N environment maps.
6. The method according to claim 5, characterized in that The step of determining, based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system, a pixel point corresponding to each pixel point in the simulated image in the M environment maps includes: Determining a position of a pixel corresponding to the third pixel in the projection plane of the first environment map according to a projection matrix corresponding to the projection plane of the first environment map and a coordinate of the third pixel in the camera coordinate system; When determining, based on the size of the first environment map and the position of the pixel point corresponding to the third pixel point in the projection plane of the first environment map, that the pixel point corresponding to the third pixel point in the projection plane of the first environment map belongs to the first environment map, the pixel point corresponding to the third pixel point in the projection plane of the first environment map is determined as the pixel point corresponding to the third pixel point in the first environment map.
7. The method according to claim 5 or 6, characterized in that The obtaining of the coordinates of each pixel in the simulated image in the camera coordinate system includes: Obtaining a distortion model of the target distortion camera; The Gauss-Newton method is used to solve the distortion model to obtain the coordinates of each pixel in the simulated image in the camera coordinate system.
8. The method according to claim 5 or 6, characterized in that The obtaining of the coordinates of each pixel in the simulated image in the camera coordinate system includes: Obtaining a back-projection lookup table; wherein the back-projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image in the simulated image and the coordinates in the camera coordinate system; According to the position of each pixel point in the simulated image, the inverse projection lookup table is searched to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
9. The method according to any one of claims 1 to 8, characterized in that The target distortion camera includes a target fisheye camera.
10. The method according to any one of claims 1 to 9, characterized in that Any pair of adjacent side surfaces among the N side surfaces of the pyramid is orthogonal.
11. The method according to any one of claims 1 to 10, characterized in that At least one pair of adjacent side surfaces among the N side surfaces intersects.
12. The method according to any one of claims 1 to 11, characterized in that The target device includes a vehicle.
13. An image generating device, characterized in that: The device comprises: The first scene acquisition module is used to acquire the three-dimensional 3D scene of the target device's driving environment; A first projection module, configured to project the 3D scene onto N sides of a pyramid to obtain N environment maps, wherein N is an integer greater than or equal to 3; a first pixel point determination module, configured to determine a corresponding pixel point in M environment maps for each pixel point in a simulated image of a target distorted camera; wherein M is a positive integer less than or equal to N, and the target device includes the target distorted camera; The first pixel value determination module is used to determine the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the M environment maps.
14. The device according to claim 13, characterized in that The first pixel value determination module is specifically used to determine the pixel value of the corresponding pixel point of the first pixel point in an environment map as the pixel value of the first pixel point when M is equal to 1; wherein, the first pixel point is any pixel point in the simulated image.
15. The device according to claim 13 or 14, characterized in that The first pixel value determination module is specifically used to, when M is greater than 1, fuse the pixel values of the corresponding pixel points of the second pixel point in M environment maps to obtain the pixel value of the second pixel point; wherein, the second pixel point is any pixel point in the simulated image.
16. The device according to any one of claims 13 to 15, characterized in that The first pixel point determination module is used to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system; determine the pixel point corresponding to each pixel point in the simulated image in the M environment maps based on the projection matrix corresponding to the projection plane of the first environment map and the coordinates of the third pixel point in the camera coordinate system; wherein the third pixel point is any pixel point in the simulated image, and the first environment map is any environment map among the N environment maps.
17. The device according to claim 16, characterized in that The first pixel point determination module is specifically used to obtain the distortion model of the target distortion camera; use the Gauss-Newton method to solve the distortion model to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
18. The device according to claim 16, characterized in that The first pixel point determination module is specifically used to obtain a back projection lookup table; wherein, the back projection lookup table includes a mapping relationship between the position of each pixel point in the simulated image in the simulated image and the coordinates in the camera coordinate system; according to the position of each pixel point in the simulated image in the simulated image, the back projection lookup table is searched to obtain the coordinates of each pixel point in the simulated image in the camera coordinate system.
19. An image generation method, characterized in that: The method includes: Obtain the three-dimensional 3D scene of the target device's driving environment; Projecting the 3D scene onto a plane to obtain an environment map; Obtaining coordinates of each pixel point in a simulated image of a target distortion camera in a camera coordinate system, wherein the coordinates of each pixel point in the simulated image in the camera coordinate system are obtained by solving a distortion model of the target distortion camera using a Gauss-Newton method, and the target device includes the target distortion camera; Determine, based on the coordinates of each pixel in the simulated image in the camera coordinate system, a pixel point corresponding to each pixel in the simulated image of the target distortion camera in the environment map; The pixel value of each pixel in the simulated image is determined according to the pixel value of a corresponding pixel in the environment map.
20. The method according to claim 19, characterized in that Determining the pixel value of each pixel in the simulated image according to the pixel value of the corresponding pixel in the environment map includes: The pixel value of the corresponding pixel point in the environment map for each pixel point in the simulated image is determined as the pixel value of each pixel point in the simulated image.
21. The method according to claim 19 or 20, characterized in that The target distortion camera is a target distortion pinhole camera.
22. An electronic device, characterized in that: include: a memory and a processor, the memory being coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 12, or executes the method according to any one of claims 19 to 21.
23. A chip, characterized in that: The method comprises one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method according to any one of claims 1 to 12 are executed, or the steps of the method according to any one of claims 19 to 21 are executed.
24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when run on a computer or a processor, enables the computer or the processor to execute the method according to any one of claims 1 to 12, or, or execute the method according to any one of claims 19 to 21.
25. A computer program product, characterized in that The computer program product comprises computer instructions, which, when executed by a computer or a processor, cause the steps of the method according to any one of claims 1 to 12 to be performed, or cause the steps of the method according to any one of claims 19 to 21 to be performed.
Citation Information
Patent Citations
Image rendering method and device, electronic equipment and storage medium
CN112215936A
Way to generate images with distortion for fisheye lens
CN112837209A
Vehicle key point information detection and vehicle control
WO2022121283A1
Cited By
Infrared image enhancement method and device, equipment and storage medium
CN121504738A