METHOD FOR CREATING A PANORAMIC IMAGE FOR A VEHICLE ALL-AROUND VISION SYSTEM

By utilizing depth information and camera parameters to accurately calculate spatial coordinates, the method addresses the Manhattan effect in vehicle surround-view systems, ensuring a realistic and undistorted panoramic image reconstruction.

DE102025138629A1Pending Publication Date: 2026-04-02ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025138629
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-29
Filing Date
2025-09-24
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Conventional vehicle surround-view systems suffer from the Manhattan effect, where objects near the vehicle are excessively magnified or stretched, leading to significant distortions in panoramic images due to the lack of depth information in image projections.

Method used

A method is employed to create panoramic images by retrieving depth information for each pixel, calculating accurate spatial coordinates using intrinsic and extrinsic camera parameters, and converting these coordinates into a vehicle body coordinate system to accurately reconstruct the vehicle's surroundings, thereby avoiding distortions.

Benefits of technology

The method accurately reconstructs the vehicle's surroundings without excessive magnification or stretching, providing a realistic and undistorted panoramic view by determining the precise spatial position of each pixel in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for creating a panoramic image for a vehicle surround-view system is disclosed. The method comprises: obtaining depth information from an image captured by each camera of at least one camera of a vehicle surround-view system; calculating first coordinates of a real-world scene position corresponding to a pixel in a first spatial coordinate system using intrinsic camera parameters and the depth value of the pixel; converting the first coordinates of the real-world scene position corresponding to the pixel into second coordinates in a second spatial coordinate system using extrinsic camera parameters; and creating a panoramic image for the vehicle surround-view system based on each captured image and the second coordinates of the real-world scene position corresponding to each pixel in each image.The panoramic image created according to the method of some embodiments of the present application can recreate the vehicle's surroundings more realistically and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF INVENTION

[0001] The present application relates to the field of image processing and in particular to a method for creating panoramic images for a vehicle surround view system, a vehicle surround view system and a computer program product. STATE OF THE ART

[0002] With the rapid development of image processing and computer vision technologies, more and more technologies are being used in automotive electronics. Conventional rear-view camera systems can only cover a limited area around the rear of the vehicle, while other areas around the vehicle (including the front and sides) remain blind spots, undoubtedly posing an increased safety risk. To expand the driver's field of vision, perceive the vehicle's surroundings in a 360-degree panorama, and gain a comprehensive overview of the vehicle's environment, the surround-view vehicle system (or parking assistance system with panoramic view) was developed.

[0003] In panoramic images from vehicle surround-view systems that use appropriate technologies, real-world objects (such as vehicles, buildings, or people) located near the vehicle cameras or imaging devices can appear infinitely enlarged within the panoramic view (particularly with increased height). This occurs because such objects are projected onto predetermined surfaces at a specific distance from the vehicle (e.g., the perimeter surfaces of the panoramic image).

[0004] This phenomenon is commonly referred to as the Manhattan effect, as in Fig. Figure 1 illustrates this. Generally, the closer an object is to the vehicle camera, the more its image is magnified or stretched in the panoramic image, and it can even deviate completely from the object's true shape and / or size, leading to significant distortions in the panoramic image results from the vehicle's surround-view system. Therefore, it is essential to find a solution for processing images captured by the vehicle's surround-view system to avoid or mitigate distortions in panoramic images caused by the Manhattan effect. REVELATION OF THE INVENTION

[0005] Against this background, the present application provides a method for creating panoramic images for a vehicle surround view system, a vehicle surround view system and a computer program product, in the hope of mitigating or overcoming some or all of the above-mentioned shortcomings and other possible shortcomings.

[0006] According to a first aspect of the present application, a method for creating a panoramic image for a vehicle surround-view system is provided, comprising: retrieving depth information for an image captured by at least one camera in the vehicle surround-view system, wherein the depth information includes a depth value for each pixel in the image; calculating, for each pixel in an image captured by each camera of the at least one camera, using intrinsic camera parameters and a depth value of the pixel, first coordinates of a real scene position corresponding to the pixel in a first spatial coordinate system, wherein the first spatial coordinate system represents a camera coordinate system corresponding to the camera;Converting, for each pixel in the image captured by each camera of the at least one camera, the first coordinates of a real-world scene position corresponding to the pixel, using an extrinsic camera parameter, into second coordinates in a second spatial coordinate system, wherein the second spatial coordinate system represents a vehicle body coordinate system corresponding to the vehicle's surround-view system; creating a panoramic image for the vehicle's surround-view system based on the images captured by at least one camera and the second coordinates of the real-world scene position corresponding to each pixel in each image.

[0007] In some embodiments of the present application, for each pixel in the image captured by each camera of the at least one camera, the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system are calculated using the intrinsic camera parameters and the depth value of the pixel, including: for each pixel in the image captured by each camera of the at least one camera, the camera ray equation of the pixel is calculated using the intrinsic camera parameters, wherein the camera ray equation of the pixel is the equation of the straight line passing through the optical center of the camera and the pixel in the first spatial coordinate system; for each pixel in the image captured by each camera of the at least one camera, the first coordinate of the real scene position corresponding to the pixel in the first spatial coordinate system is determined according to the depth value of the pixel and the camera ray equation.

[0008] In some embodiments of the present application, for each pixel in the image captured by each camera of the at least one camera, a camera ray equation of the pixel is calculated using the intrinsic camera parameters, comprising: performing the following steps for each pixel in the image captured by each camera of the at least one camera: acquiring third coordinates of the pixel in a coordinate system of a third plane, wherein the coordinate system of the third plane represents a pixel coordinate system of the image; calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinates of the pixel, wherein the normalized coordinates represent the coordinates of the projection point of the pixel in the normalized plane in the first spatial coordinate system;Determining a camera beam equation for the pixel according to the normalized coordinates corresponding to the pixel, where the optical center of the camera is the origin of the first spatial coordinate system.

[0009] In some embodiments of the present application, for each pixel in the image captured by each camera of the at least one camera, the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system are calculated using the intrinsic camera parameter and the depth value of the pixel, comprising: for each pixel in the image captured by each camera of the at least one camera, the following steps are performed: acquiring the third coordinates of the pixel in a coordinate system of a third plane, wherein the coordinate system of the third plane represents the pixel coordinate system of the image; calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinates of the pixel, wherein the normalized coordinates represent the coordinates of the projection point of the pixel in the normalized plane in the first spatial coordinate system;and determining the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system according to the depth value of the pixel.

[0010] In some embodiments of the present application, for each pixel in the image captured by each camera of the at least one camera, the first coordinates of the real scene position corresponding to the pixel are converted into second coordinates in the second spatial coordinate system using the extrinsic camera parameter, wherein the following steps are performed for each pixel in the image captured by each camera of the at least one camera: converting the first coordinates of the real scene position corresponding to the pixel into homogeneous coordinates; creating an extrinsic parameter matrix using extrinsic camera parameters, wherein the extrinsic camera parameters comprise a rotation matrix representing a rotation relation from the first coordinate system to the second coordinate system, and a translation vector representing a translation relation from the first coordinate system to the second coordinate system;Determining the second coordinates of the real scene position corresponding to the pixel in the second coordinate system, according to the extrinsic parameter matrix and the homogeneous coordinates of the real scene position corresponding to the pixel.

[0011] In some embodiments of the present application, for each pixel in the image captured by each camera of the at least one camera, first coordinates of a real scene position corresponding to the pixel in a first spatial coordinate system are calculated using intrinsic camera parameters and a depth value of the pixel, comprising: rectifying an image captured by each camera of the at least one camera using an intrinsic camera parameter to obtain a rectified image, wherein the pixels in the image correspond one-to-one to the rectified pixels in the rectified image; capturing, for each pixel in the image captured by each camera of the at least one camera, fourth coordinates of a rectified pixel corresponding to the pixel in a third-plane coordinate system, wherein the third-plane coordinate system represents a pixel coordinate system of the image;Calculate for each pixel in the image captured by each camera of the at least one camera using the intrinsic camera parameter and the depth value of the pixel from the first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system, as well as the fourth coordinates of the rectified pixel corresponding to the pixel.

[0012] In some embodiments of the present application, for an image captured by each camera of the at least one camera, a rectification process is performed on the image using an intrinsic camera parameter to obtain a rectified image, comprising: acquiring a distortion coefficient and an intrinsic parameter matrix from the intrinsic camera parameters; determining a position mapping relationship of pixels of the image to rectified pixels based on the distortion coefficient and the intrinsic parameter matrix; and mapping each pixel in the image to the position of a corresponding rectified pixel using the position mapping relationship to produce a rectified image.

[0013] In some embodiments of the present application, at least one camera of the vehicle surround view system comprises a monocular fisheye camera, and for an image captured by each camera of the at least one camera of the vehicle surround view system, depth information of the image is acquired, comprising: predictions for an image captured by each camera of the at least one camera of the vehicle surround view system of the depth information of the image using an image depth estimation model, wherein the image depth estimation model comprises a trained convolutional neural network model.

[0014] In some embodiments of the present application, depth information of the image is acquired for an image acquired by each camera of the at least one camera of a vehicle surround view system, comprising acquiring the depth information of the image acquired by each camera by at least one of the following methods: binocular or multi-camera image depth estimation; image depth estimation based on structured light; image depth estimation based on LiDAR.

[0015] In some embodiments of the present application, a panoramic image for the vehicle's surround-view system is created based on the images captured by the at least one camera and the second coordinates of the real scene position corresponding to each pixel in each image, comprising: for each image captured by the at least one camera, reconstructing the local three-dimensional image corresponding to the image based on the image and the second coordinates of each pixel contained therein; and creating a panoramic image for the vehicle's surround-view system based on the local three-dimensional images corresponding to the images captured by the at least one camera.

[0016] In some embodiments of the present application, for each image captured by the at least one camera, a local three-dimensional image corresponding to the image is reconstructed based on the second coordinates of the image and each pixel contained therein, comprising: performing the following steps for each image captured by the at least one camera: generating point cloud data based on the second coordinates of each pixel in the image, wherein the point cloud is a set of discrete points in three-dimensional space corresponding to each pixel; extracting surface information of the real scene corresponding to the image from the point cloud data to create a local three-dimensional model; performing surface optimization processing on the local three-dimensional model to obtain an optimized local three-dimensional model;Obtaining pixel values ​​of each pixel of the image and their mapping relationship with the point cloud data; and rendering processing on the surface of the optimized local three-dimensional model based on the pixel values ​​of each pixel of the image and their mapping relationship with the point cloud data to obtain a local three-dimensional image corresponding to the image.

[0017] According to a second aspect of the present application, a vehicle surround view system is provided comprising: a plurality of cameras arranged at various locations on a vehicle body, and a processor, the processor being configured to perform a method according to some embodiments of the present application.

[0018] According to a third aspect of the present application, a computer program product is provided which comprises a computer program, wherein the computer program implements the steps of a method according to some embodiments of the present application when executed by a processor.

[0019] In a method for creating a panoramic image for a vehicle surround-view system according to some embodiments of the present application, the depth information of (two-dimensional) images taken in different directions by one or more vehicle cameras is first acquired, and then the three-dimensional spatial position of the scene corresponding to each pixel of the acquired image in the real world is acquired. This is done based on the depth information and the intrinsic and extrinsic parameters of the camera (for example, the first coordinates in the coordinate system of the vehicle body), which allows the actual spatial position of the real scene corresponding to each pixel in each image to be determined more accurately, since the image depth information (i.e.,The depth value of each pixel in the image accurately represents and reflects the actual spatial distance of the object from the corresponding scene in the image to the camera. A panoramic image is then created by the vehicle's surround-view system, capturing the three-dimensional spatial position of the scene corresponding to each pixel in the image based on the depth information. Therefore, the panoramic image, created using the precise spatial position of the real scene corresponding to each pixel in the image, can more realistically and accurately restore the original appearance of the vehicle's surroundings, thus eliminating image distortion problems and the Manhattan effect (e.g., the Manhattan effect).This avoids the over-enlargement or stretching of the scene near the on-board camera, or other serious distortion problems that arise in related technologies by mapping or projecting the panoramic image onto preset imaging boundaries (such as virtual walls and / or floor) around the vehicle.

[0020] These and other advantages of the present application will become apparent from the embodiments described below and will be illustrated by means of the embodiments described below. DESCRIPTION OF THE FIGURES

[0021] Preferred embodiments of the present application are described in more detail below with reference to the figures: Fig. Figure 1 schematically shows a diagram of a panoramic imaging principle for a vehicle surround view system in the state of the art; Fig. Figure 2A schematically illustrates an exemplary application scenario or implementation environment of a method for creating panoramic images for a vehicle surround view system according to some embodiments of the present application; Fig. Figure 2B shows an interactive process of an exemplary application scenario of a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application; Fig. Figure 3 shows a flowchart of a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application; Fig. Figure 4 schematically shows a geometric model of a method for creating a panoramic image for a vehicle all-round vision system according to some embodiments of the present application; Fig. Figure 5 schematically shows a comparison of the effects of a prior art method for creating a panoramic image and a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application; Fig. Figures 6A to 6C show an example process with initial coordinate calculation steps in a method for creating a panoramic image for a vehicle all-round vision system according to some embodiments of the present application; Fig. Figure 7 shows an example process with distortion correction steps in a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application; Fig. Figure 8 shows an example process of the second coordinate conversion step in a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application; Fig. Figure 9 shows an example process with panoramic image creation steps in a method for creating a panoramic image for a vehicle all-round vision system according to some embodiments of the present application; Fig. Figure 10 shows an example process with local three-dimensional image generation steps in a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application. DETAILED DESCRIPTION OF THE EXECUTION FORMS

[0022] The exemplary embodiments are now described in more detail with reference to the figures. However, these exemplary embodiments can be implemented in various forms and should not be considered limited to the embodiments described here. Rather, these embodiments serve to present the application comprehensively and completely and to fully convey the concept of exemplary embodiments to those skilled in the field. The same reference numerals are used in different figures to refer to the same or similar parts, which is why their repeated description is omitted.

[0023] Furthermore, the described features, structures, or properties can be combined in any suitable manner in one or more embodiments. Numerous specific details are set forth in the following description to enable a comprehensive understanding of the embodiments of the present invention. However, those skilled in the field will recognize that the technical solutions of the present application can be implemented without one or more of the specific details, or that other methods, components, devices, steps, etc., can be used. In other cases, generally known methods, devices, implementations, or processes are not shown or described in detail so as not to obscure the various aspects of the present application.

[0024] The blocks depicted in the figures are merely functional units and do not necessarily correspond to physically separate units. That is, these functional units can be implemented in software, in one or more hardware modules or integrated circuits, or in various networks and / or processor devices and / or microcontroller devices.

[0025] The flowcharts shown in the figures serve only as examples and do not necessarily include all content and processes / steps, nor do they have to be executed in the described order. For example, some processes / steps can be broken down, while others can be combined or partially combined, so the actual execution sequence may change depending on the actual circumstances.

[0026] It is understood that while the terms "first," "second," "third," etc., may be used herein to describe various components, these components should not be restricted by these terms. Therefore, a first component discussed below could be referred to as a second component without departing from the principles of the concepts of this application. The terms "and / or" and similar terms used herein encompass all combinations of any, several, or all of the listed related elements.

[0027] Experts in the field will understand that the figures are merely schematic diagrams of exemplary embodiments and that the modules or processes in the figures are not necessarily required for the implementation of the present application and therefore may not be used to limit the scope of protection of the present application.

[0028] Before the embodiments of the present application are presented in detail, some related concepts will first be explained for the sake of clarity. 1. Vehicle surround view system

[0029] Also known as a vehicle surround-view panoramic parking aid system, this system provides more intuitive visual assistance for driving a motor vehicle. It uses multiple wide-angle or fisheye cameras mounted around the vehicle to capture images of the front, rear, left, and right sides. After processing, the image information is displayed on the onboard display as a panoramic bird's-eye view. This helps the driver eliminate blind spots and improves driving safety and comfort. The vehicle surround-view system uses fisheye or ultra-wide-angle cameras and utilizes the principle of transmission and reflection through physical optical spherical mirrors to simultaneously capture a 360-degree horizontal field of view and a 180-degree vertical field of view.The software supplied with the hardware is then used to convert the image and display it in a way that is familiar to the human eye. The images captured by these cameras are converted into digital information by the image capture component and sent to the video synthesis / processing component for distortion correction, perspective conversion, image stitching, and image enhancement. Finally, the data is converted into an analog signal output to generate panoramic image information of the vehicle and its surroundings on the onboard display. A vehicle surround-view system is frequently used in various driving scenarios, particularly at low speeds, during perpendicular parking, parallel parking, reversing, in narrow passages, and in complex road conditions, and can significantly improve driver control and safety.Furthermore, the vehicle's all-round vision system offers excellent support even under the conditions of dense urban traffic. 2. Geometric model of the camera imaging for the vehicle's surround view system (see Fig. 4)

[0030] In this geometric model, there are usually four different coordinate systems, namely the camera coordinate system O c -X c Y c Z c (first spatial coordinate system), the vehicle body coordinate system (or world coordinate system, second spatial coordinate system) O w -X w Y w Z w , the image coordinate system xOy and the pixel coordinate system uOv (third plane coordinate system).

[0031] The vehicle body coordinate system (also called the vehicle coordinate system) is a three-dimensional rectangular coordinate system used to describe the motion of a motor vehicle. Its origin usually coincides with the vehicle's center of mass, but other points can be chosen as the origin depending on the situation. In the vehicle coordinate system, the X w -axis normally parallel to the ground and points towards the front (or rear, depending on the definition) of the vehicle, the Z w The -axis points upwards through the vehicle's center of mass, and the Y wThe x-axis points to the left (or right, depending on the definition) of the driver. The vehicle body coordinate system is primarily used to describe parameters such as the vehicle's position, attitude, and speed while driving. The camera coordinate system is also a three-dimensional rectangular coordinate system with the origin O. c in the optical center of the camera lens. The X c - and Y c The x-axes run parallel to the x- and y-axes of the image coordinate system. c The x-axis is the optical axis of the lens and is perpendicular to the image plane. The camera coordinate system is primarily used to describe the spatial position of every point in the captured image. It is one of the most widely used coordinate systems in computer vision and image processing.

[0032] The origin O1 of the image coordinate system is the intersection of the optical camera axis with the image plane (principal point), i.e., the center of the image. The x- and y-axes run parallel to the u- and v-axes of the pixel coordinate system. The distance between the origin O c The origin O1 of the camera coordinate system and the origin O1 of the image coordinate system correspond to the focal length f of the camera. The origin O2 of the pixel coordinate system is located in the upper left corner of the image. The u and v axes run parallel to the two sides of the image plane and reflect the arrangement of the pixels in the camera's CCD / CMOS chip. Both the image coordinate system and the pixel coordinate system are two-dimensional rectangular coordinate systems defined in the image plane.

[0033] It should be noted that the units of the coordinate axes u and v in the pixel coordinate system are the number of pixels (integers), while the units of the other three coordinate systems are usually units of length such as m or mm. 3. Intrinsic and extrinsic parameters of the camera

[0034] Camera intrinsic parameters are parameters that describe the internal properties of a camera. These parameters are typically captured during camera calibration, are fixed for a specific camera model, and do not change over time. Once the intrinsic camera parameters are captured, they usually remain unchanged during the camera's use. Camera intrinsic parameters can include focal length, principal point (optical center), pixel coordinates, distortion coefficients, pixel size (dx, dy), and so on. The focal length (f) is a key parameter of a camera lens that determines the size of the image formed on the image plane after light passes through the lens. The longer the focal length, the closer the objects appear in the image; the shorter the focal length, the farther away the objects appear.The principal point is the intersection of the optical axis of the camera lens and the image plane, and is also the center of the image plane. In the pixel coordinate system, the coordinates of the principal point are usually given in pixels. Due to manufacturing and installation errors of the camera lens, light is distorted when projected onto the image plane. Distortion coefficients are used to describe the degree and type of distortion, including radial distortion coefficients (such as k1, k2, k3, etc.) and tangential distortion coefficients (such as p1, p2, etc.). Pixel size refers to the actual physical size of each pixel on the image plane, i.e., the actual length or width represented by a pixel. Furthermore, depending on the camera model and calibration procedure, the intrinsic parameters of the camera may also include other parameters such as the skew factor.

[0035] In this document, extrinsic camera parameters are parameters that describe the position and attitude of the vehicle body within the camera coordinate system. These parameters typically include a rotation matrix (R) and a translation vector (t), which are used to convert points in the camera coordinate system to the vehicle body's coordinate system. The extrinsic parameters can change depending on the camera position or the time of capture. Rotation matrix (R): The rotation matrix is ​​a 3x3 matrix that describes the rotation relationship from the camera coordinate system to the vehicle coordinate system. Each column of the rotation matrix is ​​a unit vector representing the direction of the X-axis. w -axis, Y w -axis and Z w-axis in the vehicle body coordinate system. Translation vector (t): The translation vector is a 3x1 matrix (or vector) that describes the position of the vehicle in the camera coordinate system. The three components of the translation vector represent the coordinates of the origin of the vehicle body coordinate system in the camera coordinate system.

[0036] 4. Depth information of the image captured by the camera: This usually refers to the relative distance between the object or scene represented by each pixel in the image and the viewer (or the optical center of the camera). This can be achieved using special (depth) cameras or other technical means. This depth information is widely used in 3D reconstruction, virtual reality, augmented reality, machine vision, and other fields.

[0037] Fig. Figure 1 schematically shows a diagram of a panoramic imaging principle for a vehicle surround view system in the state of the art.

[0038] As in Fig. As shown in Figure 1, a vehicle camera 120 is installed on the right side of a vehicle 110 located on the ground 100. Since the image captured by the camera 120 lacks depth information in the prior art, the panoramic image corresponding to the camera 120 on the right side of the vehicle 110 can only be projected onto a preset panoramic boundary surface; that is, the imaging area (or projection surface) comprises the ground 100 and the virtual wall 130 between the camera 120 and a preset virtual wall 130 (e.g., with a preset height). As shown in Figure 1, the image area (or projection surface) comprises the ground 100 and the virtual wall 130 between the camera 120 and a preset virtual wall 130 (e.g., with a preset height). Fig. As shown in Figure 1, a person 140 is standing near the right side of the vehicle 110. Based on state-of-the-art panoramic imaging principles, the camera 120 projects the upper half of the person 140 onto the virtual wall 130 (Figure 150a between projection points A1 and A2). Simultaneously, the camera 120 projects the lower half (the shaded area) of the person 140 onto the ground 100 (Figure 150b between A2 and A3). As shown in Fig. As shown in Figure 1, assuming that the optical center of the camera lens is C, the distance (or actual depth) from the optical center C of the camera to various points of person 140 (e.g., the upper part R1 and the waist R2) is much smaller than the distance (i.e., the image depth) from the optical center C to corresponding positions (e.g., A1 and A2) of the panoramic images 150a and 150b of the person. As shown in Fig. As shown in Figure 1, this creates the Manhattan effect, in which both the upper part 150a and the lower part 150b of the panoramic image of person 140 (relative to the actual size of person 140) are significantly stretched, resulting in a strong distortion of the panoramic display of the vehicle's surround view system, making it no longer able to accurately and faithfully reproduce the vehicle's surroundings or scene.

[0039] In response to the aforementioned problems in the prior art, the present application proposes a method for creating a three-dimensional panoramic image of a vehicle based on the depth information of a two-dimensional image captured by a camera. In the panoramic imaging method according to some embodiments of the present application, to avoid or mitigate the Manhattan effect (i.e., the problem of severe distortion in vehicle panoramic views), a vehicle panoramic image is created by capturing the depth values ​​(i.e., depth information) of each pixel from the images captured by the various cameras of the vehicle's surround-view system (since the pixel depth values ​​reflect the distance from the corresponding real-world scene to the camera). In this way, the actual spatial position (i.e.,The three-dimensional spatial coordinates of the real scene, corresponding to each pixel in each image, are determined more accurately using the depth information from the images captured by each camera. Therefore, the panoramic image created based on the precise spatial position of each pixel in the image can reconstruct the scene around the vehicle more realistically and accurately. This avoids serious distortion problems such as excessive magnification and stretching of the scene near the vehicle camera, which are caused in related technologies by mapping or projecting onto a predefined perimeter or boundary surface (such as a virtual wall and / or the ground).

[0040] Fig. Figure 2A schematically shows an exemplary implementation environment 200 for a method for creating panoramic images for a vehicle surround-view system according to some embodiments of the present application. As in Fig. As shown in Figure 2A, the implementation environment 200 can include a vehicle terminal 210. In some embodiments, the terminal 210 can be used to implement a method for creating panoramic images according to the present application. For example, the terminal 210 can be equipped with appropriate programs or commands for executing the various methods provided in this application.For example, the terminal 210 can be a vehicle surround-view system, which may include image capture devices such as cameras or video cameras, usually installed at the front, rear, left and right of the vehicle (some systems may include additional cameras to cover a wider area); an image or video processing device responsible for correcting, stitching and fusing the captured images to produce a three-dimensional panoramic image or video around the vehicle; and an on-board display responsible for showing the processed panoramic image to a user, such as the driver.

[0041] Optionally, the implementation environment 200 can be used in addition to the vehicle terminal 210, as described in Fig. Figure 2A shows that the system further comprises a server 220 and a network 230 for connecting the terminal device 210 and the server 220. The server 220 and the terminal device 210 can also cooperate to implement various methods according to the present application. Therefore, the vehicle terminal device 210 and the server 220 together form a vehicle surround-view system, wherein the terminal device 210 comprises an image acquisition device for capturing images around the vehicle and an on-board display for showing the generated panoramic image, and the server 220 can comprise an image processing device or an image processing component for processing the captured images to generate or create a panoramic image around the vehicle.

[0042] The terminal device 210 can also be any other type of mobile computing device capable of creating a panoramic image of a vehicle, including a mobile computer (e.g., a personal digital assistant (PDA), laptop, notebook, tablet, netbook, etc.), a mobile phone (e.g., a cell phone, smartphone, etc.), a wearable computing device (e.g., a smartwatch, a head-worn device, including smart glasses, etc.), or other types of mobile devices (e.g., an intelligent in-vehicle system (navigation, autonomous driving, etc.)). In some embodiments, the terminal device 210 can also be a stationary computing device equipped with a vehicle surround-view system function, such as a desktop computer, a game console, or a smart TV.

[0043] Server 220 can be a single server or a server cluster, or it can be a cloud server or cloud server cluster capable of providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functionalities, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It is understood that the servers mentioned herein are typically server computers with large memory and processing resources, but other implementations are possible. Alternatively, Server 220 can also be a standard desktop computer comprising a main unit, a display device, and other components.

[0044] Examples of network 230 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. The server 220 and the terminal device 210 can include at least one communication interface (not shown) that can communicate over network 230. Such a communication interface can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), a wired or wireless interface (such as IEEE 802.11 Wireless LAN (WLAN)), a WiMAX interface (Worldwide Interoperability for Microwave Access), an Ethernet interface, a USB interface (Universal Serial Bus), a cellular network interface, a Bluetooth interface, an NFC interface (Near Field Communication), etc.

[0045] As in Fig. As shown in Figure 2A, the terminal device 210 can include a display screen, and a terminal user can interact with a terminal application via the display screen. The terminal device 210 can interact with the server 220 via the network 230, for example, to send or receive data. End-user applications or games can be native applications, web applications, or lightweight applications such as mini-programs (e.g., mobile mini-programs). If the terminal application is a local application program that needs to be installed, the terminal application can be installed on the terminal device 210. If the terminal application is a web application, the terminal application or game can be accessed via a browser.If the end-device application is a mini-program, it can be opened directly on the 210 end-device by searching for relevant information (such as the application name, etc.) or by scanning its graphic code (such as a barcode, QR code, etc.), without installing the application. In this document, the end-device application could be an application for launching and managing a vehicle camera for image capture, for example, a mobile phone application for taking photos and / or videos. Furthermore, the end-device application could also be used to manage the entire vehicle panorama image creation process, such as image capture, image processing, etc.

[0046] Fig. Figure 2B shows an exemplary interaction flow diagram for implementing the method for creating panoramic images in a vehicle surround view system according to some embodiments of the present application within the scope of the application. Fig. 2A shows the exemplary implementation environment 200. The following refers to the example implementation environment shown in Fig. The exemplary interaction flow diagram shown in Figure 2B briefly describes the operating principle of a method for creating panoramic images for a vehicle all-round vision system according to some embodiments of the present application.

[0047] As in Fig. As shown in Figure 2B, the server 220 can first be configured to capture the depth information of each image taken by at least one camera of the vehicle's surround-view system, where the depth information includes the depth value of each pixel within the image. Second, the server 220 can be configured to use the intrinsic camera parameters and the pixel's depth value for each pixel in the image captured by each camera of the at least one camera to calculate the first coordinates of the real-world scene position corresponding to the pixel in the first spatial coordinate system, where the first spatial coordinate system represents the camera coordinate system corresponding to the camera.Here too, the Server 220 can be configured to use the external camera parameter for each pixel in the image captured by each camera of the at least one camera to convert the first coordinates of the real-world scene position corresponding to the pixel into second coordinates in a second spatial coordinate system, where the second spatial coordinate system represents the vehicle body's coordinate system corresponding to the vehicle's surround-view system. Finally, the Server 220 can be configured to create a panoramic image for the vehicle's surround-view system based on the images captured by the at least one camera and the second coordinates of the real-world scene position corresponding to each pixel in each image.

[0048] Alternatively, as indicated by the dashed lines in Fig. As shown in Figure 2B, before the process of creating a panoramic image begins on server 220, the terminal device 210 is configured to capture images from various directions around the vehicle using multiple built-in image capture devices (such as cameras or video cameras, etc.). These images serve as processing objects for subsequent image processing steps and are transmitted to server 220. Optionally, after the vehicle panoramic image has been created, terminal device 210 can be configured to receive the created panoramic image from server 220 and display the panoramic image of the vehicle on the built-in on-board display, allowing the driver or user to use it as an additional reference while driving or parking.

[0049] The implementation environment and the interaction flowcharts of Fig. 2A and Fig. Figure 2B serves only for illustration, and the street detection method according to the present application is not limited to the exemplary implementation environments and interaction flow diagrams shown. It is understood that although the server 220 and the terminal 210 are presented and described here as separate structures, they may be different components of the same computing device.

[0050] Optionally, all steps of the road detection method according to some embodiments of the present disclosure can also be implemented only on the terminal device 210, i.e., the terminal device 210 can be configured for the following: obtaining depth information of the image captured by each camera of the at least one camera of the vehicle surround view system, wherein the depth information of the image includes the depth value of each pixel in the image; calculating, for each pixel in an image captured by each camera of the at least one camera, using intrinsic camera parameters and a depth value of the pixel, first coordinates of a real scene position corresponding to the pixel in a first spatial coordinate system, wherein the first spatial coordinate system represents a camera coordinate system corresponding to the camera;Converting, for each pixel in the image captured by each camera of the at least one camera, the first coordinates of a real-world scene position corresponding to the pixel, using an extrinsic camera parameter, into second coordinates in a second spatial coordinate system, wherein the second spatial coordinate system represents a vehicle body coordinate system corresponding to the vehicle's surround-view system; creating a panoramic image for the vehicle's surround-view system based on the images captured by at least one camera and the second coordinates of the real-world scene position corresponding to each pixel in each image.

[0051] Fig. Figure 3 schematically shows a flowchart of a method for creating a panoramic image for a vehicle surround-view system according to some embodiments of the present application. The execution unit of the method described in Figure 3 is a diagram of a method for creating a panoramic image for a vehicle surround-view system according to some embodiments of the present application. Fig. The 3 methods shown for creating panoramic images can be used in Fig. 2A and Fig. The terminal device 210 shown in 2B (e.g. a vehicle surround view system) can be included or can optionally also include a server 220.

[0052] As in Fig. As shown in Figure 3, a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application may comprise the following steps: S310, step towards capturing depth information; S320, first step in calculating the coordinates; S330, second step to convert the coordinates; S340, step towards creating a panoramic image.

[0053] The execution process of steps S310 to S340 above is described in detail below in the order shown in the figure. In general, the above steps can be performed in an image or video processing unit of a vehicle surround-view system (e.g., terminal 210).

[0054] In step S310 (step to acquire depth information), depth information from an image captured by at least one camera of the vehicle's surround view system is acquired, with the depth information of the image including the depth value of each pixel in the image.

[0055] In order to avoid or mitigate the problem of excessive distortion of vehicle panorama images, corresponding to the Manhattan effect, as described in the concept of the present application, corresponding three-dimensional panorama images or reconstructions can be created for various different scenes in the images captured or photographed by at least one vehicle camera located around the vehicle (for example, four in the front, rear, left, and right directions) by determining the actual distances of these scenes to the camera lens or optical center. In other words, the vehicle panorama image can be reconstructed by determining the depth value of each pixel in the captured image (i.e., the depth information of the image) (since the pixel depth value reflects the object distance from the corresponding real-world scene to the camera). In this way, the actual spatial position (i.e.,The three-dimensional spatial coordinates of the real scene, corresponding to each pixel in each image, are determined more accurately using the depth information from the images captured by each camera. Therefore, the panoramic image created based on the precise spatial position of each pixel in the image can reconstruct the scene around the vehicle more realistically and accurately. The Manhattan effect, which arises in related technologies due to mapping onto a preset perimeter plane (i.e., serious distortion problems such as excessive magnification and stretching of the scene near the onboard camera), is avoided.

[0056] In this document, the term "camera" (on board the vehicle surround-view system) can refer to various image capture devices, such as dedicated camera devices with photographic capabilities like cameras, camcorders, and digital cameras, or it can more broadly encompass various mobile computing devices with image capture capabilities and image capture modules or components (such as cameras), such as smartphones, tablets, etc. In some embodiments, the main components of a camera may include: a light source (such as a flash lamp) that provides sufficient illumination for image capture to ensure image quality; an image sensor for converting optical signals into electrical signals; a lens for focusing light to produce a clear image; and an image capture card for converting analog signals into digital signals so that they can be processed by a computer.In general, the minimum camera in step S310 can comprise four cameras installed in the vehicle's surround-view system in four directions: front, rear, left, and right. Some vehicle surround-view systems can also add more cameras to cover a larger area.

[0057] To obtain depth information from the images captured by each vehicle camera around the vehicle, some embodiments employ an image depth estimation method, which can essentially comprise two approaches: first, a method based on an AI (artificial intelligence) machine learning model (especially deep learning), primarily used for estimating the depth of monocular images (for example, if the vehicle camera is a monocular fisheye camera, the image it captures is monocular); second, other methods. Among these is the deep learning-based AI model method, which learns the image-to-depth relationship by training neural networks and exhibits high accuracy and robustness.

[0058] In some embodiments, step S310 (obtaining depth information from an image captured by each camera of the at least one camera of the vehicle's surround-view system) may include: predicting the depth information of an image captured by each camera of the at least one camera of the vehicle's surround-view system using an image depth estimation model, wherein the image depth estimation model comprises a trained convolutional neural network (CNN) model. CNNs have achieved great success in tasks such as image classification, object detection, and image segmentation, and are also frequently used in monocular depth estimation, for example, estimating the depth of images from monocular fisheye cameras. The image depth estimation model used here (e.g., images from a monocular fisheye camera) may employ an encoder-decoder architecture.The encoder compresses the input image into a low-dimensional feature representation, and the decoder reconstructs this feature representation into a depth map. This architecture can efficiently extract depth information from images. Commonly used loss functions in depth estimation for convolutional neural networks include L1 loss, L2 loss, and scale-invariant loss. These loss functions are used to measure the difference between the depth map predicted by the model and the actual depth map, thus guiding the model training process. To improve the model's generalization ability and performance, data extension techniques (such as random cropping, mirroring, and rotating) are typically employed to increase the variety of the training data. Simultaneously, model performance is optimized through adjustments to the network architecture, hyperparameters, and training strategies.Furthermore, the image depth estimation model can adopt a pyramid structure to better capture the multiscale information in the image. By extracting features at different scales and combining these features to predict depth information, the accuracy and robustness of the depth estimation can be improved. Image depth estimation models can also incorporate attention mechanisms that help the model pay more attention to important areas in the image, thus improving depth estimation performance. In depth estimation, the attention mechanism can be used to instruct the model to pay more attention to information such as edges and textures in the image, which aids in depth inference.

[0059] In some implementations, the training process of the machine learning-based image depth estimation model can be performed as follows: Supervised learning: Using a training dataset with known depth information to train a deep neural network. During the training process, the network predicts the depth information of new images by learning the mapping relationship between the input image and the corresponding depth map. Self-supervised learning: If no large-scale annotated depth data is available, self-supervised tasks such as image reconstruction are used to train the depth estimation network. For example, network parameters are optimized by minimizing the reconstruction error so that the network can learn useful depth information.Semi-supervised learning: This approach combines the advantages of supervised and self-supervised learning, using partially annotated depth data and a large number of unannotated images to train the depth estimation network. This approach allows for the comprehensive use of limited annotated data while simultaneously learning useful feature representations from a large amount of unannotated data.

[0060] In some embodiments, step S310 may also include: obtaining depth information from the image captured by each camera using at least one of the following methods: binocular or multi-camera depth estimation; structured light-based depth estimation; or LiDAR-based depth estimation. The principle of binocular or multi-camera depth estimation is to capture the same scene with two or more cameras and convert the parallax information between images into depth information through triangulation. The advantages are high accuracy and applicability to a wide variety of scenarios. Furthermore, LiDAR-based depth estimation is achieved by sending a laser beam at the target and receiving the reflected signal. The round-trip time difference of the laser beam is then measured to calculate the target distance, and the depth estimation is subsequently performed.The principle of depth estimation based on structured light involves calculating depth information by projecting a specific pattern of structured light onto the target and analyzing the distortion of the structured light on the target surface. This method is suitable for close-range measurements but requires strict environmental control.

[0061] In step S320 (first coordinate calculation step), for each pixel in the image captured by each camera of the at least one camera, the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system are calculated using the intrinsic camera parameter and the depth value of the pixel, where the first spatial coordinate system represents the camera coordinate system corresponding to the camera.

[0062] Based on the concept of this application, after the depth information of each image captured by the vehicle camera has been acquired by a depth camera or other specialized technology, the camera's own intrinsic and extrinsic parameters can be combined with the depth information of the image to create an accurate panoramic (three-dimensional) image or video of the vehicle. This is done to determine the actual spatial position or coordinates of the spatial scene represented by each pixel in each image, and then to create a three-dimensional image of the corresponding scene. Determining the spatial position (spatial coordinates) of the real scene corresponding to each pixel in the image above can be divided into two steps (i.e.,S320 and S330): First, using the intrinsic camera parameters (in particular the focal length (or normalized focal length), the pixel coordinates of the principal point, the pixel size, etc.) and the pixel depth value, the first coordinates of the real scene represented by the pixel in the camera coordinate system (i.e., the first spatial coordinate system) are calculated (S320); then, using the extrinsic camera parameters, the first coordinates of the pixel are converted into its second coordinates in the vehicle body coordinate system (i.e., the second spatial coordinate system) (S330).

[0063] Fig. Figure 4 schematically illustrates a geometric model of vehicle panorama imaging according to some embodiments of the present application.

[0064] As in Fig. As shown in Figure 4, assuming that the pixel coordinates of a pixel P in the imaging plane or pixel plane α of the vehicle camera are in the pixel coordinate system uOv (i.e., the coordinate system of the third plane) (u1, v1) and the first coordinates of the actual position P' of the corresponding real scene in three-dimensional space are in the camera coordinate system O c -X c Y c Z c (i.e., the coordinate system of the first space) (x1, y1, z1) is that, according to the camera mapping principle, the formula for expressing the relationship between pixel coordinates and camera coordinates is: [u1v11]=1z1[fx0cx0fycy001][x1y1z1]

[0065] In formula (1) f x , f y the normalized focal length of the camera on the u-axis and v-axis in the pixel coordinate system, fx=fdx,fy=fdy, f is the focal length of the camera, dx and dy are the pixel sizes on the u-axis and v-axis respectively in the pixel coordinate system; c x , c y are the pixel coordinates of the α principal point on the pixel plane (the intersection of the optical axis of the camera and the pixel plane); it should be noted that f x , f y , c x , c y , dx and dy are only related to the properties of the camera itself and belong to the intrinsic parameters of the camera; the matrix above [fx0cx0fycy001] is called the intrinsic parameter matrix and is denoted by K.

[0066] In some embodiments, intrinsic camera parameters, as described above, are parameters that describe internal properties of a camera. These parameters are usually determined during camera calibration, are fixed for a specific camera model, and do not change over time. In some embodiments, after acquiring the depth value of each pixel of the image captured by the vehicle camera and the corresponding intrinsic camera parameters, the spatial coordinates of the actual scene represented by each pixel in the camera coordinate system can be calculated using formula (1). As in Fig. As shown in Figure 4, the depth value for a pixel P (u1, v1) in an image (i.e., an image or pixel plane) is h, i.e., the distance from the actual scene P' corresponding to pixel P to the optical center O. cThe camera's coordinate system is h. The coordinates of P' in the camera coordinate system (x1, y1, z1) can then be calculated using formula (1). In particular, z1 can first be approximated as the depth value h, i.e., if z1 = h, then formula (1) yields: x1=h⋅u1−cxfx,y1=h⋅v1−cyfy,z1=h.

[0067] In step S330 (the second step for coordinate conversion), for each pixel in the image captured by each camera of the at least one camera, the extrinsic camera parameter is used to convert the first coordinates of the real scene, corresponding to the pixel in the first spatial coordinate system, into second coordinates in the second spatial coordinate system, where the second spatial coordinate system represents the vehicle body coordinate system, corresponding to the vehicle's all-around view system.

[0068] Based on the concept of the present application, after the first coordinate of the spatial position of the real scene has been determined using intrinsic parameters and depth values, corresponding to each pixel in the image in the first spatial coordinate system (i.e., the camera coordinate system), the extrinsic camera parameters (such as the rotation matrix R and the translation vector t) can be used to convert the first coordinates into the second coordinates in the second spatial coordinate system (i.e., the vehicle body coordinate system). This is because the extrinsic camera parameters themselves can describe the conversion between the vehicle body coordinate system and the camera coordinate system.

[0069] As in Fig. As shown in Figure 4, it is assumed that the first coordinate of a pixel P in the imaging plane or pixel plane α of the vehicle camera corresponds to the real scene position P' in the first spatial coordinate system (i.e., the camera coordinate system O). c -X c Y c Z c ) corresponds to x1, y1, z1), then according to the definition of the intrinsic camera parameters, the second coordinate of pixel P is that which corresponds to the real scene position P' in the second spatial coordinate system (i.e., the vehicle body coordinate system O). w -X w Y w Z w ) corresponds to (x2, y2, z2), which can be determined by the following formula (2): [x2y2z21]=[Rt0T1][x1y1z11]

[0070] In formula (2), R and t denote the rotation matrix and the translation vector, respectively, in the coordinate transformation process from the camera coordinate system (first spatial coordinate system) to the vehicle body coordinate system (second spatial coordinate system). The matrix [Rt0T1] This is called the extrinsic parameter matrix and is denoted by T. In the extrinsic parameter matrix T, the rotation matrix R is a 3x3 matrix used to describe the rotation relationship from the camera coordinate system to the vehicle coordinate system. Each column of the rotation matrix is ​​a unit vector representing the direction of the X. w -axis, Y w -axis and Z w The -axis represents the vehicle body coordinate system. The translation vector t is a 3x1 matrix (or a vector), and the three components represent the coordinates of the origin of the vehicle body coordinate system in the camera coordinate system.

[0071] In step S340 (step to create a panoramic image) a panoramic image is created for the vehicle's surround view system, based on the images taken by at least one camera and the second coordinates of the real scene position, which correspond to each pixel in each image.

[0072] According to the concept of the present application, the vehicle surround-view system uses each vehicle camera to capture two-dimensional images and obtain depth information. It then retrieves the spatial position coordinates (i.e., the second coordinates) of the real-world scene, corresponding to each pixel in each two-dimensional image within the vehicle body coordinate system, based on the depth information and the camera's intrinsic and extrinsic parameters. Subsequently, the three-dimensional panoramic image can be reconstructed based on the texture coordinates (i.e., pixel values) of the pixels in each two-dimensional image captured by each camera and the spatial geometric coordinates of the real-world scene corresponding to those pixels.

[0073] In some embodiments, step S340 for creating the panoramic image can essentially comprise two main steps: First, for each two-dimensional image captured by each vehicle camera, a local three-dimensional image, or an image corresponding to the two-dimensional image, is created based on the texture coordinates or pixel values ​​of each pixel (corresponding to the color of the real scene) and the corresponding three-dimensional geometric coordinates (secondary coordinates) of the real scene; then, the (created) local three-dimensional images corresponding to each two-dimensional image are combined and merged to obtain a 360-degree panoramic image around the vehicle. In particular, image preprocessing may be performed first when creating a local three-dimensional image, i.e.,The original image captured by the camera or lens is corrected for lens distortion, image brightness and contrast are adjusted, and other preprocessing operations are performed to ensure image quality. The appropriate algorithm for three-dimensional reconstruction is then used. Image stitching and merging can involve the use of professional image processing software or algorithms to precisely stitch and merge images captured by multiple cameras. During the stitching process, factors such as geometric relationships and color consistency between the images must be considered to ensure that the resulting three-dimensional panorama is optically seamless. Commonly used stitching algorithms may include those based on feature point matching, image content matching, and other factors.The process is based on the initial image stitching and merging. Optionally, after stitching and merging the images, further optimization and rendering steps can be included to refine the details of the resulting three-dimensional panorama. This includes removing image noise, improving image details, adjusting image brightness and contrast, and so on. Rendering and processing can make three-dimensional panoramas more realistic and clearer.

[0074] In a method for creating a panoramic image for a vehicle surround-view system according to some embodiments of the present application, the depth information of (two-dimensional) images taken in different directions by one or more vehicle cameras is first acquired, and then the three-dimensional spatial position of the scene corresponding to each pixel of the acquired image in the real world is acquired. This is done based on the depth information and the intrinsic and extrinsic parameters of the camera (for example, the first coordinates in the coordinate system of the vehicle body), which allows the actual spatial position of the real scene corresponding to each pixel in each image to be determined more accurately, since the image depth information (i.e.,The depth value of each pixel in the image accurately represents and reflects the actual spatial distance of the object from the corresponding scene in the image to the camera. A panoramic image is then created by the vehicle's surround-view system, capturing the three-dimensional spatial position of the scene corresponding to each pixel in the image based on the depth information. Therefore, the panoramic image, created using the precise spatial position of the real scene corresponding to each pixel in the image, can more realistically and accurately restore the original appearance of the vehicle's surroundings, thus eliminating image distortion problems and the Manhattan effect (e.g., the Manhattan effect).This avoids the over-enlargement or stretching of the scene near the on-board camera, or other serious distortion problems that arise in related technologies by mapping or projecting the panoramic image onto preset imaging boundaries (such as virtual walls and / or floor) around the vehicle.

[0075] Fig. Figure 5 schematically shows a comparison of the effects of a prior art method for creating a panoramic image and a method for creating a panoramic image for a vehicle surround-view system according to some embodiments of the present application. (a) and (b) Fig. Figure 5 schematically illustrates a prior art method for constructing a panoramic image or a flowchart of a method for constructing a panoramic image according to some embodiments of the present application. As in (a) in Fig. As shown in Figure 5, the car 520 parked near vehicle 510 is clearly deformed in the panoramic image of vehicle 510 created using the prior art method. In particular, the car body is significantly stretched and distorted longitudinally (i.e., vertically). This is because the imaging process does not take into account the depth information of the image captured by the on-board camera and the actual spatial position information of the scene corresponding to each pixel. In comparison, as shown in (b) in Fig. 5 shown, in the panoramic image of the motor vehicle 510, which was created according to the panoramic image creation method of some embodiments of the present application, the image of the parked motor vehicle 530 (the same vehicle as the motor vehicle 520 in (a)) in the vicinity of the motor vehicle 510 shows no obvious deformation or stretching or other distortion phenomena and restores the original shape of the motor vehicle more realistically and accurately.

[0076] Fig. Figure 6A schematically illustrates an example process with initial coordinate calculation steps according to some embodiments of the present application.

[0077] As in Fig. As shown in Figure 6A, the first coordinate calculation steps S320 in some embodiments may include the following: Performing the following steps for each pixel in the image captured by each camera of the at least one camera: S321, Acquiring third coordinates of the pixel in a third-plane coordinate system, wherein the third-plane coordinate system represents a pixel coordinate system of the image; S322, Calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinates of the pixel, where the normalized coordinates represent the coordinates of the projection point of the pixel on the normalized plane in the first spatial coordinate system; S323, Calculating the normalized coordinates corresponding to the pixel according to the depth value of the pixel, in order to determine the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system.

[0078] As described in S321, according to the concept of the present application, it is first necessary to determine the spatial position coordinates of the spatial scene point P', which corresponds to pixel P in Fig. 4 corresponds (for example, the coordinates in the camera coordinate system (first spatial coordinate system)), first to capture the two-dimensional pixel coordinates of point P (the third coordinates in the coordinate system of the third plane), for example (u1, v1), and then to use the intrinsic camera parameters and their pixel depth value to obtain the final result by coordinate transformation.

[0079] Next, as in Fig. As shown in Figure 4, in step S322 the two-dimensional pixel coordinates of point P are converted into three-dimensional normalized coordinates of the projection point P" in the normalized plane β using the camera mapping principle. The normalized plane (also called the Z=1 plane in the camera coordinate system) is a virtual plane used to simplify the camera model and the projection process. For example, the normalized coordinates (x0, y0, z0) of the projection of point P onto the normalized plane β (point P") are: x0=u1−cxfx,y0=v1−cyfy,z0=1. Finally, as described in S323, the depth value h is used to convert the normalized coordinates (x0, y0, z0) into the first coordinates (x1, y1, z1) of the real scene position P', which corresponds to pixel P in the camera coordinate system (i.e., the first spatial coordinate system), i.e.: x1=h⋅u1−cxfx,y1=h⋅v1−cyfy,z1=h.

[0080] Fig. Figure 6B schematically illustrates an example process with initial coordinate calculation steps according to some embodiments of the present application.

[0081] In some embodiments, the first coordinate calculation steps S320 can also be implemented in the following way: for each pixel in the image captured by each camera of the at least one camera, the camera ray equation of the pixel is calculated using the intrinsic camera parameters, where the camera ray equation of the pixel represents the equation of the straight line passing through the optical center of the camera and the pixel in the first spatial coordinate system; for each pixel in the image captured by each camera of the at least one camera, the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system are determined according to the depth value of the pixel and the camera ray equation.

[0082] In particular, the first coordinate calculation steps S320, as in Fig. As shown in 6B, this includes: Performing the following steps for each pixel in the image captured by each camera of the at least one camera: S321', Acquiring third coordinates of the pixel in a third-plane coordinate system, wherein the third-plane coordinate system represents a pixel coordinate system of the image; S322', Calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinates of the pixel, where the normalized coordinates represent the coordinates of the projection point of the pixel on the normalized plane in the first spatial coordinate system; S323', Determining a camera beam equation of the pixel according to the normalized coordinates corresponding to the pixel, where the optical center of the camera is the origin of the first spatial coordinate system; S324', Calculating the normalized coordinates corresponding to the pixel according to the depth value of the pixel, in order to determine the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system.

[0083] Compared to Fig. 6A are the ones in Fig. The steps S321' to S322' shown in 6B are essentially the same as those in Fig. Steps S321 to S322 shown in 5A are not described in detail here. As described in S322, see below. Fig. As shown in Figure 4, the normalized coordinates (x0, y0, z0) of the projection point P" of pixel P in the normalized plane β are obtained, where x0=u1−cyfy,y0=v1−cyfy,z0=1. Therefore, as described in S323', knowledge of spatial analytic geometry can be used to calculate the equation of the line passing through O c and P" (or P or P') in the first spatial coordinate system, based on the coordinates (0, 0, 0) of the optical center O c and the normalized coordinates corresponding to the pixel (i.e., the normalized coordinates of the projection point P" of point P in the first spatial coordinate system (camera coordinate system)) (x0, y0, z0), i.e., the camera ray equation of pixel P: xx0=yy0=zz0

[0084] As in Fig. As shown in Figure 4, points P, P', and P" all lie on the camera beam, as described in S324'. Therefore, the first coordinate (x1, y1, z1) of the real scene position P", corresponding to point P in the first spatial coordinate system, can be calculated using equation (3) of the camera beam and the depth value h of pixel P (i.e., the distance between O and P"). In particular, since the length of the line segment OP" = h (i.e., the depth value of pixel P) and the first coordinates of P" (x1, y1, z1) satisfy the beam equation (2), the following equations can be obtained: {x12+y12+z12=h2x1x0=y1y0=z1z0

[0085] Solving the group of equations (4) yields the first coordinates of the real scene position P", which corresponds to point P in the first spatial coordinate system: x1=h⋅x0x02+y02+z02,y1=h⋅y0x02+y02+z02,z1=h⋅z0x02+y02+z02.

[0086] Fig. Figure 6C schematically illustrates an example process of the first coordinate calculation steps according to some embodiments of the present application.

[0087] According to the concept of the present application, before the intrinsic camera parameters and the pixel depth value are used to calculate the first coordinates of the image pixel point in the first spatial coordinate system (camera coordinate system), the image captured by the vehicle camera can first undergo distortion correction to convert the original image into a corrected image (to improve image quality). Subsequently, the first coordinates of the real scene position corresponding to the pixel in the corrected image are calculated in the first spatial coordinate system.

[0088] In general, image distortion refers to image disturbances caused by the characteristics of the camera lens or the imaging system. This distortion can manifest as inaccurate shapes, sizes, or positions of objects within the image. Image distortion primarily occurs in the following forms: First, radial distortion, caused by the shape of the lens, typically results in objects in the center of the image appearing normal, while objects at the edges may appear stretched or compressed. Barrel distortion and pincushion distortion are the two main forms of radial distortion. Second, tangential distortion occurs when the camera is not perfectly perpendicular to the surface of the subject, causing straight lines in the image to appear twisted or skewed.Distortion correction is a technique for improving image quality by removing or reducing the aforementioned image distortions or disturbances. Image distortion correction can be achieved through various methods, such as polynomial distortion correction, camera calibration (distortion coefficients), and specialized image processing software and algorithms, to effectively improve image quality and eliminate or mitigate the effects of distortion on the image.

[0089] As in Fig. As shown in Figure 6C, the first coordinate calculation steps S320, according to some embodiments of the present application, may include the following: S610, Correcting a distorted image taken by each camera of the at least one camera using an intrinsic camera parameter to obtain a corrected image, wherein the pixels in the image correspond one to one to the corrected pixels in the corrected image; S620, Acquiring for each pixel in the image captured by each camera of the at least one camera, of fourth coordinates of a rectified pixel corresponding to the pixel in a third-plane coordinate system, wherein the third-plane coordinate system represents a pixel coordinate system of the image; S630, Calculate for each pixel in the image captured by each camera of the at least one camera using the intrinsic camera parameter and the depth value of the pixel of first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system and the fourth coordinates of the corresponding rectified pixel.

[0090] As described in S610, to improve image quality and correct distortion in the initial two-dimensional image captured by the vehicle camera, the initial two-dimensional image can be distorted or rectified before determining the first (spatial) coordinates of the real scene corresponding to the image pixel points. This is done using the camera's intrinsic parameters (especially the distortion coefficients contained therein). This corrects the original image into a rectified image, ensuring that the original pixel set corresponds one-to-one to the rectified pixel set in the rectified image. The distortion correction process primarily involves using the distortion coefficient to correct the coordinates of the pixel points in the original image, so that the rectified image more closely reflects the actual situation.

[0091] As described in S620 to S630, after the distortion correction process, the first coordinates of the corresponding real-world scene position (the three-dimensional spatial coordinates in the camera coordinate system) can be calculated based on the pixel coordinates of the pixels in the corrected image (i.e., the fourth coordinates in the third-plane coordinate system), the depth value of the corresponding pixel in the original image, and the intrinsic camera parameter (e.g., the intrinsic parameter matrix). The calculation process for the first coordinates based on the corrected image is similar to the calculation process based on the original image (see the direct calculation method in S620 to S630). Fig. 6A or the radiation equation method in Fig. 6B) and will not be repeated here.

[0092] Fig. Figure 7 schematically illustrates an example process with distortion correction steps according to some embodiments of the present application.

[0093] Distortion coefficients are part of the camera's intrinsic parameters, describing the type and degree of distortion the camera lens can produce during the imaging process. Camera calibration methods use intrinsic camera parameters, such as calibrated distortion coefficients, to correct images and approximate their shape and proportions to those of the real world.

[0094] As in Fig. 7 shown, can be in Fig. Step S610 shown in 6C for distortion correction (for an image taken by each camera of the at least one camera, correcting the distortion of the image using the intrinsic parameters of the camera to obtain a corrected image) includes the following: S611, Determining the distortion coefficient and the intrinsic parameter matrix from the intrinsic parameters of the camera. S612, Determining a position mapping relationship of pixels of the image to rectified pixels based on the distortion coefficient and the intrinsic parameter matrix; S613, Mapping each pixel in the image to the position of a corresponding corrected pixel using the position mapping relation to create a corrected image.

[0095] First, as described in S611, calibrated intrinsic camera parameters are required during the distortion correction process, in particular the distortion coefficients and the intrinsic camera parameter matrix. The distortion coefficients include radial distortion coefficients (such as k1, k2, k3, etc.) and tangential distortion coefficients (such as p1, p2, etc.), as well as internal parameter matrices. K=[fx0cx0fycy001].

[0096] Secondly, the distortion correction process of the camera calibration method, as described in S612 to S613, actually proceeds as follows: First, the distortion coefficient and the intrinsic parameter matrix are used to calculate the exact position of each pixel in the original image in the undistorted state (i.e., the mapping relationship of the (corrected) pixel position from the original image to the corrected image, including two mapping matrices for the x-coordinate and the y-coordinate, respectively); then, each pixel in the original (distorted) image is shifted (i.e., mapped) to the corresponding (corrected pixel) position in the corrected image, producing a corrected image.

[0097] Fig. Figure 8 schematically illustrates an example process of second coordinate transformation steps according to some embodiments of the present application. As in Fig. As shown in Figure 8, the second coordinate conversion steps S330 can include the following: Performing the following steps for each pixel in the image captured by each camera of the at least one camera: S331, Converting the first coordinates of the real scene position corresponding to the pixel into homogeneous coordinates; S332, Creating an extrinsic parameter matrix using extrinsic camera parameters, wherein the extrinsic camera parameters comprise a rotation matrix representing a rotation relation from the first coordinate system to the second coordinate system, and a translation vector representing a translation relation from the first coordinate system to the second coordinate system; S333, Determining the second coordinates of the real scene position corresponding to the pixel in the second coordinate system, according to the extrinsic parameter matrix and the homogeneous coordinates of the real scene position corresponding to the pixel.

[0098] To convert camera coordinates into vehicle body coordinates, the three-dimensional camera coordinates must first be converted into four-dimensional homogeneous coordinates, as described in S331. As in Fig. As shown in Figure 4, for each pixel P in the pixel plane α, it is assumed that the corresponding real scene position P' has its first coordinates in the first spatial coordinate system (i.e., the camera coordinate system O). c -X c Y c Z c If (x1, y1, z1) are the coordinates, then the corresponding homogeneous coordinates are (x1, y1, z1, 1).

[0099] Secondly, the extrinsic camera parameters mainly comprise the rotation matrix R (3x3 matrix), which describes the rotation relationship from the camera coordinate system to the vehicle coordinate system, and the translation vector t (3x1 matrix or column vector), which represents the translation relationship from the origin of the camera coordinate system to the origin of the vehicle body coordinate system. In S332, the corresponding matrix of extrinsic parameters is: T=[Rt0T1], where R=[r11r12r13r21r22r23r31r32r33],t=[t0t1t2],0T=(0,0,0): Finally, as described in S333, the transformation of the first coordinates (camera coordinate system O) can be carried out based on formula (2). c -X c Y c Z c ) into the second coordinates (vehicle body coordinate system O w -X w Y w Z w ) the real scene position P', which corresponds to any pixel P in the Fig. 4 pixel plane α shown corresponds to the following formula (5): [x2y2z21]=[r11r12r13t0r21r22r23t1r31r32r33t20001][x1y1z11]

[0100] In other words, first the homogeneous coordinates of the second coordinates can be determined by multiplying the extrinsic parameter matrix T on the left-hand side with the homogeneous coordinates of the first coordinates, and then the second coordinates can be determined from the first three elements.

[0101] Fig. Figure 9 schematically illustrates an example process of the steps for creating a panoramic image according to some embodiments of the present application. As in Fig. 9 shown, can be found in Fig. The 3 steps shown in S340 for creating the panorama image include the following: S341, Reconstruct for each image taken by the at least one camera a local three-dimensional image corresponding to the image based on the image and the second coordinates of each pixel contained therein; S342, Creating a panoramic image for a vehicle surround view system based on the local three-dimensional images corresponding to the images captured by the at least one camera.

[0102] As mentioned above, creating a three-dimensional panoramic image of a vehicle involves several important steps, two of which are the most crucial: One is to convert the two-dimensional image into a local three-dimensional image using the pixel values ​​and the three-dimensional spatial coordinates of the corresponding real-world scene; the other is to integrate (seam and / or merge) the local three-dimensional images corresponding to each camera into the overall three-dimensional panoramic image of the vehicle. Since the three-dimensional spatial coordinates are captured based on the depth values ​​corresponding to the pixels, this creation method allows for the creation of a realistic, clear, and practical three-dimensional panoramic image of the vehicle.

[0103] Fig. Figure 10 schematically illustrates an example process of local three-dimensional image generation steps according to some embodiments of the present application. As in Fig. 10 shown, can the in Fig. Step 9 S341 (reconstruction of a local three-dimensional image corresponding to the image, based on the image and the second coordinates of each pixel therein, for each image captured by the at least one camera) includes the following: Performing the following steps for each image captured by the at least one camera: S341a, Generating point cloud data according to the second coordinates of each pixel in the image, wherein the point cloud is a set of discrete points in a three-dimensional space corresponding to each pixel; S341b, Extracting surface information of the real scene corresponding to the image from the point cloud data to create a local three-dimensional model; S341c, Performing surface optimization processing on the local three-dimensional model to obtain an optimized local three-dimensional model; S341d, Acquiring the pixel value of each pixel of the image and its mapping relationship to the point cloud data; S341e, based on the pixel value of each pixel of the image and its mapping relationship to the point cloud data, renders the surface of the optimized local three-dimensional model to obtain a local three-dimensional image corresponding to the image.

[0104] In some implementation examples, the process of creating a local three-dimensional image mainly comprises two subprocesses: surface shape reconstruction (as in S341a to S341c) and surface texture reconstruction (as in S341d to S341e). In the first subprocess (shape reconstruction), the object's surface information is extracted from the point cloud data, and a local three-dimensional model, i.e., the original surface shape, is created. This can be achieved using various algorithms such as Delaunay triangulation, Poisson surface reconstruction, spherical harmonics, etc. Secondly, the local three-dimensional model is optimized; that is, the constructed original surface shape is refined to eliminate noise, fill holes, smooth the surface, etc., to improve the model's quality and realism.In the second subprocess (texture reconstruction), the local three-dimensional model is rendered after thickness optimization using the mapping relationship between pixels and point cloud data to create an image or animation (local three-dimensional image) with realistic light, shadow, and texture effects.

[0105] In some embodiments of the present application, a vehicle surround view system is provided comprising: a plurality of cameras arranged at various locations on the vehicle body; and a processor configured to execute a method for creating a panoramic image for a vehicle surround view system according to some embodiments of the present application.

[0106] In the vehicle surround-view system according to some embodiments of the present application, the vehicle cameras, which are arranged at various locations on the vehicle body, can be different image acquisition devices. Generally, the multiple vehicle cameras in the vehicle surround-view system according to the present application can be arranged in four directions around the vehicle: front, rear, left, and right, or optionally installed in more directions to cover a larger area around the vehicle body. In the vehicle surround-view system according to some embodiments of the present application, the processor can comprise other logic components, which are implemented in hardware as a dedicated integrated circuit or formed using one or more semiconductors. A processor can consist of semiconductors and / or transistors (e.g., electronic integrated circuits (ICs)).In such a context, processor-executable instructions can be electronically executable instructions.

[0107] In a vehicle surround-view system according to some embodiments of the present application, the processor for images taken by several vehicle cameras arranged at different locations on the vehicle body can be configured to perform the various steps of a method for creating panoramic images for the vehicle surround-view system according to some embodiments of the present application based on these images and their depth information as well as the intrinsic and extrinsic parameters of each camera, in order to achieve a true and accurate reconstruction of the vehicle body's surroundings by the vehicle surround-view system and thereby avoid the problem of image distortion and the Manhattan effect.which are caused in related technologies by the mapping or projection of the panoramic image onto preset mapping boundaries (such as virtual walls and / or the ground) around the vehicle.

[0108] In particular, according to embodiments of the present application, the processes described above with reference to flowcharts can be implemented as a computer program. For example, one embodiment of the present application provides a computer program product comprising a computer program stored on a computer-readable medium, and the computer program contains program code for performing at least one step in a method embodiment of the present application. In the explanatory notes to this description, the terms “an embodiment”, “some embodiments”, “example”, “specific example”, or “some examples” mean that the specific features, structures, materials, or properties described in connection with the embodiment or example are included in at least one embodiment or example of the present application.In this description, the schematic expressions of the terms mentioned above do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or properties described may be combined in any suitable manner in one or more embodiments or examples. Moreover, provided there are no contradictions, those skilled in the field may combine and integrate the features of different embodiments or examples described herein.

[0109] Any process or procedure description presented in flowcharts or otherwise described herein may be understood as a representation of a module, fragment, or part of code comprising executable instructions containing one or more steps for implementing user-defined logic functions or processes. The scope of preferred embodiments of the present application includes alternative implementations in which functions need not necessarily be executed in the sequence shown or described (including substantially simultaneous execution according to the functions involved or in reverse order). This should be clear to those skilled in the art in the field to which the embodiments of the present application relate.

[0110] The logic and / or steps depicted in flowcharts or otherwise described herein may, for example, be considered a sequenced list of executable instructions for implementing the logical functions and may be embodied in any computer-readable medium for use by or in conjunction with an instruction execution system, apparatus, or device (such as a computer-based system, a system with a processor, or any other system capable of receiving and executing instructions from an instruction execution system, apparatus, or device). For the purposes of this specification, a "computer-readable medium" may be any device capable of containing, storing, communicating, distributing, or transporting the program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0111] It is understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the embodiments mentioned above, several steps or procedures can be implemented by software or firmware stored in memory and executed by a suitable instruction execution system. For example, if the implementation is carried out using hardware, it can be implemented using one of the following technologies known in the art, or a combination thereof: a discrete logic circuit with a logic gate circuit for implementing logical functions on data signals, a dedicated integrated circuit with a suitable combinational logic gate circuit, a programmable gate array, a field-programmable gate array, etc.

[0112] Experts in the field will recognize that all or some of the steps of the method of the above embodiments can be performed by hardware linked to program instructions, and that the program can be stored on a computer-readable storage medium which, when executed, includes the execution of one or a combination of the steps of the method embodiments.

[0113] Furthermore, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into one module. The integrated modules mentioned above can be implemented in the form of hardware or software functional modules. If the integrated module is implemented in the form of a software functional module and marketed or used as a standalone product, it can also be stored on a computer-readable storage medium.

Claims

[1] Method for creating a panoramic image for a vehicle surround view system, characterized by , that it includes the following: Retrieving depth information for an image captured by each camera of the at least one camera in the vehicle's surround-view system, wherein the depth information of the image includes a depth value for each pixel in the image; Calculate, for each pixel in an image captured by each camera of the at least one camera, using intrinsic camera parameters and a depth value of the pixel, first coordinates of a real scene position corresponding to the pixel in a first spatial coordinate system, wherein the first spatial coordinate system is a camera coordinate system corresponding to the camera; for each pixel in the image captured by each camera of the at least one camera, convert first coordinates of a real scene position corresponding to the pixel, using an extrinsic camera parameter, into second coordinates in a second spatial coordinate system, wherein the second spatial coordinate system is a vehicle body coordinate system corresponding to the vehicle's surround-view system; Creating a panoramic image for the vehicle's surround view system based on images taken by at least one camera and the second coordinates of the real scene position, corresponding to each pixel in each image. [2] Method according to claim 1, characterized by , that calculating first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system for each pixel in the image captured by each camera of the at least one camera, using an intrinsic camera parameter and a depth value of the pixel, includes the following: Calculate for each pixel in an image captured by each camera of the at least one camera using an intrinsic camera parameter a camera ray equation for the pixel, wherein the camera ray equation for the pixel is an equation in the first spatial coordinate system of a straight line passing through the optical center of the camera and the pixel; Determine for each pixel in the image, which is captured by each camera of the at least one camera, first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system, according to a depth value of the pixel and a camera ray equation. [3] The method of claim 2, wherein calculating the camera beam equation for each pixel in the image captured by each camera of the at least one camera using an intrinsic camera parameter comprises: performing the following steps for each pixel in the image captured by each camera of the at least one camera: Obtaining a third coordinate of the pixel in a coordinate system of a third plane, where the coordinate system of the third plane represents a pixel coordinate system of the image; Calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinate of the pixel, where the normalized coordinates represent the coordinates of the projection point of the pixel in the normalized plane in the first spatial coordinate system; determining a camera ray equation of the pixel according to the normalized coordinates corresponding to the pixel, where the optical center of the camera is the origin of the first spatial coordinate system. [4] Method according to claim 1, characterized by, that for each pixel in the image captured by each camera of the at least one camera, calculating the first coordinate of the real scene position corresponding to the pixel in the first spatial coordinate system, using the intrinsic camera parameter and the depth value of the pixel, comprises the following: Performing the following steps for each pixel in the image captured by each camera of the at least one camera: Obtaining a third coordinate of the pixel in a coordinate system of a third plane, where the coordinate system of the third plane represents a pixel coordinate system of the image; Calculating the normalized coordinates corresponding to the pixel using the intrinsic camera parameter and the third coordinates of the pixel, where the normalized coordinates represent the coordinates of the projection point of the pixel in the normalized plane in the first spatial coordinate system; Calculate the normalized coordinates corresponding to the pixel according to the depth value of the pixel, in order to determine the first coordinates of the real scene position corresponding to the pixel in the first spatial coordinate system. [5] The method of claim 1, wherein, for each pixel in the image captured by each camera of the at least one camera, the conversion of first coordinates of a real scene position corresponding to the pixel into second coordinates in a second spatial coordinate system using an extrinsic camera parameter comprises: performing the following steps for each pixel in the image captured by each camera of the at least one camera; converting the first coordinates of the real scene position corresponding to the pixel into homogeneous coordinates; Creating an extrinsic parameter matrix using extrinsic camera parameters, wherein the extrinsic camera parameters include a rotation matrix representing a rotation relation from the first coordinate system to the second coordinate system, and a translation vector representing a translation relation from the first coordinate system to the second coordinate system; Determining the second coordinates of the real scene position corresponding to the pixel in the second coordinate system, according to the extrinsic parameter matrix and the homogeneous coordinates of the real scene position corresponding to the pixel. [6] Method according to claim 1, wherein calculating first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system for each pixel in the image captured by each camera of the at least one camera, using an intrinsic camera parameter and a depth value of the pixel, comprises the following: De-distorting an image captured by each camera of the at least one camera using an intrinsic camera parameter to obtain a de-distorted image, wherein the pixels in the image correspond one-to-one to the de-distorted pixels in the de-distorted image; Capturing, for each pixel in the image captured by each camera of the at least one camera, fourth coordinates of a rectified pixel corresponding to the pixel in a third-plane coordinate system, wherein the third-plane coordinate system represents a pixel coordinate system of the image; Calculate for each pixel in the image captured by each camera of the at least one camera using the intrinsic camera parameter and the depth value of the pixel from first coordinates of a real scene position corresponding to the pixel in the first spatial coordinate system and the fourth coordinates of the corresponding rectified pixel. [7] Method according to claim 6, wherein the step of performing a rectification processing on the image recorded by each camera of the at least one camera using an intrinsic camera parameter to obtain the rectified image comprises the following: Determining the distortion coefficient and the intrinsic parameter matrix from the intrinsic camera parameters. Determining a position mapping relationship of pixels of the image to corrected pixels based on the distortion coefficient and the intrinsic parameter matrix; Map each pixel in the image to the position of a corresponding corrected pixel using the position mapping relationship to create a corrected image. [8] Method according to claim 1, characterized by , that the at least one camera comprises a monocular fisheye camera and the acquisition of depth information from an image captured by each camera of the at least one camera of the vehicle's surround view system comprises the following: Predictions of depth information for an image captured by each camera of the at least one camera of the vehicle surround view system using an image depth estimation model, wherein the image depth estimation model includes a trained convolutional neural network model. [9] Method according to claim 1, characterized by, that the acquisition of depth information from an image captured by each camera of the at least one camera of the vehicle's surround-view system includes obtaining depth information from the image captured by each camera by at least one of the following methods: Depth estimation using binocular or multi-camera images; Image depth estimation based on structured light; Image depth estimation based on LiDAR. [10] Method according to claim 1, wherein creating a panoramic image for the vehicle surround view system based on each image taken by the at least one camera and the second coordinates of the real scene position corresponding to each pixel in each image comprises: Reconstruct, for each image taken by the at least one camera, a local three-dimensional image corresponding to the image, based on the image and the second coordinates of each pixel contained therein. Creating a panoramic image for a vehicle surround view system based on the local three-dimensional images corresponding to the images captured by the at least one camera. [11] The method of claim 10, wherein, for each image captured by the at least one camera, the reconstruction of the local three-dimensional image corresponding to the image, based on the image and the second coordinates of each pixel therein, comprises: performing the following steps for each image captured by the at least one camera: Generating point cloud data according to the second coordinates of each pixel in the image, where the point cloud is a set of discrete points in a three-dimensional space corresponding to each pixel; Extracting surface information of the real-world scene corresponding to the image from the point cloud data to create a local three-dimensional model; Performing surface optimization processing on the local three-dimensional model to obtain an optimized local three-dimensional model; Capturing the pixel value of each pixel in the image and its mapping relationship to the point cloud data; Based on the pixel value of each pixel of the image and its mapping relationship to the point cloud data, the surface of the optimized local three-dimensional model is rendered to obtain a local three-dimensional image corresponding to the image. [12] Vehicle surround view system, comprising: several cameras distributed across the entire vehicle body, and a processor configured to perform a method according to any one of claims 1 to 11. [13] Computer program product comprising a computer program wherein, when the computer program is executed by a processor, the steps of a method according to any one of claims 1 to 11 are implemented.