Surround-view image generation method and apparatus, electronic device, and computer-readable storage medium
By performing 3D spatial transformation and constructing a projection ray array from multiple fisheye images, the problems of insufficient 3D surround view and projection distortion in existing panoramic parking imaging solutions are solved, achieving full-view coverage and visualization of three-dimensional obstacles, thus improving parking safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOTOVIS TECH SHANGHAI CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-06-02
AI Technical Summary
Existing panoramic parking imaging solutions cannot provide a comprehensive 3D surround view, suffer from projection distortion, and have a high failure rate in matching weak texture areas, failing to effectively display three-dimensional obstacles around the vehicle.
By performing three-dimensional spatial transformation on multiple fisheye images, a three-dimensional surface is constructed to determine the image information of the projection ray array, a panoramic image is constructed, and the three-dimensional surface is used to reconstruct and restore the real spatial geometry, eliminating stitching blind spots and projection distortion.
It achieves full-view coverage acquisition, eliminates stitching blind spots, visualizes three-dimensional obstacles, provides realistic surround view images, and improves parking safety.
Smart Images

Figure CN121329759B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to a method, apparatus, electronic device, and computer-readable storage medium for generating panoramic images. Background Technology
[0002] With the development of automotive intelligence, panoramic parking systems have become a key feature for improving parking safety. Currently, panoramic parking images are typically provided by a surround-view bird's-eye view stitching scheme. This scheme is based on the planar assumption of the IPM (Inverse Perspective Mapping) method, which cannot restore the real spatial geometry and can only provide a 2.5D top-down view. It cannot show three-dimensional obstacles around the vehicle and is prone to projection distortion, resulting in image distortion. Summary of the Invention
[0003] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for generating surround view images, in order to solve the problems of current panoramic parking image solutions that cannot provide a comprehensive 3D surround view and suffer from projection distortion.
[0004] In a first aspect, the present invention provides a method for generating panoramic images, the method comprising:
[0005] Acquire multiple fisheye images of the target vehicle, and perform three-dimensional spatial transformation on each of the multiple fisheye images to obtain the three-dimensional surface corresponding to the multiple fisheye images;
[0006] Based on the three-dimensional surface corresponding to the multi-channel fisheye images, the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined respectively.
[0007] Based on the image information corresponding to each projection ray in the projection ray array, a surround view image of the target vehicle is constructed, wherein the projection rays in the projection ray array correspond one-to-one with the pixels in the surround view image.
[0008] In one optional implementation, the multiple fisheye images are respectively subjected to three-dimensional spatial transformation to obtain three-dimensional surfaces corresponding to the multiple fisheye images, including:
[0009] The depth prediction model is input into the multi-channel fisheye images respectively to obtain the depth information of each pixel in the multi-channel fisheye images;
[0010] For each fisheye image, a 3D coordinate transformation is performed based on the depth information of each pixel in the fisheye image to construct the corresponding 3D surface.
[0011] In one optional implementation, a three-dimensional coordinate transformation is performed based on the depth information of each pixel in the fisheye image to construct a three-dimensional surface corresponding to the fisheye image, including:
[0012] Based on the depth information of each pixel in the fisheye image and the corresponding camera parameters, a three-dimensional coordinate transformation is performed on each pixel in the fisheye image to obtain the three-dimensional spatial coordinates of each pixel in the fisheye image.
[0013] Based on the three-dimensional spatial coordinates of each pixel in the fisheye image, adjacent pixels are connected to form a triangular network to obtain the three-dimensional surface corresponding to the fisheye image.
[0014] In one optional implementation, based on the three-dimensional surface corresponding to the multi-channel fisheye images, the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined, including:
[0015] Determine the intersection points of the projection rays between the three-dimensional curved surfaces corresponding to the multi-channel fisheye images and each projection ray in the projection ray array;
[0016] Calculate the projection position of the intersection of the projected rays in the fisheye image;
[0017] Based on the image information of the projection position, determine the image information corresponding to the projection ray.
[0018] In one optional implementation, calculating the projection position of the intersection of the projected rays in the fisheye image includes:
[0019] Based on the camera parameters of the fisheye image corresponding to the intersection of the projection rays, the intersection of the projection rays is projected onto the corresponding fisheye image to obtain the projection position of the intersection of the projection rays in the fisheye image.
[0020] In one optional implementation, determining the image information corresponding to the projection ray based on the image information of the projection position includes:
[0021] Bicubic interpolation is used to interpolate the image information at each projection position to obtain the image information corresponding to the intersection of each projection ray.
[0022] The weight of the projection ray intersection point relative to each fisheye image is determined based on the angle between the intersection point of the projection rays and the camera optical axis corresponding to each fisheye image.
[0023] Based on the weight of the intersection points of the projection rays relative to the fisheye image, the image information corresponding to the intersection points of the projection rays is weighted and calculated to obtain the image information corresponding to the projection rays.
[0024] In one alternative implementation, the projection ray array corresponding to the target vehicle is constructed as follows:
[0025] Using the center of the target vehicle as the center of the projection ray array, construct a spherical spatial coordinate system corresponding to the projection ray array;
[0026] Based on a preset angular resolution, a projection ray array is obtained by using a first preset angle as the projection coverage area in the first direction and a second preset angle as the projection coverage area in the second direction, and constructing projection rays in a spherical spatial coordinate system.
[0027] In a second aspect, the present invention provides a surround-view image generation apparatus, the apparatus comprising:
[0028] The 3D surface construction module is used to acquire multiple fisheye images of the target vehicle and perform 3D spatial transformation on the multiple fisheye images to obtain the 3D surface corresponding to the multiple fisheye images.
[0029] The image information determination module is used to determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle based on the three-dimensional curved surface corresponding to the multi-channel fisheye images.
[0030] The surround view image construction module is used to construct a surround view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array. The projection rays in the projection ray array correspond one-to-one with the pixels in the surround view image.
[0031] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the surround view image generation method of the first aspect or any corresponding embodiment described above.
[0032] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the surround view image generation method of the first aspect or any corresponding embodiment described above.
[0033] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the surround view image generation method of the first aspect or any corresponding embodiment described above.
[0034] The surround view image generation method provided in this invention performs three-dimensional spatial transformation on multiple fisheye images to obtain three-dimensional surfaces corresponding to the multiple fisheye images. Based on the three-dimensional surfaces corresponding to the multiple fisheye images, the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined. Based on the image information corresponding to each projection ray in the projection ray array, a surround view image of the target vehicle is constructed. Thus, on the one hand, full-view coverage acquisition is achieved through multiple fisheye images, and a surround view image is constructed using a projection ray array. The combination of the two eliminates stitching blind spots. On the other hand, the three-dimensional surface is used to reconstruct and restore the real spatial geometry, visualize three-dimensional obstacles, and eliminate projection distortion. Attached Figure Description
[0035] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0036] Figure 1 This is a schematic diagram of the first process of a surround view image generation method according to an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of a second process for generating surround view images according to an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram of the third process of the surround view image generation method according to an embodiment of the present invention;
[0039] Figure 4 This is a structural block diagram of a surround view image generation apparatus according to an embodiment of the present invention;
[0040] Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0043] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0044] With the development of automotive intelligence, panoramic parking systems have become a key feature for improving parking safety. Currently, panoramic parking images are typically provided by a surround-view bird's-eye view stitching scheme. This scheme is based on the planar assumption of the IPM (Inverse Perspective Mapping) method, which cannot restore the real spatial geometry and can only provide a 2.5D top-down view. It cannot show three-dimensional obstacles around the vehicle and is prone to projection distortion, resulting in image distortion.
[0045] In addition, there are other related technologies that use multi-view stereo vision to provide panoramic parking images. However, this solution requires additional configuration of a multi-camera synchronization module, and the matching failure rate increases significantly in weak texture areas, resulting in limited realism of the generated surround view images.
[0046] Based on this, embodiments of the present invention provide a method for generating surround-view images. This method involves performing three-dimensional spatial transformation on multiple fisheye images to obtain three-dimensional surfaces corresponding to the multiple fisheye images. Based on these three-dimensional surfaces, image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined. Based on the image information corresponding to each projection ray in the projection ray array, a surround-view image of the target vehicle is constructed. This achieves full-view coverage acquisition through multiple fisheye images and constructs a surround-view image using a projection ray array, eliminating stitching blind spots. Furthermore, it utilizes three-dimensional surfaces to reconstruct and restore the true spatial geometry, visualize three-dimensional obstacles, and eliminate projection distortion.
[0047] The surround-view image generation method provided in this invention can be applied to in-vehicle terminals, where the vehicle is equipped with multiple fisheye lenses, such as a four-channel fisheye lens, to achieve full-view coverage of the vehicle's surroundings. The in-vehicle terminal acquires the fisheye images captured by the multiple fisheye lenses in real time and generates surround-view images based on the surround-view image generation method of this invention, thereby providing users with real-time surround-view images of the vehicle's surroundings. Furthermore, it can be linked with other intelligent driving functions to ensure driving safety for users during driving or parking.
[0048] According to an embodiment of the present invention, a method for generating surround view images is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0049] This embodiment provides a method for generating surround-view images, which can be used in the aforementioned vehicle-mounted terminal. Figure 1 This is a first flowchart of a surround view image generation method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:
[0050] Step S101: Obtain multi-channel fisheye images of the target vehicle, and perform three-dimensional spatial transformation on the multi-channel fisheye images to obtain the three-dimensional surface corresponding to the multi-channel fisheye images.
[0051] In this embodiment of the invention, multiple fisheye images acquired by a multi-channel fisheye lens configured on a target vehicle are obtained, and each fisheye image is transformed into a three-dimensional spatial surface from a two-dimensional planar image, resulting in a three-dimensional surface corresponding to each fisheye image. Specifically, during the three-dimensional spatial transformation of the multiple fisheye images, for each pixel in each fisheye image, its position in the corresponding planar coordinate system (i.e., pixel coordinates) is used in conjunction with camera parameters to perform the three-dimensional spatial transformation, obtaining the spatial coordinates of each pixel. Then, for each fisheye image, a three-dimensional surface is constructed based on the spatial coordinates of the pixels it contains.
[0052] Step S102: Based on the three-dimensional surface corresponding to the multi-channel fisheye images, determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle.
[0053] In this embodiment of the invention, the projection ray array corresponding to the target vehicle corresponds to the image range of the surround view image of the target vehicle. The projection rays in the projection ray array adopt a uniform angular resolution distribution, thereby ensuring consistent imaging accuracy in each region of the surround view image. The intersection points of the projection rays between the three-dimensional curved surface corresponding to the multi-channel fisheye image and the projection rays in the projection ray array corresponding to the target vehicle are determined. The image information of the projection position corresponding to the projection ray intersection point in the fisheye image is used as the image information corresponding to the projection ray. The image information includes depth information and color information. The depth information is used to characterize the three-dimensional spatial position of the pixel or the object corresponding to the projection ray intersection point, and the color information is used to characterize the color of the pixel or the object corresponding to the projection ray intersection point.
[0054] In one optional implementation, the projection ray array corresponding to the target vehicle is constructed as follows: a spherical spatial coordinate system corresponding to the projection ray array is constructed with the center of the target vehicle as the center of the projection ray array; based on a preset angular resolution, a first preset angle is used as the projection coverage area in the first direction, and a second preset angle is used as the projection coverage area in the second direction, and projection rays are constructed in the spherical spatial coordinate system to obtain the projection ray array. The first direction refers to the horizontal direction corresponding to the target vehicle, with the plane where the target vehicle is located as the reference, and the positive direction of the first direction corresponds to the direction in which the target vehicle is traveling forward. The second direction refers to the vertical direction corresponding to the target vehicle, with the positive direction of the second direction corresponding to the direction away from the ground, and the vertical direction is based on a plane perpendicular to the plane where the target vehicle is located. The first preset angle is 360°, thus covering a 360° field of view in the horizontal direction, ensuring coverage of the entire vehicle. The second preset angle is 180°, thus covering a 180° field of view looking upwards in the vertical direction, but not covering a 180° field of view looking downwards in the vertical direction. The 180° field of view looking downwards in the vertical direction corresponds to the field of view below the vehicle, which is a field of view with no practical application value. Only covering the 180° field of view looking upwards in the vertical direction reduces the amount of data processing while meeting the requirements of vehicle surround view imaging.
[0055] In one optional implementation, the projection coverage area in the first direction and the projection coverage area in the second direction can be adjusted by the user. Specifically, the size of the first preset angle and the second preset angle and the angle relative to the spherical spatial coordinate system can be adjusted. For example, the second preset angle can be adjusted to 90°, and the second preset angle corresponds to the positive direction of the first direction and the second direction, that is, it corresponds to the direction of looking forward at 90°.
[0056] In one optional implementation, a preset angular resolution is set based on the required resolution of the surround-view image. Specifically, the resolution of the surround-view image is... Pixels pixels, then the default angular resolution is .
[0057] In one optional implementation, projection rays are constructed in a spherical spatial coordinate system. After obtaining the projection ray array, the vertex coordinates and ray direction corresponding to each projection ray are calculated to locate each projection ray. This facilitates the continuous construction of a surround-view image of the target vehicle based on the projection ray array. The vertex coordinates corresponding to a projection ray are the coordinates of the pixel point corresponding to that projection ray in the surround-view image. The ray direction is used to characterize the deviation angle of the projection ray relative to the spherical spatial coordinate system. It can be represented by the following formula (1):
[0058] Formula (1)
[0059] in, Let be the vertex coordinates of the pixel corresponding to the projected ray in the panoramic image along the first direction. Let be the vertex coordinates of the pixel corresponding to the projected ray in the panoramic image in the second direction. The pixel corresponding to the projection ray in the panoramic image is the pixel in the panoramic image along the first direction. 1 pixel The pixel corresponding to the projection ray in the panoramic image is the pixel in the second direction of the panoramic image. 1 pixel The resolution of the surround view image in the first direction. This represents the resolution of the panoramic image in the second direction.
[0060] The direction of the projected ray can be represented by the following formula (2):
[0061] Formula (2)
[0062] in, This represents the direction of the projected ray.
[0063] Step S103: Construct a surround view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array.
[0064] In this embodiment of the invention, the projection rays in the projection ray array correspond one-to-one with the pixels in the surround view image. The image information corresponding to each projection ray in the projection ray array is used as the image information of its pixel in the surround view image. Thus, the image information of the pixels in the surround view image is filled one by one and 3D rendered to obtain the surround view image of the target vehicle.
[0065] This embodiment provides a method for generating surround-view images, which can be used in the aforementioned vehicle-mounted terminal. Figure 2 This is a second flowchart of the surround view image generation method according to an embodiment of the present invention, as follows: Figure 2 As shown, the process includes the following steps:
[0066] Step S201: Obtain multi-channel fisheye images of the target vehicle, and perform three-dimensional spatial transformation on the multi-channel fisheye images to obtain the three-dimensional surface corresponding to the multi-channel fisheye images.
[0067] Specifically, step S201 includes:
[0068] Step S2011: Input the multi-channel fisheye images into the depth prediction model to obtain the depth information of each pixel in the multi-channel fisheye images.
[0069] In this embodiment of the invention, multiple fisheye images are input into a depth prediction model. The depth prediction model predicts and outputs the depth information of each pixel in the fisheye image, and the output of the depth prediction model is used as the depth information of the pixel. The depth prediction model can be a lightweight network such as the EfficientNet-B4 backbone. The input of the depth prediction model is the pixel coordinates and pixel value of each pixel in the fisheye image, and the output is the pixel coordinates and predicted depth information of each pixel in the fisheye image. When training the depth prediction model, sparse ground truth data generated by the LiDAR camera system and fisheye image-ground truth depth pairs generated by the simulation platform can be used as sample data.
[0070] Step S2012: For each fisheye image, perform three-dimensional coordinate transformation based on the depth information of each pixel in the fisheye image to construct the three-dimensional surface corresponding to the fisheye image.
[0071] In this embodiment of the invention, for each fisheye image, a three-dimensional coordinate transformation is performed based on the depth information of the pixels contained therein, in conjunction with its corresponding camera parameters, to construct a three-dimensional surface corresponding to the fisheye image. The camera parameters include intrinsic and extrinsic parameters. Intrinsic parameters include the camera's own imaging parameters, such as focal length, distortion parameters, and pixels. Extrinsic parameters include parameters of the camera's position and orientation in the world coordinate system, such as rotation and translation parameters.
[0072] In one optional implementation, for each fisheye image, the three-dimensional coordinates of each pixel in the fisheye image are transformed to obtain the three-dimensional point cloud data corresponding to the fisheye image. Then, a three-dimensional surface is constructed based on the three-dimensional point cloud data. Specifically, based on the depth information of each pixel in the fisheye image and the camera parameters corresponding to the fisheye image, the three-dimensional coordinates of each pixel in the fisheye image are transformed to obtain the three-dimensional spatial coordinates of each pixel in the fisheye image. Based on the three-dimensional spatial coordinates of each pixel in the fisheye image, adjacent pixels are connected to form a triangular network to obtain the three-dimensional surface corresponding to the fisheye image. The three-dimensional coordinate transformation of the pixel can be shown by the following formula (3):
[0073] Formula (3)
[0074] in, These are the coordinates of the pixel in the world coordinate system. That is, the three-dimensional coordinates of the pixel. For pixels in the camera The corresponding coordinates in the fisheye image For camera rotation parameters, For camera Translation parameters, For camera The intrinsic parameter matrix, This represents the depth value of a pixel.
[0075] Step S202: Based on the 3D surface corresponding to the multi-channel fisheye images, determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle. For details, please refer to [link to details]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0076] Step S203: Construct a surround-view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array. For details, please refer to [link to relevant documentation]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.
[0077] This embodiment provides a method for generating surround-view images, which can be used in the aforementioned vehicle-mounted terminal. Figure 3 This is a third flowchart of the surround view image generation method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:
[0078] Step S301: Acquire multi-channel fisheye images of the target vehicle, and perform 3D spatial transformation on each of the multi-channel fisheye images to obtain the corresponding 3D surfaces. For details, please refer to [link to relevant documentation]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0079] Step S302: Based on the three-dimensional surface corresponding to the multi-channel fisheye images, determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle.
[0080] Specifically, step S302 includes:
[0081] Step S3021: Determine the intersection points of the projection rays between the three-dimensional curved surfaces corresponding to the multi-channel fisheye images and each projection ray in the projection ray array.
[0082] In this embodiment of the invention, the intersection points of each projection ray in the projection ray array with the corresponding 3D surfaces of the multiple fisheye images are found. An accelerated algorithm can be used, such as the Möller-Trumbore algorithm based on the boundary volume hierarchy method. A spatial index is constructed for the 3D surface corresponding to each fisheye image to build a boundary volume hierarchy tree. The triangular meshes in the 3D surface are divided according to their spatial position, with each node corresponding to a bounding box. The top-level node encloses the entire 3D surface, and the bottom-level node corresponds to a single triangular mesh. Then, the boundary volume hierarchy tree corresponding to each fisheye image for each projection ray in the projection ray array is traversed starting from the top-level node. It is determined whether the projection ray intersects with the bounding box of the current node. If they do not intersect, all child nodes under that node are skipped. If they intersect, the child nodes are recursively traversed until the bottom-level node, finally obtaining the set of triangular meshes intersecting with the projection ray. For each triangular mesh in the set, the intersection point between the projection ray and the triangular mesh is calculated based on the Möller-Trumbore algorithm to obtain the projection ray intersection point. The specific calculation method is not elaborated here.
[0083] In one alternative implementation, if a projection ray intersects with a three-dimensional surface corresponding to multiple fisheye images or with multiple triangular meshes in a three-dimensional surface, the intersection point closest to the center of the sphere is taken as the projection ray intersection point, thereby retaining the intersection point closest to the vehicle and filtering out the intersection points farther away that are blocked by the closer intersection points, thus retaining only the intersection points with actual imaging significance.
[0084] Step S3022: Calculate the projection position of the intersection of the projection rays in the fisheye image.
[0085] In this embodiment of the invention, the intersection of the projection rays is back-projected onto the fisheye image to obtain the corresponding projection position of the intersection of the projection rays in the fisheye image. Specifically, based on the camera parameters of the fisheye image corresponding to the intersection of the projection rays, the intersection of the projection rays is projected onto the corresponding fisheye image to obtain the corresponding projection position of the intersection of the projection rays in the fisheye image, as shown in the following formula (4):
[0086] Formula (4)
[0087] in, This represents the projection position of the intersection of the projected rays in the fisheye image. The camera that represents the fisheye image. For camera The intrinsic parameter matrix, For camera rotation parameters, For camera Translation parameters, The function corresponds to the distortion intrinsic parameters calibrated for fisheye images.
[0088] Step S3023: Based on the image information of the projection position, determine the image information corresponding to the projection ray.
[0089] In this embodiment of the invention, the image information of the projection position is used as the image information corresponding to the projection ray.
[0090] In one optional implementation, the pixel coordinates of the projection position calculated in step S3022 may be non-integer. In this case, the image information of the pixel cannot be directly read from the fisheye image. Therefore, interpolation calculation can be performed on the image information of a preset number of pixels near the calculated projection position to obtain the image information corresponding to the intersection of the projection rays. Specifically, bicubic interpolation is used to interpolate the image information of each projection position to obtain the image information corresponding to each intersection of the projection rays.
[0091] In one optional implementation, within the overlapping area of the field of view of multiple fisheye images, the intersection point of the projection rays may be the intersection point of the same projection ray and the three-dimensional curved surface corresponding to the multiple fisheye images. The image information corresponding to the projection ray has the image information source of the multiple fisheye images. Therefore, the weight of the projection ray intersection point relative to each fisheye image can be set based on the angular deviation between the projection ray and the camera optical axis corresponding to each fisheye image. Specifically, the weight of the projection ray intersection point relative to each fisheye image is determined based on the angle between the projection ray intersection point and the camera optical axis corresponding to each fisheye image, as shown in the following formula (5):
[0092] Formula (5)
[0093] in, The camera that represents the fisheye image. For camera The corresponding weights of the fisheye image, The intersection of the projected rays corresponds to the camera. The corresponding fisheye image information, The intersection of the projected rays corresponds to the average value of the image information from the multiple fisheye images. Intersection of the projected rays and the camera The connection between the optical center and the camera The angle formed by the optical axis, For consistency weight, The viewpoint weights are defined as follows: The weight of the intersection of the projection rays relative to each fisheye image includes the weights of depth information and color information, which are calculated using the depth information and color information corresponding to the intersection of the projection rays, respectively. The consistency weight includes depth consistency weight and color consistency weight.
[0094] After calculating the weights of the projection ray intersections relative to each fisheye image, the image information corresponding to the projection ray intersections contained in the projection ray is weighted based on the weights of the projection ray intersections relative to each fisheye image to obtain the image information corresponding to the projection ray. Before performing the weighting calculation, the weights of the projection ray intersections relative to each fisheye image can be normalized.
[0095] Step S303: Construct a surround view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array.
[0096] The surround view image generation method provided in this invention performs three-dimensional spatial transformation on multiple fisheye images to obtain three-dimensional surfaces corresponding to the multiple fisheye images. Based on the three-dimensional surfaces corresponding to the multiple fisheye images, the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined. Based on the image information corresponding to each projection ray in the projection ray array, a surround view image of the target vehicle is constructed. Thus, on the one hand, full-view coverage acquisition is achieved through multiple fisheye images, and a surround view image is constructed using a projection ray array. The combination of the two eliminates stitching blind spots. On the other hand, the three-dimensional surface is used to reconstruct and restore the real spatial geometry, visualize three-dimensional obstacles, and eliminate projection distortion.
[0097] This embodiment also provides a surround-view image generation apparatus, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0098] This embodiment provides a surround view image generation device, such as... Figure 4 As shown, it includes:
[0099] The 3D surface construction module 401 is used to acquire multiple fisheye images of the target vehicle and perform 3D spatial transformation on the multiple fisheye images to obtain the 3D surface corresponding to the multiple fisheye images.
[0100] The image information determination module 402 is used to determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle based on the three-dimensional curved surface corresponding to the multi-channel fisheye images.
[0101] The surround view image construction module 403 is used to construct a surround view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array, wherein the projection ray in the projection ray array corresponds one-to-one with the pixel in the surround view image.
[0102] In one alternative implementation, the three-dimensional surface construction module 401 includes:
[0103] The depth information determination unit is used to input the multi-channel fisheye images into the depth prediction model to obtain the depth information of each pixel in the multi-channel fisheye images;
[0104] The 3D surface construction unit is used to perform 3D coordinate transformation based on the depth information of each pixel in the fisheye image for each channel of fisheye image, and construct the corresponding 3D surface of the fisheye image.
[0105] In one alternative implementation, the three-dimensional surface construction unit includes:
[0106] The spatial coordinate transformation subunit is used to perform three-dimensional coordinate transformation on each pixel in the fisheye image based on the depth information of each pixel in the fisheye image and the camera parameters corresponding to the fisheye image, so as to obtain the three-dimensional spatial coordinates of each pixel in the fisheye image.
[0107] The triangular network connection subunit is used to connect adjacent pixels into a triangular network based on the three-dimensional spatial coordinates of each pixel in the fisheye image, so as to obtain the three-dimensional surface corresponding to the fisheye image.
[0108] In one optional implementation, the image information determination module 402 includes:
[0109] The projection ray intersection point determination unit is used to determine the intersection points of the projection rays between the three-dimensional curved surface corresponding to the multi-channel fisheye image and each projection ray in the projection ray array;
[0110] The projection position calculation unit is used to calculate the projection position of the intersection of the projection rays in the fisheye image.
[0111] The image information determination unit is used to determine the image information corresponding to the projection ray based on the image information of the projection position.
[0112] In one optional implementation, the projection position calculation unit includes:
[0113] The intersection projection subunit is used to project the intersection of the projection rays onto the corresponding fisheye image based on the camera parameters of the fisheye image corresponding to the intersection of the projection rays, so as to obtain the projection position of the intersection of the projection rays in the fisheye image.
[0114] In one optional implementation, the base image information determination unit includes:
[0115] The interpolation calculation subunit is used to perform interpolation calculations on the image information of each projection position using bicubic interpolation to obtain the image information corresponding to the intersection of each projection ray.
[0116] The weight calculation subunit is used to determine the weight of the projection ray intersection point relative to each fisheye image based on the angle between the projection ray intersection point and the camera optical axis corresponding to each fisheye image.
[0117] The weighted calculation subunit is used to perform weighted calculation on the image information corresponding to the intersection points of the projection rays contained in the projection ray based on the weight of the intersection points of the projection rays relative to the fisheye image, so as to obtain the image information corresponding to the projection ray.
[0118] In one alternative embodiment, the device further includes:
[0119] The coordinate system construction module is used to construct a spherical spatial coordinate system corresponding to the projection ray array, with the center of the target vehicle as the sphere center of the projection ray array;
[0120] The projection ray construction module is used to construct projection rays in a spherical spatial coordinate system based on a preset angular resolution, with a first preset angle as the projection coverage area in the first direction and a second preset angle as the projection coverage area in the second direction, to obtain a projection ray array.
[0121] The surround view image generation apparatus provided in this embodiment of the invention can execute the surround view image generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0122] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0123] The following is a detailed reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0124] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0125] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the surround-view image generation method of the embodiments of the present invention.
[0126] Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0127] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the surround-view image generation method shown in the above embodiments is implemented.
[0128] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0129] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the scope defined by the invention.
Claims
1. A method for generating panoramic images, characterized in that, The method includes: Acquire multiple fisheye images of the target vehicle, and perform three-dimensional spatial transformation on each of the multiple fisheye images to obtain the three-dimensional surface corresponding to the multiple fisheye images; Based on the three-dimensional surface corresponding to the multiple fisheye images, the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle is determined respectively. Based on the image information corresponding to each projection ray in the projection ray array, a surround view image of the target vehicle is constructed, wherein the projection rays in the projection ray array correspond one-to-one with the pixels in the surround view image; The step of determining the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle based on the three-dimensional curved surface corresponding to the multiple fisheye images includes: Determine the intersection points of the three-dimensional curved surfaces corresponding to the multiple fisheye images and the projection rays between each projection ray in the projection ray array; Calculate the projection position of the intersection of the projected rays in the fisheye image; Based on the image information at the projection positions, the image information corresponding to the projection ray is determined. Specifically, bicubic interpolation is used to interpolate the image information at each projection position to obtain the image information corresponding to each projection ray intersection point. Based on the angle between the projection ray intersection point and the camera optical axis corresponding to each fisheye image, the weight of the projection ray intersection point relative to each fisheye image is determined. Based on the weight of the projection ray intersection point relative to each fisheye image, the image information corresponding to the projection ray intersection points included in the projection ray is weighted to obtain the image information corresponding to the projection ray. The weight of the projection ray intersection point relative to each fisheye image is shown in the following formula: ; in, The camera that represents the fisheye image. For camera The corresponding weights of the fisheye image, The intersection of the projected rays corresponds to the camera. The corresponding fisheye image information, The intersection of the projected rays corresponds to the average value of the image information from the multiple fisheye images. Intersection of the projected rays and the camera The connection between the optical center and the camera The angle formed by the optical axis, For consistency weight, For viewpoint weights.
2. The method according to claim 1, characterized in that, The step of performing three-dimensional spatial transformation on the multiple fisheye images to obtain the three-dimensional surfaces corresponding to the multiple fisheye images includes: The multiple fisheye images are input into the depth prediction model to obtain the depth information of each pixel in the multiple fisheye images. For each fisheye image, a three-dimensional coordinate transformation is performed based on the depth information of each pixel in the fisheye image to construct a three-dimensional surface corresponding to the fisheye image.
3. The method according to claim 2, characterized in that, The step of performing three-dimensional coordinate transformation based on the depth information of each pixel in the fisheye image to construct a three-dimensional surface corresponding to the fisheye image includes: Based on the depth information of each pixel in the fisheye image and the camera parameters corresponding to the fisheye image, a three-dimensional coordinate transformation is performed on each pixel in the fisheye image to obtain the three-dimensional spatial coordinates of each pixel in the fisheye image. Based on the three-dimensional spatial coordinates of each pixel in the fisheye image, adjacent pixels are connected to form a triangular network to obtain the three-dimensional surface corresponding to the fisheye image.
4. The method according to claim 1, characterized in that, The calculation of the projection position of the intersection of the projection rays in the fisheye image includes: Based on the camera parameters of the fisheye image corresponding to the intersection of the projection rays, the intersection of the projection rays is projected onto the corresponding fisheye image to obtain the projection position of the intersection of the projection rays in the fisheye image.
5. The method according to any one of claims 1-4, characterized in that, The projection ray array corresponding to the target vehicle is constructed in the following manner: Using the center of the target vehicle as the center of the projection ray array, a spherical spatial coordinate system corresponding to the projection ray array is constructed; Based on a preset angular resolution, a projection ray array is obtained by constructing projection rays in the spherical spatial coordinate system, using a first preset angle as the projection coverage area in the first direction and a second preset angle as the projection coverage area in the second direction.
6. A surround-view image generation device, characterized in that, The device includes: The three-dimensional surface construction module is used to acquire multiple fisheye images of the target vehicle and perform three-dimensional spatial transformation on the multiple fisheye images to obtain the three-dimensional surfaces corresponding to the multiple fisheye images. The image information determination module is used to determine the image information corresponding to each projection ray in the projection ray array corresponding to the target vehicle based on the three-dimensional curved surface corresponding to the multiple fisheye images; The surround view image construction module is used to construct a surround view image of the target vehicle based on the image information corresponding to each projection ray in the projection ray array, wherein the projection ray in the projection ray array corresponds one-to-one with the pixel in the surround view image; The image information determination module is used to determine the intersection points of the projection rays between the three-dimensional curved surfaces corresponding to the multiple fisheye images and each projection ray in the projection ray array; calculate the projection position of the projection ray intersection point in the fisheye image; and determine the image information corresponding to the projection ray based on the image information of the projection position. Specifically, bicubic interpolation is used to interpolate the image information at each projection position to obtain the image information corresponding to each projection ray intersection point. Based on the angle between the projection ray intersection point and the camera optical axis corresponding to each fisheye image, the weight of the projection ray intersection point relative to each fisheye image is determined. Based on the weight of the projection ray intersection point relative to each fisheye image, the image information corresponding to the projection ray intersection points included in the projection ray is weighted to obtain the image information corresponding to the projection ray. The weight of the projection ray intersection point relative to each fisheye image is shown in the following formula: ; in, The camera that represents the fisheye image. For camera The corresponding weights of the fisheye image, The intersection of the projected rays corresponds to the camera. The corresponding fisheye image information, The intersection of the projected rays corresponds to the average value of the image information from the multiple fisheye images. Intersection of the projected rays and the camera The connection between the optical center and the camera The angle formed by the optical axis, For consistency weight, For viewpoint weights.
7. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the surround view image generation method according to any one of claims 1 to 5 by executing the computer instructions.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the surround view image generation method according to any one of claims 1 to 5.