Image processing method, controller, device, storage medium and product

By acquiring multiple images and three-dimensional perception data of the target device's surroundings and stitching the images using the three-dimensional perception data, the distortion and warping problems in panoramic image stitching are resolved, high-quality panoramic images are generated, and driving safety is improved.

CN120634847APending Publication Date: 2025-09-12BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510642109.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing panoramic image stitching methods suffer from distortion and distortion, which leads to large visual errors and poor stitching effects.

Method used

By acquiring multiple images and 3D perception data of the target device's surroundings, the 3D perception data is used to stitch the images together, and feature matching and homography matrix calculation are used, combined with deep learning and mesh optimization technology to generate high-quality panoramic images.

Benefits of technology

It improves the accuracy and effect of panoramic image stitching, reduces geometric distortion, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634847A_ABST
    Figure CN120634847A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method, a controller, equipment, a storage medium and a product. The method comprises the following steps: acquiring a plurality of images and three-dimensional perception data collected for the surrounding environment of target equipment; and splicing the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. Therefore, by introducing the three-dimensional perception data collected for the surrounding environment of the target equipment, the multiple images collected for the surrounding environment of the target equipment are spliced, the panoramic image which is corresponding to the surrounding environment of the target vehicle and has a better splicing effect and is more accurate can be obtained, and the splicing effect of the panoramic image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to an image processing method, controller, device, storage medium and product. Background Art

[0002] An Around View Monitor (AVM) system is an in-vehicle driver assistance system designed to provide a panoramic view of the vehicle's surroundings using multi-camera stitching technology. This system is primarily used to help drivers gain a more intuitive understanding of their surroundings when parking, driving at low speeds, or maneuvering in complex environments, thereby improving driving safety.

[0003] During the research and practice of existing technologies, it was found that the panoramic image obtained by stitching images around the vehicle captured by multiple cameras using existing image processing methods is very prone to distortion or warping, resulting in large visual errors, making the panoramic image stitching effect poor. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, controller, device, storage medium and product, which can generate a panoramic image with better and more accurate stitching effect corresponding to the surrounding environment of the target vehicle, thereby improving the panoramic image stitching effect.

[0005] In order to achieve the above object, according to a first aspect of the present application, there is provided an image processing method, the method comprising:

[0006] Acquire multiple images and three-dimensional perception data of the surrounding environment of the target device;

[0007] The images are stitched based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0008] According to a second aspect of the present application, there is provided an image processing apparatus, the apparatus comprising:

[0009] An acquisition module, used to acquire multiple images and three-dimensional perception data of the surrounding environment of the target device;

[0010] A stitching module is used to stitch the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0011] According to a third aspect of the present application, a controller is provided, comprising a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to implement the image processing method provided in an embodiment of the present application.

[0012] According to a fourth aspect of the present application, a device is provided, wherein the vehicle includes the controller provided by the third aspect of the present application.

[0013] According to a fifth aspect of the present application, the device provided in the fourth aspect of the present application includes a vehicle.

[0014] According to a sixth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and the computer program is suitable for loading by a processor to execute the steps of any image processing method provided in the embodiments of the present application.

[0015] According to the seventh aspect of the present application, a computer program product is provided, which includes a computer program, and the computer program is stored in a computer-readable storage medium; when the processor of the controller reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the controller executes the steps in the image processing method provided in the embodiment of the present application.

[0016] In the image processing method, controller, device, storage medium, and product of the embodiments of the present application, multiple images of the target vehicle's surroundings and three-dimensional perception data are acquired; the images are stitched together based on the three-dimensional perception data to obtain a panoramic image corresponding to the surroundings. Thus, by incorporating the three-dimensional perception data acquired for the target vehicle's surroundings and stitching together the multiple images acquired for the target vehicle's surroundings, a more accurate and better stitched panoramic image corresponding to the target vehicle's surroundings can be obtained, thereby improving the panoramic image stitching effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 This is a schematic diagram of an implementation scenario of an image processing method provided in an embodiment of the present application;

[0019] Figure 2 This is a flowchart of an image processing method provided by an embodiment of the present application;

[0020] Figure 3a This is an image processing schematic diagram of an image processing method provided by an embodiment of the present application;

[0021] Figure 3bis another image processing schematic diagram of an image processing method provided by an embodiment of the present application;

[0022] Figure 4 This is a schematic diagram of the overall process of an image processing method provided by an embodiment of the present application;

[0023] Figure 5 is a structural diagram of an image processing device provided in an embodiment of the present application;

[0024] Figure 6 It is a schematic diagram of the structure of the controller provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0026] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0027] The embodiments of the present application provide an image processing method, apparatus, controller, device, storage medium, and product. The image processing apparatus can be integrated into a controller, which can be a server or a terminal.

[0028] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as basic cloud computing services such as big data and artificial intelligence platforms. Terminals may include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Terminals and servers can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.

[0029] The controller may be integrated into a device, and the device may include a vehicle, which may be a fuel vehicle, a plug-in hybrid vehicle, or a new energy vehicle, etc. This application does not make any specific limitations on this.

[0030] See also Figure 1 , taking the image processing device integrated into the controller as an example, Figure 1 A schematic diagram of an implementation scenario of the image processing method provided in an embodiment of the present application, wherein the controller can be integrated in a device or be communicatively connected to the device, the device can be a vehicle, and can obtain multiple images and three-dimensional perception data collected for the surrounding environment of the target device; the images are spliced ​​based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0031] It should be noted that Figure 1 The schematic diagram of the implementation environment scenario of the image processing method shown is merely an example. The implementation environment scenario of the image processing method described in the embodiments of this application is intended to more clearly illustrate the technical solutions of the embodiments of this application and does not constitute a limitation of the technical solutions provided by the embodiments of this application. Persons skilled in the art will appreciate that with the evolution of image processing and the emergence of new business scenarios, the technical solutions provided in this application will also be applicable to similar technical problems.

[0032] The solutions provided in the embodiments of the present application are specifically described by the following embodiments. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0033] This embodiment will be described from the perspective of an image processing apparatus, which may be integrated into a device.

[0034] See also Figure 2 , Figure 2 : is a flow chart of an image processing method provided in an embodiment of the present application. The image processing method includes:

[0035] Step S101: Acquire multiple images and three-dimensional perception data collected from the surrounding environment of the target device.

[0036] The target device may be a device that requires panoramic image stitching of its surroundings, and may include a vehicle, aircraft, ship, or other device. The image may be an image captured of the surroundings of the target device. For example, the target device may include multiple cameras for capturing the surroundings of the target device. For example, the camera may be a vehicle-mounted fisheye camera that can be used to capture fisheye images.

[0037] The 3D perception data can be 3D data collected from the surrounding environment of the target device, and can include 3D data such as point clouds and grids. For example, taking the target device as a vehicle, the 3D perception data can include multi-source 3D information collected by the intelligent driving vehicle, specifically including 3D sensors (including lidar, millimeter-wave radar, ultrasonic radar, etc.), as well as 3D information (including point clouds, occupancy grids, etc.) calculated by intelligent driving computing power based on multi-view cameras.

[0038] Optional, taking the target device as a vehicle as an example, please refer to Figure 3a , Figure 3a This is an image processing diagram of an image processing method provided in an embodiment of the present application. Multiple cameras can be configured around the vehicle. For example, the vehicle can be configured with a front camera, a rear camera, a left camera, and a right camera, so that multiple images of the vehicle's surrounding environment can be collected through these four cameras.

[0039] Step S102: stitching the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0040] The panoramic image may be an image including the surrounding environment of the target device, or may be an image obtained by stitching together multiple images collected of the surrounding environment of the target device.

[0041] Among them, there can be multiple ways to stitch images based on three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. For example, each image can be converted into a target coordinate system corresponding to the target device to obtain a first image; the first image can be stitched based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0042] The target coordinate system may be a coordinate system for stitching multiple images, for example, a cylindrical coordinate system, a spherical coordinate system, etc. The first image may be an image converted into the target coordinate system.

[0043] For example, if the target device is a vehicle, please refer to Figure 3b , Figure 3b This is another image processing schematic diagram of an image processing method provided in an embodiment of the present application. Taking the target coordinate system as a cylindrical coordinate system as an example, the front camera, rear camera, left camera and right camera of the vehicle can be used to capture the front image, rear image, left image and right image, so that these four images can be spliced ​​into the cylinder corresponding to the cylindrical coordinate system to realize the stitching of panoramic images of multiple images.

[0044] Among them, there are multiple ways to convert each image into the target coordinate system corresponding to the target device to obtain the first image. For example, the camera parameters of the camera that captures the image can be obtained; based on the camera parameters, each image is converted into the target coordinate system corresponding to the target device to obtain the first image.

[0045] The camera parameters may include intrinsic and extrinsic parameters of the camera.

[0046] Optionally, multiple cameras of the target device can be calibrated to obtain the intrinsic and extrinsic parameters of each camera to obtain camera parameters. The camera parameters can then be used to transform the multiple images captured by the multiple cameras into a cylindrical coordinate system and unfold them into a cylindrical plane to obtain a first image, which can form the initial reference for panoramic image stitching.

[0047] In one embodiment, using a fisheye camera and a cylindrical coordinate system as an example, a process for converting each image into the target coordinate system corresponding to the target device is used as a pre-processing step for panoramic stitching. This process aims to convert images captured by the fisheye camera into images captured by a virtual cylindrical camera. This conversion is linear in the vertical direction, thereby effectively preserving the geometric shapes of vertical objects in the target device's surroundings. However, due to the characteristics of geometric projection, a certain degree of distortion still exists in the horizontal direction. By performing cylindrical expansion on the fisheye camera images, the distortion in the common view area of ​​the cameras to be stitched can be partially corrected, thus providing a good foundation for subsequent feature point matching and panoramic image stitching.

[0048] Optionally, the process of converting each image into the target coordinate system corresponding to the target device can be designed as a reentrant mode, that is, it can support the cylindrical expansion operation of images captured by multiple cameras concurrently, thereby improving processing efficiency. The cylindrical expansion process of the fisheye image essentially realizes the conversion of the fisheye camera image to the virtual cylindrical camera image. In the implementation process, first, the origin of the vehicle coordinate system of the target device is used as the center of the bottom surface of the cylinder to construct a unified cylindrical coordinate system, so that multiple images captured by multiple cameras in the front, back, left and right can be uniformly projected onto the cylinder corresponding to the cylindrical coordinate system through coordinate transformation.

[0049] In one embodiment, a vehicle coordinate system can be used as an intermediate medium during the conversion of each image into a target coordinate system corresponding to the target device. First, each pixel in the captured image can be transformed into the vehicle coordinate system through inverse projection based on the intrinsic and extrinsic parameters of the camera. Then, through the projection process of the virtual cylindrical camera, the points in the vehicle coordinate system can be mapped to the cylindrical coordinate system corresponding to the virtual cylinder, thereby achieving the mapping of the camera image to the virtual cylindrical camera image. This step can effectively reduce the geometric distortion caused by perspective differences during the panoramic stitching process, provide high-quality initial input for subsequent stitching optimization, and thus improve the effect of the stitched panoramic image.

[0050] In one embodiment, please refer to Figure 4 , Figure 4 This is a schematic diagram of the overall flow of an image processing method provided in an embodiment of the present application. Matching points can be preset for the cylinder. When it is impossible to obtain a feature matching point pair that meets the requirements, the preset matching points can be directly superimposed to obtain the matching point pairs for calculating the homography matrix. At this time, there is no texture or depth feature information.

[0051] Among them, there can be multiple ways to stitch the first images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. For example, based on the three-dimensional perception data, corresponding feature matching point pairs can be extracted for adjacent first images; based on the feature matching point pairs, the first images can be stitched to obtain a panoramic image corresponding to the surrounding environment.

[0052] Among them, the feature matching point pair can be a pair composed of matching feature points extracted from adjacent first images respectively. For example, feature point 1 is extracted from first image 1, and feature point 2 is extracted from first image 2 and belongs to the same point in the surrounding environment as feature point 1. Then, feature point 1 and feature point 2 constitute a feature matching point pair.

[0053] The feature matching point pairs may include at least one of texture feature matching point pairs and depth matching point pairs.

[0054] The texture feature matching point pair may be feature points extracted based on texture information in the first image, and the depth matching point pair may be feature points extracted based on three-dimensional perception data.

[0055] Optionally, when the feature matching point pairs include depth matching point pairs, the step of extracting the depth matching point pairs may include: projecting the three-dimensional perception data onto the image plane corresponding to the first image; and determining multiple depth matching point pairs of the three-dimensional perception data in the adjacent first image based on the projection result.

[0056] The image plane may be the image plane where the first image is located, or may be the camera plane corresponding to the camera that captures the first image.

[0057] Optionally, in the case where the feature matching point pairs include texture feature matching point pairs, the step of extracting the texture feature matching point pairs may include: extracting texture feature points in adjacent first images to obtain a plurality of texture feature matching point pairs.

[0058] Among them, there are many ways to extract texture feature points in adjacent first images and obtain multiple texture feature matching point pairs. For example, initial texture feature matching point pairs can be extracted in adjacent first images; the initial texture feature matching point pairs can be filtered to obtain multiple texture feature matching point pairs.

[0059] The initial texture feature matching point pair may be a feature point pair extracted from an adjacent first image, and filtering has not yet been performed.

[0060] Among them, there can be many ways to filter the initial texture feature matching point pairs to obtain multiple texture feature matching point pairs. For example, based on the distribution of the initial texture feature matching point pairs in the first image to which they belong, the abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs can be determined; the abnormal initial texture feature matching point pairs can be removed from the initial texture feature matching point pairs to obtain multiple texture feature matching point pairs.

[0061] Among them, the abnormal initial texture feature matching point pair may be an initial texture feature matching point pair that does not meet the requirements, for example, it may be an incorrect noise point that interferes with the accuracy of the feature point matching result. In this way, by eliminating the abnormal initial texture feature matching point pair, the accuracy of the feature matching point pair extraction can be improved, thereby improving the panoramic image stitching effect.

[0062] Among them, there are many ways to determine the abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs based on the distribution of the initial texture feature matching point pairs in the first image to which they belong. For example, the adjacent first images can be grid-divided separately to obtain multiple first grid areas of each first image; based on the initial texture feature matching points distributed in the matching first grid areas in the adjacent first images, the abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs can be determined.

[0063] The specific division of the grid area can be set according to actual needs, and the first grid area can be a grid area obtained by grid dividing the first image.

[0064] Among them, there can be multiple ways to determine the abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs based on the initial texture feature matching points distributed in the first grid area matched in the adjacent first image. For example, abnormal initial texture feature matching points can be determined from the initial texture feature matching points distributed in the first grid area matched in the adjacent first image; and abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs can be determined based on the abnormal initial texture feature matching points.

[0065] Among them, there can be multiple ways to determine abnormal initial texture feature matching points from the initial texture feature matching points distributed in the matched first grid areas in the adjacent first images. For example, for each group of matched first grid areas in the adjacent first images, based on the number of initial texture feature matching points in one first grid area and another first grid area, the matching point distribution similarity information corresponding to the first grid area can be determined; based on the matching point distribution similarity information, the abnormal initial texture feature matching points can be determined.

[0066] The matching point distribution similarity information may be information describing the distribution similarity of initial texture feature matching points distributed in a group of first grid areas matched in adjacent first images.

[0067] Optionally, when the matching point distribution similarity information is not within a preset threshold range, the initial texture feature matching point in the first grid area is determined to be an abnormal initial texture feature matching point.

[0068] Among them, the specific numerical range of the preset threshold range can be set according to actual conditions, for example, it can be a numerical range such as (0.8, 1.2), which is not limited in the embodiment of the present application. When the matching point distribution similarity information is close to 1, it can indicate that the initial texture feature matching points distributed in a group of first grid areas are similar, indicating that the initial texture feature matching points distributed in this group of first grid areas are reliable. If the matching point distribution similarity information is far from 1, it can indicate that the gap between the initial texture feature matching points distributed in the group of first grid areas is large, indicating that the initial texture feature matching points distributed in this group of first grid areas are unreliable and can be eliminated as abnormal initial texture feature matching point pairs.

[0069] In a specific implementation, the feature matching point pairs of two adjacent views with better robustness and accuracy can be obtained by constructing texture descriptions and mapping depth values ​​in multiple views. Figure 4, where texture feature matching point pairs can be extracted using traditional feature detection algorithms or deep learning methods. For example, the Scale-Invariant Feature Transform (SIFT) algorithm can be used to extract key points and feature descriptors of images I1 and I2, which can be expressed as:

[0070] p i =(x i ,y i ), q i =(x′ i , y′ i )

[0071] Among them, p i =(x i ,y i ) can represent the coordinates of the i-th key point in image I1, q i =(x′ i ,y′ i ) represents the coordinates of the i-th key point in image I2, which is p i The corresponding matching points in image I2.

[0072] Then, the nearest neighbor search matching feature descriptor can be used to obtain a set of texture feature matching point pairs, which can be expressed as:

[0073]

[0074] Optionally, a deep learning model (such as SuperGlue) can be used to further enhance the robustness of the extracted texture feature matching point pairs by learning global context information.

[0075] In order to ensure the accuracy of the matching point pairs in the set of texture feature matching point pairs, unreliable outliers can be removed. For example, the two adjacent images can be grid-divided by the fast and robust feature matching filtering algorithm based on motion statistics (Grid-based Motion Statistics, referred to as GMS), and the distribution consistency of the initial texture feature matching point pairs in each grid can be counted. Specifically, for the set of initial texture feature matching point pairs P initial , divide the corresponding first image grid into m×n grid areas, and count the number of initial texture feature matching points N contained in each grid i ,Then, the distribution similarity of the initial texture feature matching points in the grid area can be used as the criterion for abnormal initial texture feature matching points, which can be specifically expressed as:

[0076]

[0077] in, It can be expressed as the number of initial texture feature matching points contained in the i-th grid area of ​​the first image 1, It can be expressed as the number of initial texture feature matching points contained in the i-th grid area of ​​the first image 2, R i It can be expressed as the matching point distribution similarity information. If R i Within the preset threshold range, the initial texture feature matching points contained in the grid area can be determined as initial texture feature matching points that meet the requirements, that is, internal points, otherwise they are abnormal initial texture feature matching points, that is, external points. In this way, we can eliminate external points efficiently without relying on global model assumptions, while retaining more internal points and improving the extraction accuracy of texture feature matching point pairs.

[0078] Optionally, for depth matching point pairs, the coordinates X of data points such as 3D point clouds in 3D perception data can be used. i =(X i ,Y i ,Z i ), the camera intrinsic parameter matrix K and extrinsic parameter matrix [R|t] of the target vehicle's multi-channel camera are projected onto the image plane to obtain the pixel coordinates of each data point in the 3D perception data:

[0079] p i =K·[R|t]·X i .

[0080] Then, according to the depth value corresponding to each data point, it can be mapped to the corresponding pixel point through bilinear interpolation, so that the point projected onto the image plane can be obtained. Since a data point in the three-dimensional perception data can be projected into two images, a data point can be projected onto the image plane to obtain two corresponding matching points, that is, a depth matching point pair. In this way, a depth matching point set can be obtained:

[0081]

[0082] in, It can be a depth matching point set, that is, a set consisting of depth matching point pairs.

[0083] Optionally, there are multiple ways to stitch the first images based on feature matching point pairs to obtain a panoramic image corresponding to the surrounding environment. For example, the first transformation parameters between image planes corresponding to adjacent first images can be calculated based on the feature matching point pairs; and the first images can be stitched based on the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

[0084] The first transformation parameter may be a parameter for stitching the first images, for example, a homography matrix corresponding to the planes where the two adjacent first images are located. The homography matrix may be a 3×3 matrix that can be used to describe the perspective transformation relationship between the two planes.

[0085] Among them, there can be multiple ways to calculate the first transformation parameters between image planes corresponding to adjacent first images based on feature matching point pairs. For example, the homography matrix between image planes corresponding to adjacent first images can be calculated based on the coordinate information of the feature matching point pairs to obtain the first transformation parameters.

[0086] The coordinate information may be the coordinate position of the feature point in the feature matching point pair in each first image.

[0087] Specifically, since the two matching feature matching point pairs satisfy the homography matrix transformation: q i =Hp i , where H is a 3×3 matrix. To determine the value of H, the following linear equations can be constructed for each pair of feature matching points:

[0088] x′ i (h7x i +h8y i +h9)-(h1x i +h2y i +h3)=0

[0089] y′ i (h7x i +h8y i +h9)-(h4x i +h5y i +h6)=0

[0090] Where h1-h9 are the elements in the homography matrix corresponding to the image planes of the two adjacent first images. For n pairs of feature matching points, the matrix equation A·H=0 can be constructed according to the Direct Linear Transformation (DLT) algorithm, where A can be expressed as:

[0091]

[0092] Then, A can be solved by the Singular Value Decomposition (SVD) method:

[0093] A=U·Σ·V T

[0094] Among them, H is the eigenvector corresponding to the smallest singular value, which is reshaped into a 3×3 matrix to obtain the global homography matrix, that is, the first transformation parameter. U can be expressed as an m×m orthogonal matrix, and its column vectors are called left singular vectors. These vectors constitute the orthogonal basis of the column space of A. Σ can be expressed as an m×n diagonal matrix, and the elements on the diagonal are non-negative singular values, and are arranged in order from large to small. V can be expressed as an n×n orthogonal matrix, and its column vectors are called right singular vectors. These vectors constitute the orthogonal basis of the row space of A. Geometrically, V represents the rotation of the original matrix A on its row space. V T It can be expressed as the transposed matrix of V, which is an n×n matrix.

[0095] In this way, the quality of matching point pairs is significantly improved by removing noise points through the GMS algorithm and combining them with depth matching point pairs. Moreover, more accurate matching point pairs can be retained by removing noise points to enhance the description ability of global geometric relationships. The DLT algorithm and SVD decomposition algorithm based on matching points further ensure the robustness of the global homography matrix, laying a good initial foundation for subsequent mesh optimization, thereby improving the accuracy and stitching effect of panoramic image stitching.

[0096] Optionally, there are multiple ways to stitch the first images together according to the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment. For example, the geometric surface corresponding to the target coordinate system can be gridded to obtain multiple second grid areas; the second transformation parameters corresponding to each second grid area in the first image can be calculated according to the first transformation parameters and three-dimensional perception data; and the first images can be stitched together based on the second transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

[0097] The second grid area may be a plurality of grid areas obtained by meshing a geometric surface corresponding to the target coordinate system. The second transformation parameter may be a transformation parameter corresponding to each set of second grid areas matched on adjacent first images, and may be a local transformation parameter relative to the first transformation parameter. Optionally, the second transformation parameter may also be a homography matrix.

[0098] Optionally, the target coordinate system may include a cylindrical coordinate system, and the geometric surface may include a cylindrical surface.

[0099] Among them, there can be multiple ways to calculate the second transformation parameters corresponding to each second grid area in the first image based on the first transformation parameters and the three-dimensional perception data. For example, the three-dimensional perception data can be mapped to the geometric surface corresponding to the target coordinate system to obtain the first position information of each data point in the three-dimensional perception data in the geometric surface; a three-dimensional model can be constructed based on the three-dimensional perception data to obtain the second position information of each data point in the three-dimensional perception data in the three-dimensional model; based on the first position information and the second position information, the second transformation parameters corresponding to the second grid area can be determined.

[0100] The 3D perception data may include multiple 3D data points, for example, multiple point clouds. The first position information may be the position of each data point in the 3D perception data on a geometric surface, and the second position information may be the position of each data point in the 3D perception data on a 3D model surface. Optionally, the 3D model may be a 3D surface model.

[0101] Among them, there can be multiple ways to determine the second transformation parameters corresponding to the second grid area based on the first position information and the second position information. For example, the initial second transformation parameters corresponding to the second grid area can be determined based on the first position information and the second position information; based on the initial second transformation parameters and the first transformation parameters, the second transformation parameters corresponding to the second grid area can be calculated.

[0102] There are multiple ways to calculate the second transformation parameters corresponding to the second grid area based on the initial second transformation parameters and the first transformation parameters. For example, a correction matrix can be calculated based on the initial second transformation parameters and the first transformation parameters; and the second transformation parameters corresponding to the second grid area can be calculated based on the correction matrix.

[0103] The correction matrix may be a parameter used to correct the initial second transformation parameters.

[0104] Among them, there can be multiple ways to calculate the second transformation parameters corresponding to the second grid area based on the correction matrix. For example, for each second grid area, the weight information corresponding to the feature matching point pair can be calculated based on the coordinate information of the grid center point and the feature matching point pair in the second grid area; based on the weight information and the correction matrix, the second transformation parameters corresponding to the second grid area can be determined.

[0105] The grid center point may be the center point of the second grid area, and the weight information may be information representing the degree of influence of the feature matching point on the current grid area.

[0106] In one embodiment, by mapping the three-dimensional perception data to the cylindrical surface corresponding to the cylindrical coordinate system, combined with the feature point matching results, an initial estimate is provided for the local homography matrix of the grid area. First, the surface of the three-dimensional perception data can be reconstructed. Based on the input three-dimensional perception data, a surface interpolation or reconstruction algorithm is used to generate a three-dimensional model, calibrate the spatial position (X, Y, Z) of each point in the three-dimensional model, and mesh the cylindrical surface to obtain multiple grids, i.e., the second grid area. Then, the vertices of the cylindrical mesh can be mapped to the surface of the three-dimensional model through cylindrical parameterization, and the local normal vector n = (n x ,n y ,n z ). Then, the vertex offset correction value d = n·(pq) can be calculated using the point-to-surface distance formula. Where p is the cylindrical point coordinate, q is the point coordinate on the 3D model surface, and n is the normal vector. Based on the normal vector offset and the geometric characteristics of the mapping point, as well as the first transformation parameters, the initial second transformation parameters can be determined, which can be expressed as:

[0107] H local =H0+ΔH

[0108] Among them, H local It can be expressed as the initial second transformation parameter, H0 can be the first transformation parameter, and ΔH is the matrix corrected based on the local characteristics, that is, the correction matrix.

[0109] Then, the preliminary optimization process of the homography matrix corresponding to the second grid area can be combined with feature point matching and 3D perception data weights to improve the local stitching accuracy through iterative optimization. For example, the weight information can be calculated based on feature matching, and the grid center point (x c ,y c ) and the source image matching inliers (x i ,y i ), calculate the Euclidean distance weight, that is, the weight information, which can be expressed as:

[0110]

[0111] Among them, the weight information w i The larger the value, the greater the influence of the feature matching point on the current grid area. The weight information is introduced into the DLT algorithm to construct the weighted A matrix, that is, the weighted matrix A w , which can be expressed as:

[0112]

[0113] Among them, a 11 -a n3 It can be the original matrix element in the A matrix.

[0114] Then, the weighted matrix can be decomposed by SVD to calculate the modified local homography matrix, that is, the second transformation parameter, which can be expressed as:

[0115] H optimized =argmin||A w ·h||

[0116] Among them, H optimized It can be expressed as a second transformation parameter, and h can be expressed as a correction matrix, for example, it can be the result of summing, concatenating or multiplying the correction matrices of all grid areas.

[0117] In this way, after the above optimization of the initial second transformation parameters corresponding to all grid areas, the homography matrix of each grid area is guaranteed to have higher accuracy in the local area, providing a reliable foundation for the subsequent global optimization of the manifold space.

[0118] Optionally, before stitching the first image based on the second transformation parameters to obtain a panoramic image corresponding to the surrounding environment, the second transformation parameters can also be optimized to obtain optimized second transformation parameters to minimize the coordination error between all grid areas in the first image, improve the stitching continuity and overall consistency between the grid areas in the first image during stitching, and thus improve the panoramic stitching effect.

[0119] Among them, there are many ways to optimize the second transformation parameters and obtain the optimized second transformation parameters. For example, the second transformation parameters can be converted from the current geometric space to the manifold space for optimization processing; the result after optimization processing is converted back to the current geometric space to obtain the optimized second transformation parameters.

[0120] Among them, the current geometric space can be the geometric space in which the second transformation parameter is calculated, for example, it can be Euclidean space, and the popular space can be a structure in mathematics that describes a locally seemingly flat high-dimensional space. It is a more general geometric space used to process curves and surfaces.

[0121] In one embodiment, the second transformation parameters can be globally optimized by mapping to the manifold space. The global optimization of the manifold space can handle the coupling relationship between local regions and ensure the continuity and overall consistency of the image stitching result. Specifically, the second transformation parameters of each grid region, that is, the local homography matrix H i , through the mapping function of the manifold space Embed it into the manifold space, which can be expressed as:

[0122] M i =f(H i )

[0123] Wherein, M can represent a specific manifold where the homography matrix is ​​located, for example, it can include the special Euclidean group "SE(3)", the affine transformation manifold, etc. i It can be expressed as the local homography matrix H i For global optimization of the manifold space, the compatibility error between all grid regions can be minimized by constructing a global objective function. For example, it can be expressed as:

[0124]

[0125] Among them, (i, j)∈E represents the adjacency relationship of the grid area, ω ij is the weight coefficient, which is used to adjust the influence of the error between different matrix pairs. i and M j are two local homography matrices in the set M, P ij is the adjacency transformation matrix, ||·||F represents the F-norm (Frobenius norm), which is a metric for matrices. It is equivalent to squaring each element in the matrix, summing the results, and taking the square root to measure the overall size or change of the matrix.

[0126] Optionally, an algorithm based on gradient descent or Riemannian optimization can be used to perform iterative updates in the manifold space to ensure that the solution is always on the manifold. The Riemannian optimization algorithm, also known as the Riemannian gradient optimization algorithm, is an optimization method performed in the manifold space. It extends the optimization method of the Euclidean space and makes it applicable to spaces with nonlinear geometric structures. Specifically, it can be expressed as:

[0127]

[0128] Where exp(·) is the exponential mapping on the manifold, η is the learning rate, Indicates that the matrix M in the kth iteration i The gradient of the energy function (E) at , where the gradient is the gradient on the Riemannian manifold. For the optimized manifold space matrix M i , through the reverse mapping function g:M→R 3×3 , converted back to Euclidean space, it can be expressed as:

[0129]

[0130] Among them, H i final It can be expressed as the optimized local homography matrix M iThe result after converting the local homography matrix (i.e., the second transformation parameter) back to the Euclidean space is the optimized second transformation parameter. This conversion ensures that all local homography matrices are continuous and consistent in the Euclidean space. The final calculated local homography matrix, i.e., the optimized second transformation parameter, is used to transform and fuse the input images to generate a panoramic stitching result. The obtained panoramic image with a high stitching effect can be expressed as:

[0131]

[0132] Among them, I final (x, y) can be expressed as a stitched panoramic image, W i (x, y) can be a weighted fusion function used to transform and fuse the input image. (x, y) can represent the coordinates of the pixel points in the image to generate a panoramic stitching result. The specific weighted fusion function can be determined according to the actual image processing process required. The embodiment of the present application is not limited here. i is the i-th input image, for example, the i-th first input image, n can be the total number of input images, for example, the total number of first images, or the number of cameras on the target vehicle for capturing images of the vehicle's surroundings, ∑ can be represented by a summation symbol, and T can be represented by a transpose symbol.

[0133] Optionally, the panoramic image can be projected onto the surface of a three-dimensional model constructed based on the three-dimensional perception data to obtain a target three-dimensional model; the target three-dimensional model is rendered to obtain a panoramic effect image corresponding to the surrounding environment.

[0134] The target three-dimensional model may be a textured three-dimensional model.

[0135] In this way, the generated panoramic image I final Projected onto the surface of the 3D model. For each point (X, Y, Z) on the 3D model surface, its coordinates (u, v) in the corresponding texture image can be calculated based on the cylindrical projection relationship. Specifically, it can be expressed as:

[0136]

[0137] Here, R is the radius of the cylinder in the cylindrical coordinate system, and h is the height of the cylinder. Mapping the color information of the panoramic image to corresponding points on the 3D model surface generates a textured target 3D model. This textured target 3D model can then be rendered by a rendering engine to generate and output the final panoramic stitching effect.

[0138] In the existing panoramic image stitching methods, the physical distance between multiple fisheye cameras is large. Due to the different observation angles, the positions of points in the same scene in the two images are greatly different. In addition, when the vehicle is driving, the structures of objects in the open scene are very different. At the same time, there are rigid buildings and non-rigid pedestrians, and static and dynamic obstacles coexist. As a result, in the application context of vehicle-mounted surround view images, the existing panoramic stitching technology solutions have the following problems: (1) Feature matching errors caused by geometric mismatch. When the parallax is large, the positions of objects in the scene in the two images change significantly, making it difficult to accurately find the corresponding points between the images based on texture alone, resulting in a situation where the feature point matching is a matter of a thousand miles; (2) The large parallax and small overlapping area make it easy to have a large mismatch in weak texture scenes; (3) The mismatch is more serious in the near field due to the lack of depth information. Ignoring depth information during stitching can lead to logical errors in image fusion, such as misalignment of foreground and background. (4) Due to the limitations of the projection model, the image distortion is large. During the image stitching process, it is usually necessary to project the image into a unified coordinate system (such as a top-down projection or a spherical projection). When the parallax is large, the unified projection model may not be able to take into account the differences in all perspectives, and areas with excessive distortion will result in unnatural stitching effects.

[0139] To this end, an embodiment of the present application provides an image processing method for panoramic stitching based on multi-channel vehicle-mounted cameras and three-dimensional perception data. The camera parameters are obtained through camera calibration, and the multiple images captured by the camera are expanded into a cylindrical plane. The global homography matrix and the local homography matrix are calculated by combining the texture feature matching point pairs and the depth matching point pairs. Then, the local homography matrix is ​​optimized based on the grid points, and the global homography matrix is ​​optimized in the manifold space to ensure the continuity and consistency of the homography matrix. Finally, the optimized panoramic images are fused and projected onto the three-dimensional model to achieve a high-precision, low-distortion panoramic stitching effect while improving robustness.

[0140] Specifically, the embodiment of the present application adds the input of three-dimensional perception data to the traditional panoramic stitching task. The existing common on-board AVM panoramic stitching solutions often directly apply static images for stitching. When performing panoramic image stitching, the embodiment of the present application takes into account the existence of multi-source three-dimensional information of intelligent driving vehicles, specifically including three-dimensional sensors (including lidar, millimeter-wave radar, ultrasonic radar, etc.), and three-dimensional information obtained by intelligent driving computing power based on multi-view cameras (including multi-view generated point clouds, occupancy grids, etc.); furthermore, the embodiment of the present application introduces a matching pair of depth information. By utilizing the three-dimensional perception data commonly found in passenger car intelligent perception systems, the three-dimensional points with known calibration external parameters are reprojected to the multi-view camera plane to construct a depth information matching pair. Since the depth matching pair comes from the same data source, it has a high degree of accuracy and can significantly improve the robustness and accuracy of feature matching; then, the embodiment of the present application projects the three-dimensional perception data (such as point cloud) to a cylindrical stitching map for initializing the local homography matrix. Traditional panoramic stitching methods usually use feature Matching is used to obtain a global single initial value of the homography matrix. However, due to the difference in depth information, the plane assumption is often not valid, resulting in a decrease in the accuracy of the global homography matrix. The present application describes directly calculating the local accurate homography matrix based on three-dimensional data surface reconstruction, thereby effectively solving the distortion and distortion problems in panoramic stitching and significantly improving the visual quality of the stitching effect. In addition, the embodiment of the present application adopts a global optimization method of manifold space in the grid optimization stage. Compared with directly performing panoramic stitching in Euclidean space, the manifold space optimization can directly handle the coupling relationship between local areas through linear approximation of local homography matrices or similarity matrices. Through this method, not only the overall optimization coordination is improved, but also the consistency and continuity of the homography matrix in the entire scene are ensured, thereby significantly improving the overall quality of the panoramic stitching effect.

[0141] As can be seen from the above, the embodiments of the present application obtain multiple images of the target device's surrounding environment and three-dimensional perception data; then stitch the images together based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. Thus, by incorporating the three-dimensional perception data collected about the target device's surrounding environment and stitching together the multiple images collected about the target device's surrounding environment, a more accurate and better stitched panoramic image corresponding to the target vehicle's surrounding environment can be obtained, thereby improving the panoramic image stitching effect.

[0142] To facilitate better implementation of the image processing method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above image processing method. The meanings of the terms herein are the same as those in the above image processing method, and the specific implementation details can be referred to the description in the method embodiment.

[0143] For example, Figure 5FIG. 2 is a schematic diagram of the structure of an image processing device provided in an embodiment of the present application. The image processing device may include an acquisition module 201 and a stitching module 202, as follows:

[0144] An acquisition module 201 is configured to acquire multiple images and three-dimensional perception data of the surrounding environment of a target device;

[0145] The stitching module 202 is configured to stitch images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0146] In one embodiment, the splicing module 202 is configured to:

[0147] Convert each image into a target coordinate system corresponding to the target device to obtain a first image;

[0148] The first images are stitched together based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0149] In one embodiment, the above-mentioned stitching of the first image based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment is used to:

[0150] Extracting corresponding feature matching point pairs for adjacent first images based on the three-dimensional perception data;

[0151] Based on the feature matching point pairs, the first images are stitched together to obtain a panoramic image corresponding to the surrounding environment.

[0152] In one embodiment, the feature matching point pairs include at least one of texture feature matching point pairs and depth matching point pairs.

[0153] In one embodiment, when the feature matching point pairs include depth matching point pairs, the step of extracting the depth matching point pairs includes:

[0154] Projecting the three-dimensional perception data onto an image plane corresponding to the first image;

[0155] A plurality of depth matching point pairs of the three-dimensional perception data in the adjacent first image are determined based on the projection result.

[0156] In one embodiment, when the feature matching point pairs include texture feature matching point pairs, the step of extracting the texture feature matching point pairs includes:

[0157] Texture feature points are extracted from adjacent first images to obtain a plurality of texture feature matching point pairs.

[0158] In one embodiment, the extraction of texture feature points in the adjacent first images to obtain multiple texture feature matching point pairs is specifically used for:

[0159] Extracting initial texture feature matching point pairs in the adjacent first image;

[0160] The initial texture feature matching point pairs are filtered to obtain multiple texture feature matching point pairs.

[0161] In one embodiment, the initial texture feature matching point pairs are filtered to obtain multiple texture feature matching point pairs, which are specifically used for:

[0162] Determining abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs based on the distribution of the initial texture feature matching point pairs in the first image to which they belong;

[0163] Abnormal initial texture feature matching point pairs are removed from the initial texture feature matching point pairs to obtain multiple texture feature matching point pairs.

[0164] In one embodiment, the above-mentioned determining abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs based on the distribution of the initial texture feature matching point pairs in the first image to which they belong is specifically used to:

[0165] Performing grid division on adjacent first images respectively to obtain a plurality of first grid areas of each first image;

[0166] Based on the initial texture feature matching points distributed in the matched first grid area in the adjacent first image, abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs are determined.

[0167] In one embodiment, the above-mentioned determination of abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs based on the initial texture feature matching points distributed in the matched first grid area in the adjacent first image is specifically used to:

[0168] Determining abnormal initial texture feature matching points from initial texture feature matching points distributed in the first grid area matched in the adjacent first image;

[0169] Based on the abnormal initial texture feature matching points, abnormal initial texture feature matching point pairs in the initial texture feature matching point pairs are determined.

[0170] In one embodiment, determining abnormal initial texture feature matching points from among the initial texture feature matching points distributed in the first grid area matched from the adjacent first image is specifically used to:

[0171] For each group of matched first grid areas in adjacent first images, determining matching point distribution similarity information corresponding to the first grid areas based on the number of initial texture feature matching points in one first grid area and another first grid area;

[0172] Based on the distribution similarity information of matching points, the abnormal initial texture feature matching points are determined.

[0173] In one embodiment, when the matching point distribution similarity information is not within a preset threshold range, the initial texture feature matching point in the first grid area is determined to be an abnormal initial texture feature matching point.

[0174] In one embodiment, the above-mentioned stitching of the first image based on the feature matching point pairs to obtain a panoramic image corresponding to the surrounding environment is specifically used for:

[0175] Calculating first transformation parameters between image planes corresponding to adjacent first images based on the feature matching point pairs;

[0176] The first images are stitched together according to the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

[0177] In one embodiment, the calculation of the first transformation parameters between image planes corresponding to adjacent first images based on the feature matching point pairs is specifically used to:

[0178] According to the coordinate information of the feature matching point pairs, the homography matrix between the image planes corresponding to the adjacent first images is calculated to obtain the first transformation parameters.

[0179] In one embodiment, the above-mentioned stitching of the first images according to the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment is specifically used for:

[0180] Meshing the geometric surface corresponding to the target coordinate system to obtain a plurality of second mesh regions;

[0181] Calculating a second transformation parameter corresponding to each second grid area in the first image according to the first transformation parameter and the three-dimensional perception data;

[0182] The first images are stitched based on the second transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

[0183] In one embodiment, the target coordinate system includes a cylindrical coordinate system, and the geometric surface includes a cylindrical surface.

[0184] In one embodiment, the calculation of the second transformation parameters corresponding to each second grid area in the first image based on the first transformation parameters and the three-dimensional perception data is specifically used to:

[0185] Mapping the three-dimensional perception data to a geometric surface corresponding to the target coordinate system to obtain first position information of each data point in the three-dimensional perception data in the geometric surface;

[0186] constructing a three-dimensional model based on the three-dimensional perception data, and obtaining second position information of each data point in the three-dimensional perception data in the three-dimensional model;

[0187] A second transformation parameter corresponding to the second grid area is determined based on the first position information and the second position information.

[0188] In one embodiment, the determining of the second transformation parameter corresponding to the second grid area based on the first position information and the second position information is specifically used for:

[0189] Determining initial second transformation parameters corresponding to the second grid area based on the first position information and the second position information;

[0190] Based on the initial second transformation parameters and the first transformation parameters, second transformation parameters corresponding to the second grid area are calculated.

[0191] In one embodiment, the calculation of the second transformation parameters corresponding to the second grid area based on the initial second transformation parameters and the first transformation parameters is specifically used to:

[0192] Calculating a correction matrix based on the initial second transformation parameters and the first transformation parameters;

[0193] According to the correction matrix, second transformation parameters corresponding to the second grid area are calculated.

[0194] In one embodiment, the calculation of the second transformation parameter corresponding to the second grid area according to the correction matrix is ​​specifically used for:

[0195] For each second grid area, calculating weight information corresponding to the feature matching point pair based on the grid center point and the coordinate information of the feature matching point pair in the second grid area;

[0196] Based on the weight information and the correction matrix, a second transformation parameter corresponding to the second grid area is determined.

[0197] In one embodiment, the image processing apparatus further includes an optimization module configured to:

[0198] The second transformation parameters are optimized to obtain optimized second transformation parameters.

[0199] In one embodiment, the above-mentioned optimization of the second transformation parameters to obtain the optimized second transformation parameters is specifically used for:

[0200] Converting the second transformation parameter from the current geometric space to the manifold space for optimization;

[0201] The optimized result is converted back to the current geometric space to obtain the optimized second transformation parameters.

[0202] In one embodiment, the above-mentioned conversion of each image into a target coordinate system corresponding to the target device to obtain a first image is used for:

[0203] Get the camera parameters of the camera that captures the image;

[0204] Based on the camera parameters, each image is converted into a target coordinate system corresponding to the target device to obtain a first image.

[0205] In one embodiment, the image processing apparatus further includes a three-dimensional display module configured to:

[0206] Projecting the panoramic image onto the surface of a three-dimensional model constructed based on the three-dimensional perception data to obtain a target three-dimensional model;

[0207] The target three-dimensional model is rendered to obtain a panoramic effect image corresponding to the surrounding environment.

[0208] In one embodiment, the target device comprises a vehicle.

[0209] As can be seen from the above, in the embodiment of the present application, acquisition module 201 acquires multiple images and three-dimensional perception data of the target device's surrounding environment; stitching module 202 stitches the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. Thus, by incorporating the three-dimensional perception data acquired for the target device's surrounding environment and stitching the multiple images acquired for the target device's surrounding environment, a more accurate and better stitched panoramic image corresponding to the target vehicle's surrounding environment can be obtained, thereby improving the panoramic image stitching effect.

[0210] Accordingly, the embodiment of the present application also provides a controller, such as Figure 6 As shown, Figure 6 Schematic diagram of the structure of the controller provided in an embodiment of the present application. The controller 300 includes a processor 301 having one or more processing cores, a memory 302 having one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 is electrically connected to the memory 302. It will be understood by those skilled in the art that the controller structure shown in the figure does not constitute a limitation of the controller, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0211] The processor 301 is the control center of the controller 300. It connects the various parts of the entire controller 300 using various interfaces and lines. It executes various functions of the controller 300 and processes data by running or loading software programs and / or units stored in the memory 302 and calling data stored in the memory 302. The processor 301 can be a processor CPU, a graphics processor GPU, a network processor (NP), etc., and can implement or execute the various methods, steps, and logic blocks disclosed in the embodiments of this application.

[0212] In the embodiment of the present application, the processor 301 in the controller 300 loads instructions corresponding to one or more application processes into the memory 302 according to the following steps, and the processor 301 runs the application stored in the memory 302 to implement various functions, such as:

[0213] Acquire multiple images and three-dimensional perception data of the surrounding environment of the target device; stitch the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0214] Furthermore, various functions implemented by running the application stored in the memory 302 can also be described in the aforementioned embodiments and will not be repeated here.

[0215] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0216] Optional, such as Figure 6 As shown, the controller 300 further includes: a touch screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. Among them, the processor 301 is electrically connected to the touch screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307 respectively. Those skilled in the art will understand that Figure 6 The controller structure shown in the figure does not constitute a limitation to the controller, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0217] The touch display screen 303 can be used for displaying a graphical user interface and receiving the operation instructions generated by the user acting on the graphical user interface. The touch display screen 303 may include a display panel and a touch panel. Among them, the display panel can be used for displaying the information input by the user or the information provided to the user and various graphical user interfaces of the controller, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light emitting diode (OLED), or the like. The touch panel can be used for collecting the touch operation of the user thereon or near it (such as the user uses any suitable object or accessory such as a finger, a stylus on the touch panel or near the touch panel), and generates corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 301, and can receive the command sent by the processor 301 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 301 to determine the type of touch event, and then the processor 301 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 303 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 303 can also be used as part of the input unit 306 to realize the input function.

[0218] The RF circuit 304 may be used to transmit and receive RF signals, so as to establish wireless communication with a network device or other controllers through wireless communication, and transmit and receive signals with the network device or other controllers.

[0219] The audio circuit 305 can be used to provide an audio interface between the user and the controller via a speaker and microphone. The audio circuit 305 can convert received audio data into electrical signals and transmit them to the speaker, which then converts them into sound signals for output. The microphone, on the other hand, converts collected sound signals into electrical signals, which are then received by the audio circuit 305 and converted into audio data. The audio data is then output to the processor 301 for processing, and then sent to another controller via the RF circuit 304. Alternatively, the audio data can be output to the memory 302 for further processing. The audio circuit 305 may also include an earphone jack to allow communication between an external headset and the controller.

[0220] The input unit 306 may be configured to receive input target video and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0221] Power supply 307 is used to supply power to various components of controller 300. Optionally, power supply 307 can be logically connected to processor 301 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 307 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0222] although Figure 6 Not shown in the figure, the controller 300 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be repeated here.

[0223] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in one embodiment, please refer to the relevant descriptions of other embodiments. It should be noted that the controller provided in the embodiment of the present application and the image processing method in the above embodiment are based on the same concept. The specific implementation process is detailed in the above method embodiment and will not be repeated here.

[0224] As can be seen from the above, the controller provided in the embodiments of the present application can obtain multiple images and three-dimensional perception data collected about the surrounding environment of the target device; and stitch the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment. Thus, by incorporating the three-dimensional perception data collected about the surrounding environment of the target device and stitching the multiple images collected about the surrounding environment of the target device, a more accurate and better stitched panoramic image corresponding to the surrounding environment of the target vehicle can be obtained, thereby improving the panoramic image stitching effect.

[0225] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0226] To this end, an embodiment of the present application provides a computer-readable storage medium including a computer program. When the computer program is executed on a controller, the computer program is used to cause the controller to perform any of the image processing methods provided in the embodiments of the present application. For example, the computer program may perform the following steps of the image processing method:

[0227] Acquire multiple images and three-dimensional perception data of the surrounding environment of the target device; stitch the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

[0228] Furthermore, for the detailed steps of the above method steps, please refer to the description in the above embodiments, which will not be repeated here.

[0229] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0230] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0231] Since the computer program stored in the computer-readable storage medium can execute any image processing method provided in the embodiments of the present application, the beneficial effects that can be achieved by any image processing method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0232] According to one aspect of the present application, a computer program product is also provided, including a computer program, which is stored in a computer-readable storage medium; when a processor of a controller reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the controller executes the methods provided in various optional implementations of the above embodiments.

[0233] In the above-described embodiments of the image processing apparatus, computer-readable storage medium, controller, device, and computer program product, the descriptions of each embodiment have different focuses. For portions not described in detail in a particular embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the above-described image processing apparatus, computer-readable storage medium, computer program product, controller, and their corresponding units can be referred to in the description of the image processing method in the above embodiments, and the details will not be repeated here.

[0234] The above is a detailed introduction to an image processing method, device, controller, equipment, computer-readable storage medium and computer program product provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An image processing method, characterized in that: include: Acquire multiple images and three-dimensional perception data of the surrounding environment of the target device; The images are stitched based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

2. The image processing method according to claim 1, wherein: The stitching of the images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment includes: Converting each of the images into a target coordinate system corresponding to the target device to obtain a first image; The first images are stitched based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment.

3. The image processing method according to claim 2, wherein: The step of stitching the first images based on the three-dimensional perception data to obtain a panoramic image corresponding to the surrounding environment includes: Extracting corresponding feature matching point pairs for adjacent first images based on the three-dimensional perception data; Based on the feature matching point pairs, the first images are stitched together to obtain a panoramic image corresponding to the surrounding environment.

4. The image processing method according to claim 3, wherein: The feature matching point pairs include at least one of texture feature matching point pairs and depth matching point pairs.

5. The image processing method according to claim 4, wherein: When the feature matching point pairs include depth matching point pairs, the step of extracting the depth matching point pairs includes: Projecting the three-dimensional perception data onto an image plane corresponding to the first image; A plurality of depth matching point pairs of the three-dimensional perception data in adjacent first images are determined based on the projection result.

6. The image processing method according to claim 4, wherein: When the feature matching point pairs include texture feature matching point pairs, the step of extracting the texture feature matching point pairs includes: Texture feature points are extracted from adjacent first images to obtain a plurality of texture feature matching point pairs.

7. The image processing method according to claim 6, wherein: The extracting of texture feature points in the adjacent first image to obtain a plurality of texture feature matching point pairs includes: Extracting initial texture feature matching point pairs in the adjacent first image; The initial texture feature matching point pairs are filtered to obtain multiple texture feature matching point pairs.

8. The image processing method according to claim 7, wherein: The filtering of the initial texture feature matching point pairs to obtain a plurality of texture feature matching point pairs includes: Determining abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs based on the distribution of the initial texture feature matching point pairs in the first image to which they belong; The abnormal initial texture feature matching point pairs are removed from the initial texture feature matching point pairs to obtain a plurality of texture feature matching point pairs.

9. The image processing method according to claim 8, wherein: The determining, based on the distribution of the initial texture feature matching point pairs in the first image, abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs includes: Performing grid division on adjacent first images respectively to obtain a plurality of first grid areas of each first image; Based on the initial texture feature matching points distributed in the matched first grid area in the adjacent first image, abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs are determined.

10. The image processing method according to claim 9, wherein: The determining of abnormal initial texture feature matching point pairs among the initial texture feature matching point pairs based on the initial texture feature matching points distributed in the matched first grid area in the adjacent first image includes: Determining abnormal initial texture feature matching points from the initial texture feature matching points distributed in the matched first grid area in the adjacent first image; Based on the abnormal initial texture feature matching points, an abnormal initial texture feature matching point pair among the initial texture feature matching point pairs is determined.

11. The image processing method according to claim 10, wherein: The determining of abnormal initial texture feature matching points from the initial texture feature matching points distributed in the first grid area matched in the adjacent first image includes: For each group of matched first grid areas in the adjacent first images, determining matching point distribution similarity information corresponding to the first grid areas based on the number of initial texture feature matching points in one first grid area and another first grid area; Based on the matching point distribution similarity information, abnormal initial texture feature matching points are determined.

12. The image processing method according to claim 11, wherein: When the matching point distribution similarity information is not within a preset threshold range, the initial texture feature matching point in the first grid area is determined to be an abnormal initial texture feature matching point.

13. The image processing method according to claim 3, wherein: The step of stitching the first images based on the feature matching point pairs to obtain a panoramic image corresponding to the surrounding environment includes: Calculating first transformation parameters between image planes corresponding to adjacent first images based on the feature matching point pairs; The first images are stitched together according to the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

14. The image processing method according to claim 13, wherein: The calculating, based on the feature matching point pairs, first transformation parameters between image planes corresponding to adjacent first images includes: According to the coordinate information of the feature matching point pair, a homography matrix between image planes corresponding to adjacent first images is calculated to obtain first transformation parameters.

15. The image processing method according to claim 13, wherein: The step of stitching the first images according to the first transformation parameters to obtain a panoramic image corresponding to the surrounding environment includes: Meshing the geometric surface corresponding to the target coordinate system to obtain a plurality of second mesh areas; Calculating a second transformation parameter corresponding to each second grid area in the first image according to the first transformation parameter and the three-dimensional perception data; The first images are stitched based on the second transformation parameters to obtain a panoramic image corresponding to the surrounding environment.

16. The image processing method according to claim 15, wherein: The target coordinate system includes a cylindrical coordinate system, and the geometric surface includes a cylindrical surface.

17. The image processing method according to claim 15, wherein: The calculating, according to the first transformation parameter and the three-dimensional perception data, a second transformation parameter corresponding to each second grid area in the first image includes: Mapping the three-dimensional perception data to a geometric surface corresponding to the target coordinate system to obtain first position information of each data point in the three-dimensional perception data in the geometric surface; constructing a three-dimensional model based on the three-dimensional perception data, and obtaining second position information of each data point in the three-dimensional perception data in the three-dimensional model; Based on the first position information and the second position information, a second transformation parameter corresponding to the second grid area is determined.

18. The image processing method according to claim 17, wherein: The determining, based on the first position information and the second position information, a second transformation parameter corresponding to the second grid area includes: determining, based on the first position information and the second position information, initial second transformation parameters corresponding to the second grid area; Based on the initial second transformation parameters and the first transformation parameters, second transformation parameters corresponding to the second grid area are calculated.

19. The image processing method according to claim 18, wherein: The calculating, based on the initial second transformation parameters and the first transformation parameters, the second transformation parameters corresponding to the second grid area includes: Calculating a correction matrix based on the initial second transformation parameters and the first transformation parameters; According to the correction matrix, second transformation parameters corresponding to the second grid area are calculated.

20. The image processing method according to claim 19, wherein: Calculating the second transformation parameter corresponding to the second grid area according to the correction matrix includes: For each second grid area, calculating weight information corresponding to the feature matching point pair based on the grid center point in the second grid area and the coordinate information of the feature matching point pair; Based on the weight information and the correction matrix, a second transformation parameter corresponding to the second grid area is determined.

21. The image processing method according to claim 15, wherein: Before stitching the first images based on the second transformation parameters to obtain a panoramic image corresponding to the surrounding environment, the method further includes: The second transformation parameters are optimized to obtain optimized second transformation parameters.

22. The image processing method according to claim 21, wherein: Optimizing the second transformation parameters to obtain optimized second transformation parameters includes: Converting the second transformation parameters from the current geometric space to the manifold space for optimization; The optimized result is converted back to the current geometric space to obtain optimized second transformation parameters.

23. The image processing method according to claim 2, wherein: The converting each of the images into a target coordinate system corresponding to the target device to obtain a first image includes: Obtaining camera parameters of a camera that captures the image; Based on the camera parameters, each of the images is converted into a target coordinate system corresponding to the target device to obtain a first image.

24. The image processing method according to any one of claims 1 to 23, wherein: The method further comprises: Projecting the panoramic image onto a surface of a three-dimensional model constructed based on the three-dimensional perception data to obtain a target three-dimensional model; The target three-dimensional model is rendered to obtain a panoramic effect image corresponding to the surrounding environment.

25. The image processing method according to any one of claims 1 to 23, wherein: The target device includes a vehicle.

26. A controller, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is enabled to perform the steps of the method according to any one of claims 1 to 25.

27. A device, characterized in that The device includes the controller of claim 26.

28. The apparatus of claim 27, comprising a vehicle.

29. A computer-readable storage medium, characterized in that The method comprises a computer program, and when the computer program is run on a controller, the computer program is used to make the controller perform the steps of the method according to any one of claims 1 to 25.

30. A computer program product, characterized in that The method comprises a computer program or instructions, which implements the steps of the method according to any one of claims 1 to 25 when executed by a processor.