A single vehicle surround view image real-time splicing method based on water drop state perception field fusion

The method of real-time dynamic stitching of single-vehicle surround view images by water droplet-state perception field fusion solves the problem of blind spots in surround view technology for large vehicles, realizes real-time, blind-spot-free environmental perception, and improves vehicle driving safety and operational convenience.

CN115439324BActive Publication Date: 2026-07-31BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2022-09-05
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing automotive surround view technology is difficult to adapt to the body rotation of large non-rigid vehicles such as trucks and trailers, resulting in blind spots and failing to provide real-time, blind-spot-free environmental perception.

Method used

A real-time dynamic stitching method for single-vehicle surround view images is adopted using a waterdrop-state perception field-of-view fusion. Through feature point generation and matching, image stitching, and parameter table updating, a waterdrop-state perception field-of-view model is established to achieve real-time dynamic stitching and 3D perception of camera images.

Benefits of technology

It provides real-time, blind-spot-free panoramic environmental display for large non-rigid vehicles, improving the vehicle's environmental perception capabilities and enhancing driving safety and ease of operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115439324B_ABST
    Figure CN115439324B_ABST
Patent Text Reader

Abstract

This invention proposes a real-time dynamic stitching method for single-vehicle surround view images based on droplet-state perceptual field-of-view fusion. This method utilizes panoramic image information provided by the vehicle's surround-view vision system, enabling the perception of a large range of environmental information. It provides real-time dynamic stitched surround view images for rotatable non-rigid special vehicles such as trucks, and can also provide droplet-state 3D perceptual field-of-view fusion images at the same scale, achieving blind-spot-free, large-area real-time environmental display. By employing a feature point extraction algorithm based on the Superpoint architecture and a feature matching algorithm based on the SuperGlue architecture, the feature matching results can be effectively improved, eliminating matching point pairs that cannot be used for pose estimation, and enhancing the accuracy and reliability of relative position parameters between cameras. By combining video stabilization and stitching within a unified framework, stable dynamic panoramic stitched images can be obtained. Combined with the constructed droplet model and related projection matrices, real-time 3D panoramic surround view images can be provided for vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, specifically relating to a real-time stitching method for single-vehicle surround view images based on water droplet state perception field fusion. Background Technology

[0002] With the continuous development of technology and the increasing demands of people for work and life, automobiles have become a necessity in countless households. Drivers hope that automotive imaging systems can provide a more comprehensive, expansive, and blind-spot-free view of the vehicle's surroundings, thereby assisting drivers in achieving more precise and safer vehicle control. Currently, the latest generation of automotive surround-view technology can achieve 3DVM display of the vehicle's surrounding environment, and more traditional reversing camera technology can also provide some assistance to drivers. However, these technologies are designed for passenger vehicles such as sedans. While they can better meet the perception needs of sedan drivers, when faced with the characteristics of large commercial vehicles, such as large distances between vehicles, high seating positions, and non-fixed front and rear connections, traditional surround-view technology struggles to compensate for the blind spots around large vehicles, causing numerous inconveniences for drivers. Therefore, the research need for automotive surround-view perception technology specifically for large vehicles is becoming increasingly urgent, especially in addressing the perception challenges of non-fixed front and rear connections and the perception challenges of mutual occlusion between large vehicles.

[0003] Automotive surround view technology uses 4 to 8 wide-angle cameras surrounding the vehicle to cover the entire field of view around the vehicle. These cameras simultaneously capture multiple video images, which are then processed into a 360-degree top-down view of the vehicle's surroundings. This view is then displayed on the center console screen, allowing the driver to clearly see if there are any obstacles around the vehicle and understand their relative positions and distances, thus assisting in vehicle operation. It is not only highly intuitive but also eliminates blind spots, helping drivers to confidently maneuver the vehicle into parking spaces or navigate complex road surfaces, effectively reducing the occurrence of accidents such as scratches, collisions, and getting stuck. Current automotive surround view technology assumes that the cameras are fixed to the vehicle, treating the entire vehicle as a rigid body. The overall technical approach uses a static stitching method. First, the fixed multiple cameras are individually calibrated to obtain distortion-free images from each camera. Second, the transformation parameters from each camera image to the projection plane are calculated, and a lookup table is created. Then, all images are projected onto the same projection plane using the lookup table, and image fusion processing is performed to form a unified projection plane image. Finally, methods such as virtual viewpoints are used to transform the projection plane onto a virtual imaging surface for display. Currently, static stitching methods are favored by industry professionals due to their advantages such as simple camera setup, easy parameter calculation, and high real-time imaging performance. However, when faced with changes in parameters such as shaking or rotation, the lookup parameters in static stitching methods cannot be changed accordingly. This makes current automotive surround view technology unsuitable for large, non-rigid special vehicles such as trucks and trailers that experience body rotation.

[0004] To address the aforementioned issues, researchers have conducted extensive studies on the upstream work of image stitching—feature point generation and matching. As of 2020, the latest feature generation and matching methods have achieved 15 FPS, initially meeting the real-time requirements of image preprocessing. Dynamic stitching of surround view images of vehicles such as trucks and trailers, which undergo vehicle rotation, has become possible.

[0005] Therefore, the real-time dynamic stitching method of single-vehicle surround view image for waterdrop-state perception field fusion provides real-time dynamic stitched images of rotatable non-rigid vehicles, and at the same time provides a three-dimensional perception field fusion model of the vehicle at the same scale. It is of great significance for the perception and display of special vehicles and has extremely high practical value. Summary of the Invention

[0006] In view of this, the present invention provides a real-time dynamic stitching method for single-vehicle surround view images based on droplet-state perception field fusion. It improves the vehicle surround view system in aspects such as feature point generation and matching, including system design, feature generation, feature matching, image stitching and shaking reduction, parameter table updating, and projection plane fusion. This enables the vehicle surround view system parameter table to be updated, and the surround view capability covers large non-rigid special vehicles such as trucks and trailers that rotate. It provides a three-dimensional perception field fusion model for the vehicle at the same scale, enabling the vehicle surround view technology to balance accuracy and efficiency, achieving real-time dynamic image stitching without blind spots.

[0007] A method for real-time dynamic stitching of single-vehicle surround view images based on water droplet-state perception field fusion includes the following steps:

[0008] Step S1: Set up the vehicle surround vision system and calibrate the parameters of the vehicle surround vision system;

[0009] Step S2: The vehicle surround vision system adopts a water droplet-state perception field-of-view model, which conforms to general quadratic surface constraints, i.e.:

[0010]

[0011] In the formula, f(p) is the field-of-view height function related to the perceived depth p; P vf To represent the perceived depth at a height of 0, P v For a height of vehicle height h v The corresponding horizontal perception depth of the vehicle at that time;

[0012] Establish a two-dimensional image coordinate system and a three-dimensional 3D projection coordinate system for the camera. Project each pixel in the camera image onto the 3D projection coordinate system of the water droplet-state 3D perception field of view model. Then:

[0013] When the height is 0, the correspondence between the pixels in the 2D image coordinate system and the 3D projected coordinate system is as follows:

[0014]

[0015] Where (x′,y′) represents the coordinates of a pixel in the two-dimensional image coordinate system, and (x,y,z) represents the corresponding coordinates of a pixel in the 3D projection coordinate system; α and β are constants, where α is the lift coefficient and β is the curvature coefficient, and m represents the image width, i.e., the number of pixels;

[0016] When the height is not 0, the correspondence between the pixels in the 2D image coordinate system and the 3D projected coordinate system is as follows:

[0017]

[0018] Where γ is the resolution adjustment coefficient;

[0019] The coordinates (x, y, z) of each pixel point mapped to the water droplet-state perception field of view, and the corresponding perception depth p, are obtained by solving equations (1) and (2).

[0020] Step S3: Based on the vehicle surround vision system parameters calibrated in step S1 and the correspondence between the pixel points in the two-dimensional image coordinate system and the 3D projection coordinate system obtained in step S2, solve the relative motion of the two-dimensional pixel points in the image coordinate system to the corresponding pixels in the 3D projection coordinate system where the 3D model is located, and establish the initial parameter table of the vehicle surround vision system.

[0021] Step S4: Synchronously acquire video data from each camera in the vehicle surround vision system, extract feature points in parallel, perform feature matching, and then update the parameter table.

[0022] Step S5: Based on the updated parameter table, project the images from each camera on the vehicle onto the water droplet perception model, and then stitch the images together.

[0023] Preferably, in step S4, the video data from each camera is stabilized before the parameter table is updated.

[0024] Preferably, the water droplet-state perception field of view model is a bowl-shaped model, with the bowl wall being a curved surface formed by rotating a set type curve around the central axis, the bowl bottom being a circle surrounded by the bowl wall, and the center of the bowl bottom being a rectangle representing the vehicle's projection onto the horizontal plane.

[0025] Preferably, the set type curve is a parabola, a circular arc, or another quadratic curve.

[0026] Ideally, the Superpoint architecture should be used for feature point extraction.

[0027] Ideally, feature matching should be implemented using the SuperGlue architecture.

[0028] Preferably, the vehicle surround vision system includes four cameras, which are respectively positioned above the front license plate, in the middle of the left rear door frame, above the rear license plate, and in the middle of the right rear door frame.

[0029] Preferably, the vehicle surround vision system is based on an Ackerman motion model vehicle.

[0030] Preferably, the camera imaging model adopts the Taylor expansion camera imaging model proposed by Scaramuzza.

[0031] Preferably, the overlapping field of view of the vehicle surround vision system faces both sides of the vehicle.

[0032] The present invention has the following beneficial effects:

[0033] This invention proposes a real-time dynamic stitching method for single-vehicle surround view images based on water droplet-state perception field fusion. This method uses panoramic image information provided by the vehicle surround view vision system, which can perceive a large range of environmental information. It provides real-time dynamic stitching images of the vehicle surround view for rotatable non-rigid special vehicles such as trucks, and can also provide water droplet-state three-dimensional perception field fusion images of the vehicle at the same scale, realizing real-time environmental display with no blind spots and a large range.

[0034] By combining the feature point extraction algorithm based on the Superpoint architecture with the feature matching algorithm based on the SuperGlue architecture, the feature matching results can be effectively improved, matching point pairs that cannot be used for pose estimation can be eliminated, and the accuracy and reliability of the relative position parameters between cameras can be improved.

[0035] By combining video stabilization and stitching within a unified framework, stable dynamic panoramic stitched images can be obtained; by combining the constructed water droplet model with the relevant projection matrix, real-time 3D panoramic surround view images can be provided for vehicles. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the hardware structure of the vehicle surround vision system in the perception method of the present invention;

[0037] Figure 2 This is a schematic diagram of the installation method of the automotive surround vision system of the present invention on a carrier.

[0038] Figure 3 This is a flowchart of the algorithm for real-time image stitching and fusion of vehicles according to the present invention;

[0039] Figure 4 This is a diagram of the single-vehicle water droplet state perception field of view model of the present invention;

[0040] Figure 5 This is a diagram of the single-vehicle water droplet-state surround perception model of the present invention. Detailed Implementation

[0041] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Those skilled in the art can understand the advantages and functions of the present invention based on the content described herein. The present invention can also be implemented and applied in other different ways.

[0042] This invention relates to a real-time dynamic stitching method for single-vehicle surround view images using droplet-state perception field fusion, including the hardware structure and software algorithm of an automotive surround view vision system. The hardware structure of the automotive surround view vision system consists of several identical cameras with identical image acquisition capabilities. The overall perception range of the automotive surround view vision system can cover the surrounding area of ​​the system carrier in the horizontal direction, achieving environmental information perception without blind spots in the horizontal direction. In this embodiment, the hardware of the automotive surround view vision system consists of four cameras with hardware synchronization triggering function, providing a 360° field of view in the horizontal direction.

[0043] The algorithm flow for the real-time dynamic stitching method of single-vehicle surround view images is as follows:

[0044] Step S1: Based on the field of view of each camera in the vehicle surround vision system, the required depth of the perceived environment, and the carrier motion constraints, determine the camera arrangement, the number of cameras, and the baseline length between cameras according to certain standards. Then, select an appropriate camera imaging model according to the camera type to calibrate the parameters of the vehicle surround vision system. The parameters to be calibrated include the intrinsic parameters of the camera imaging model and the initial extrinsic parameters representing the pose relationship between the cameras.

[0045] The vehicle surround-view vision system is based on an Ackerman motion model vehicle. The overlapping fields of view of the system are oriented towards the sides of the vehicle, which have richer texture features. During vehicle operation, this allows more environmental features to be observed by different cameras, which is beneficial for calculating the relationships between cameras during pose estimation. The camera imaging model adopts the Taylor expansion camera imaging model proposed by Scaramuzza. This model is applicable to large field-of-view cameras with significant distortion. Its intrinsic parameters include inverse projection parameters and projection parameters. The calibrated inverse projection parameters can be used to calculate the direction vector of the corresponding spatial point in the camera coordinate system, i.e., the observation direction vector. The calibrated projection parameters can be used to calculate the pixel coordinates of the imaging position of the point based on its three-dimensional coordinates in the camera coordinate system.

[0046] Step S2: Establish a single-vehicle waterdrop field-of-view model and determine the correspondence between the planar image coordinates and the waterdrop field-of-view model. The waterdrop-shaped sensing field-of-view model is a bowl-like model. The bowl wall can be a surface formed by rotating a parabola, arc, or other quadratic curve around a central axis. The bowl bottom is a circle enclosed by the bowl wall, and the center of the bowl bottom is a rectangle projected onto the horizontal plane by the vehicle. The image data collected by the vehicle will be projected onto the bowl wall according to the 360-degree latitude and longitude distribution, supplementing the 360-degree surround view of the vehicle. The upper part of the sensing field of view may not be closed.

[0047] To meet vehicle perception requirements, the bottom of the perception field of view should be an image with low curvature, approximating a top-down view of the vehicle, thus providing a relatively clear image of the vehicle's surrounding environment; the side walls of the perception field of view should be images with high curvature, approximating a side view of the vehicle, thus providing a detailed image of the vehicle's surrounding environment. Assume the vehicle length is l. v , vehicle width w v vehicle height h v At vehicle height h v At this location, the vehicle's horizontal perception depth p v Vehicle bottom horizontal sensing depth P vf Then at the vehicle height h v The location should be the center of the side view of the vehicle, corresponding to the center of the sidewall of the perceived field of view. The perceived field of view model should conform to general quadratic surface constraints, i.e.:

[0048]

[0049] In the formula, f(p) is the field-of-view height function related to the perceived depth p; it can be seen from the formula that when the height is 0, i.e., the ground, the perceived depth is P. vf At a height of vehicle height h v At that time, the corresponding vehicle horizontal perception depth is P v .

[0050] Taking a camera on one side of the vehicle as an example, its planar field-of-view image resolution is m×n. The lower left pixel of the image is selected as the origin. A two-dimensional Cartesian coordinate system is established with the image length as the x-axis and the width as the y-axis, serving as the two-dimensional image coordinate system. The correspondence between each point (x′, y′) in the image and the original image coordinates (u, v) in the two-dimensional image coordinate system is as follows:

[0051]

[0052] Where n is the total number of pixels in the image width.

[0053] Secondly, a three-dimensional Cartesian coordinate system is established in three-dimensional space as the 3D projection coordinate system. The xOy plane of this coordinate system coincides with the ground, where the origin O is located at the geometric center of the vehicle's top-view image, the x-axis points towards the front of the vehicle, and the z-axis points vertically upward. Based on the initially determined quadratic curve constraints, the proposed droplet-state three-dimensional perception field of view model can be represented in the 3D projection coordinate system as follows:

[0054]

[0055] α and β are constants, where α is the lift coefficient, determining the overall height of the field of view model; and β is the curvature coefficient, determining the degree of curvature of the sidewalls of the projection model. These parameters can be dynamically adjusted according to the actual usage scenario.

[0056] Based on the two-dimensional image coordinate system and the 3D projection coordinate system, pixels in the two-dimensional image are projected one by one into the waterdrop-shaped 3D perception field-of-view model according to the mapping relationship between pixel coordinates. Assuming the camera's field of view is 90°, it is mapped to the first octant of the 3D projection coordinate system. When the height is 0 (ground), the tangent between the waterdrop-shaped perception field-of-view model and the ground is a radius P. vf The coordinate relationships of the circle are as follows:

[0057]

[0058] When the height is not 0, the coordinate correspondence satisfies:

[0059]

[0060] Where γ is the resolution adjustment coefficient, its function is to adjust the vertical resolution of the planar image to adapt to the height of the droplet-state perception field of view model, so that at the vehicle height h v The location should be the side-view frontal image of the vehicle, i.e., the center of the sidewall of the corresponding sensing field of view. Based on the four equations shown, the coordinates (x, y, z) of each pixel in the planar image mapped to the waterdrop-shaped sensing field of view, and the corresponding sensing depth p, can be solved.

[0061] Assuming the camera's center pixel corresponds to the frontal viewpoint of the sidewall of the field of view. If the camera's field of view is w, then the number of cameras installed on the vehicle should be greater than [a certain value]. (Where, [·] indicates rounding up). With the vehicle stationary, calculate the projection correspondence between the images from each camera and the vehicle's waterdrop-state perception field of view, i.e., the relationship between camera image pixels (m*n) and waterdrop-state perception field of view pixels (2π*P). vf *2h v *θ, where θ is the resolution adjustment coefficient, corresponds to the projection mapping.

[0062] Step S3: Solve for the relative motion of the two-dimensional pixel in the image coordinate system to the corresponding pixel in the three-dimensional Cartesian coordinate system of the 3D model. Given the camera intrinsic parameters, this can be transformed into solving for the pose of the pixel in the camera coordinate system relative to the coordinate system of the 3D model, including the rotation matrix R and translation vector t from the 3D model coordinate system to the camera coordinate system. Its mathematical description is as follows:

[0063]

[0064] Where p is the two-dimensional coordinate of the point in the image coordinate system, P C Let P be the coordinates of a point in the camera coordinate system, K be the camera intrinsic parameter matrix, and P be the coordinates of the point in the camera coordinate system. W Let ω be the coordinates of the point in the world coordinate system of the 3D model, and ω be the depth of the point. This is the pose we need to solve for. Let P W =[x W ,y W ,z W ] T ,P C =[x C ,y C ,z C ] T ,P′=[x C / z C ,y C / z C ,1] T =[u1,v1,1] T Where P′ represents the pixel's representation in the camera's normalized coordinate system. p=[u,v,1] represents the pixel's coordinate representation after expanding to homogeneous coordinates in the pixel coordinate system, resulting in the complete mathematical expression:

[0065]

[0066] Given the camera intrinsic parameters, the 2D plane pixel coordinates are first transformed into camera normalized coordinates. Then, the equations for the pose parameters are obtained according to the corresponding normalization constraints. To solve the 12 unknown parameters in the corresponding rotation and translation matrices, at least 6 pairs of points need to be selected to solve the equations, and the mapping matrix between the plane image coordinates and the corresponding point coordinates of the water droplet model is obtained. Combined with the initial extrinsic parameters representing the pose relationship between cameras, an initial parameter table is established.

[0067] Step S4: Multiple cameras are synchronously triggered via hardware to acquire video data simultaneously. The Superpoint architecture feature extraction algorithm is used in multiple threads to extract feature points from multiple images in parallel. Feature point information includes the location of keypoints and local descriptors. Subsequently, the matching results are filtered based on the angle between the observation direction vectors obtained from the inverse projection of the two-dimensional feature points. A Sampson error check is introduced to verify whether the matching results meet epipolar geometric constraints, and matching point pairs with errors exceeding a threshold are removed to improve matching accuracy. The Sampson error calculation method is as follows:

[0068]

[0069] in, E is the unit direction vector in the camera coordinate system after the inverse projection of the two-dimensional feature points. 12 Let be the essential matrix between the two images.

[0070] The Superpoint architecture feature point extraction algorithm is a feature point extraction algorithm that uses a neural network method to generate feature points. First, virtual features such as points, lines, and surfaces are constructed in the simulation space, and the corresponding feature point coordinates are recorded to build a virtual dataset as completely as possible. Then, the neural network is trained. The method of generating data points through the neural network makes the feature points more evenly distributed in the image. An image pyramid is constructed by downsampling images, and feature points are extracted from multiple images, making the feature points scale-invariant. In the programming implementation, multi-threaded programming is used to process the image feature extraction process from multiple cameras in parallel, enabling efficient feature extraction.

[0071] At the same keyframe, the SuperGlue architecture feature matching algorithm is used to process images acquired by different cameras at the same time, calculating the relative positional relationship between the cameras. Assume the extrinsic parameter matrix of camera i is T. i The extrinsic parameter matrix of camera j is T. j Then the relative pose of cameras i and j is T. ij The relative pose at frame n is The relative pose at frame (n+1) is Calculate the pose change function between cameras, if:

[0072]

[0073] Then the camera pose is updated for frame n+1, thereby updating the parameter table.

[0074] Step S5: For the same camera, the SuperGlue architecture feature matching algorithm is used to process two adjacent frames in the video stream acquired by that camera, calculating the relative positional relationship between the two frames, thereby performing image stabilization on the video streams acquired by each camera. Let C be the image calculation matrix for camera i in the (n-1)th frame. i (n-1), the relative positional relationship between the (n-1)th frame and the nth frame is F. i (n), where the image matrix calculated in the nth frame is C. i (n)=F i (n)·C i (n-1). By optimizing the energy function

[0075]

[0076] Obtain the image stabilization transformation matrix P of camera i from frame (n-1) to frame n. i (n), where r is a certain number of images before the nth frame, which can generally be set between 30 and 50 depending on the situation, and ω n,r is a constant, representing the weight between the two items.

[0077] The SuperGlue architecture feature point matching algorithm is a graph neural network-based feature point matching algorithm. During the feature matching process, when matching two-dimensional feature points from different images, the feature points are constructed into a graph according to the traversal order. Matching is performed within the graph set corresponding to each feature point, improving the efficiency and accuracy of feature matching. The angle between the observation direction vectors obtained from the inverse projection of the two-dimensional feature points is used to filter the matching results, avoiding the introduction of excessive errors and reducing false matches, thus improving the accuracy of camera pose construction. Furthermore, Sampson error is used to check whether the matching results meet epipolar geometric constraints, eliminating matching point pairs with errors exceeding a threshold, thereby improving the accuracy of feature matching.

[0078] Step S6: Assume that the relative poses of camera i and camera j in the nth frame are... The homography matrix between the nth frame images of camera i and camera j can be obtained. This allows for image stitching between two cameras operating at the same time. This is achieved by optimizing the following energy function:

[0079]

[0080] To obtain the optimal P i ,P j , Parameters, where P i ,P j Let P be the stabilization transformation matrix calculated in the previous step, where β is a constant representing the weight between the stabilization and stitching terms. Through the optimized P... i,P j The simulated camera pose after image stabilization correction is obtained, combined with Update the parameters by adjusting the relative pose between cameras and updating the optimized parameter table.

[0081] Step S7: In multiple camera processing threads, the parameter table is called respectively to project the preprocessed images from each camera onto the waterdrop-shaped sensing field of view model. The overlapping areas of multiple images on the sensing field of view are fused to eliminate ghosting and stitching seams. Global brightness normalization and white balance processing are performed on the entire waterdrop-shaped sensing field of view to output the vehicle's waterdrop-shaped surround-view sensing field of view data.

[0082] In this process, both the projection of images acquired by each camera onto the droplet-shaped perception field of view model and the projection of the droplet-shaped perception field of view acquired by the virtual viewpoint onto the virtual projection surface can employ perspective transformation and inverse perspective transformation methods. First, key feature points in the image are extracted, including but not limited to traditional checkerboard corner points and lane lines. Second, based on the correspondence between the front and rear of the vehicle, the extracted feature points are evenly distributed on the projection plane. Then, based on the correspondence between the image feature points and the feature points on the projection plane, the homography transformation matrix between the image plane and the projection plane is calculated. Finally, all pixels on the image plane are transformed onto the projection plane using the homography transformation matrix.

[0083] Example:

[0084] This implementation provides a real-time dynamic stitching method for multi-vehicle surround view images using droplet-state perception field fusion. A schematic diagram of the hardware structure used in this method is shown below. Figure 1 As shown in the figure, the car surround vision system consists of four cameras, which are respectively placed above the front license plate, in the middle of the left rear door frame, above the rear license plate, and in the middle of the right rear door frame, forming four perspectives of the vehicle: front, left, rear, and right. Figure 2 This is a schematic diagram of the installation and arrangement of the automotive surround vision system on the carrier. The horizontal field of view of each camera is approximately 120°. According to the requirements of the field of view of the automotive surround vision system in this invention, the field of view of the automotive surround vision system in the horizontal direction completely covers the area around the carrier, and the arrangement of the cameras makes the four overlapping fields of view of the multi-camera vision system face towards both sides of the carrier.

[0085] like Figure 3 As shown, the real-time dynamic stitching method for multi-vehicle surround view images includes the following steps:

[0086] Step S1: Image Acquisition by the Automotive Surround View Vision System. The circuit synchronously triggers four cameras at a frequency of 15Hz, acquiring image data from the four cameras with a consistent time base and timestamp information. Simultaneously, pre-calibrated Taylor expansion intrinsic parameters and pose transformation extrinsic parameters are acquired and associated with the corresponding camera images. In this embodiment, the four cameras are arranged according to the design described above. They receive trigger signals and return image information via the GMSL interface. The Taylor expansion intrinsic parameters include five affine transformation parameters, four inverse projection Taylor expansion coefficients, and nine projection Taylor expansion coefficients. The pose transformation extrinsic parameters include six parameters: Cayley rotation parameters and translation vector parameters.

[0087] Step S2: Initialize the parameter table of the vehicle surround view vision system. Use a 9*9 checkerboard calibration board as feature markers, with a checkerboard size of 10cm*10cm. Adjust the position of the checkerboard on the front camera so that the checkerboard occupies more than one-third of the central area of ​​the image. Select any four corner points of the checkerboard center in the image and record the image coordinates corresponding to each corner point. Select four corresponding position points on the front projection surface of the teardrop-shaped perception field of view, so that the field of view angle occupied by the checkerboard in the vehicle surround view is the same as the field of view angle occupied by the position points in the teardrop-shaped field of view model. Calculate the homography matrix between the coordinates of the image feature points and the coordinates of the selected points in the teardrop model. Calculate the homography matrix parameters corresponding to the four cameras in sequence and store them in the parameter table to complete the initialization. If an initialization parameter table already exists, skip directly to step S3.

[0088] Step S3: The parameter table of the vehicle surround-view vision system is updated in real time. In this implementation case, the Superpoint architecture feature extraction algorithm is used to extract image features. The neural network dataset is generated in virtual simulation software. Classic geometric objects such as cubes, cuboids, and spheres are constructed in the dataset, and special corner points such as vertices and interior points on the objects are used as feature points. This feature point data is recorded as the network training set. After extracting the corner points, the quadtree algorithm is used to segment the image, and 1000 key points are extracted evenly in the image to make the feature points evenly distributed in the image. Feature extraction is performed on multiple images within the pyramid of downsampled images to make the feature points scale-invariant. The feature point information obtained by the feature extraction algorithm includes the key point positions and the descriptors obtained by the convolutional network. Feature points of each camera image are extracted in parallel through multi-threaded programming.

[0089] The SuperGlue architecture feature matching algorithm is used for real-time pose estimation between camera images. This process first uses a graph neural network method to match point pairs. Based on the feature extraction algorithm, a feature point map is first constructed. An arbitrary feature point is selected from the feature point set, and its surrounding feature points are recorded counter-clockwise to form a map. The map contains the center point's location information and the surrounding point's location information. The maps from different images are then fed into the neural network to match feature point pairs. Considering the initialization and real-time update states, in the initialization state, only 2D feature points from different camera images within the overlapping field of view are matched. Matching point pairs with an angle between the observation direction vectors less than 1.5° or greater than 60° are discarded. A Sampson error threshold of 0.01 is used to check whether the matching point pairs satisfy the epipolar geometry constraint. Then, the map point positions are calculated using triangulation, and the camera poses of the vehicle surround-view vision system are initialized simultaneously. In the real-time update state, matching is first performed using 2D feature points from different cameras. Based on the initial pose direction, all matching point pairs with an angle less than 15° or greater than 178.5° are discarded. Then, a nonlinear optimization method is used to solve for the current camera of the multi-camera vision system. In this implementation, the nonlinear optimization is based on the Levenberg-Marquardt iterative optimization algorithm. After the camera pose is calculated, the parameter table can be updated.

[0090] Step S4: Image stabilization preprocessing for each camera in the surround-view system. In this embodiment, based on the feature point extraction and matching method in step S3, inter-frame relative pose estimation is performed on the images acquired by each camera to obtain the relative positional relationship between each frame, resulting in a relative transformation matrix F. Based on all the obtained relative transformation matrices, the current image calculation matrix C can be calculated, and the image calculation matrix of the first frame can be set as an identity matrix. Based on the obtained image calculation matrix C, the stabilization term E is optimized iteratively, and the stabilization transformation matrix P for each frame of each camera is calculated, thereby removing the pose changes caused by the camera's own shaking, while simultaneously optimizing and smoothing adjacent frames to reduce image shaking and improve video viewing quality.

[0091] Step S5: Image Stabilization and Stitching under a Unified Framework. Using the initial parameter table obtained in Step S3, the pose relationships between each camera are acquired, thereby calculating the homography matrix H corresponding to the images between cameras. The stabilization transformation matrix P and the corresponding homography matrix H of two adjacent cameras are taken, and stabilization and stitching optimization are performed under a unified framework. Based on the calculated stabilization transformation matrix P, the optimal homography matrix between the two camera images in the current frame is obtained by minimizing the energy function. Compared to the homography matrix H before optimization, the optimized... On the one hand, it can alleviate the jitter between adjacent images after stitching; on the other hand, it can reduce ghosting and distortion in the stitched images. This is achieved through the optimized homography matrix. The poses between the cameras are recalculated, and the parameter table is updated again by combining the homography matrix from the corresponding image plane to the waterdrop-state perception field of view.

[0092] Step S6: Display of single-vehicle surround view images. In this implementation case, after updating the parameter table, the images from the vehicle's four cameras can be projected onto the droplet-state perception model using an inverse perspective transformation method. The droplet-state perception field of view model is as follows: Figure 4 As shown. In the overlapping image region, the center line of the overlapping region is calculated. Using the center line as the distance center, the weighting coefficient of the overlapping region on the left image decreases from 1 to 0 from left to right, and the weighting coefficient of the overlapping region on the right image decreases from 1 to 0 from right to left. The left and right images are then weighted and fused in the overlapping region to eliminate the stitching seam. Then, the image intensity is normalized to prevent local over-brightness. This achieves the display of a single-vehicle surround view image. The single-vehicle water droplet state perception model is as follows: Figure 5 As shown.

[0093] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A real-time dynamic splicing method for single vehicle surround view image based on water drop state perception field fusion, characterized in that, Includes the following steps: Step S1: Set up the vehicle surround vision system and calibrate the parameters of the vehicle surround vision system; Step S2: The vehicle surround vision system adopts a water droplet-state perception field-of-view model, which conforms to general quadratic surface constraints, i.e.: In the formula, To perceive depth The relevant field of view height function; To represent the perceived depth at a height of 0. To be at the height of the vehicle The corresponding horizontal perception depth of the vehicle at that time; Establish a two-dimensional image coordinate system and a three-dimensional 3D projection coordinate system for the camera. Project each pixel in the camera image onto the 3D projection coordinate system of the water droplet-state 3D perception field of view model. Then: When the height is 0, the correspondence between the pixels in the 2D image coordinate system and the 3D projected coordinate system is as follows: (1) in, This represents the coordinates of a pixel in a two-dimensional image coordinate system. () represents the coordinates of a pixel in the 3D projected coordinate system; All are constants. The lifting coefficient, The curvature coefficient, This indicates the image width, i.e., the number of pixels. When the height is not 0, the correspondence between the pixels in the 2D image coordinate system and the 3D projected coordinate system is as follows: (2) wherein is a resolution adjustment factor; According to equations (1) and (2), the coordinates of each pixel point mapping to the water droplet state perception field are solved , and the corresponding perception depth ; Step S3: Based on the vehicle surround vision system parameters calibrated in step S1 and the correspondence between the pixel points in the two-dimensional image coordinate system and the 3D projection coordinate system obtained in step S2, solve the relative motion of the two-dimensional pixel points in the image coordinate system to the corresponding pixels in the 3D projection coordinate system where the 3D model is located, and establish the initial parameter table of the vehicle surround vision system. Step S4: Synchronously acquire video data from each camera in the vehicle surround vision system. Under the same keyframe, use the SuperGlue architecture feature matching algorithm to extract feature points in parallel from the images acquired by different cameras at the same time, perform feature matching, and calculate the relative positional relationship between the cameras. Assume the relative poses of cameras i and j are... The relative pose at frame n is The relative pose at frame n+1 is Calculate the pose change function between cameras, if: Then the camera pose is updated for frame n+1, thereby updating the parameter table; Step S5: Based on the updated parameter table, project the images from each camera on the vehicle onto the water droplet perception model, and then stitch the images together.

2. The method for real-time dynamic stitching of single-vehicle surround view images based on water droplet state perception field fusion as described in claim 1, characterized in that, In step S4, the video data from each camera is stabilized, and then the parameter table is updated.

3. The real-time dynamic stitching method of surround view images for a vehicle based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, The water droplet-state perception field of view model is a bowl-shaped model. The bowl wall is a curved surface formed by rotating a set type curve around the central axis, the bowl bottom is a circle surrounded by the bowl wall, and the center of the bowl bottom is a rectangle projected onto the horizontal plane as a vehicle.

4. The single vehicle surround view image real-time dynamic stitching method based on water-drop state perception field fusion according to claim 3, characterized in that, The specified curve type is a parabola, circular arc, or other quadratic curve.

5. The real-time dynamic stitching method of surround view images for a vehicle based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, Feature point extraction is performed using the Superpoint architecture.

6. The real-time dynamic stitching method of surround view images for a vehicle based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, Feature matching is implemented using the SuperGlue architecture.

7. A method for real-time dynamic stitching of single-vehicle surround view images based on water droplet-state perception field fusion as described in claim 1 or 2, characterized in that, The vehicle surround vision system includes four cameras, which are respectively located above the front license plate, in the middle of the left rear door frame, above the rear license plate, and in the middle of the right rear door frame.

8. The single vehicle surround view image real-time dynamic stitching method based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, The vehicle surround vision system is based on an Ackerman motion model vehicle.

9. The real-time dynamic stitching method of surround view images for a vehicle based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, The camera imaging model adopts the Taylor expansion camera imaging model proposed by Scaramuzza.

10. The real-time dynamic stitching method of surround view images for a vehicle based on water-drop state perception field fusion according to claim 1 or 2, characterized in that, The overlapping field of view of the vehicle surround vision system faces both sides of the vehicle.