Adaptive pose loop-around stitching method, device, equipment, storage medium and product

By acquiring initial and current pose parameters in the vehicle for dynamic compensation, the problems of image misalignment and ghosting caused by changes in vehicle posture are solved, thereby improving the imaging quality and environmental perception capabilities of the surround view system.

CN122434728APending Publication Date: 2026-07-21GUANGDONG SFOUNDINT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG SFOUNDINT TECHNOLOGY CO LTD
Filing Date
2026-06-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing vehicle surround view systems suffer from camera pose drift due to drastic changes in vehicle body posture in construction machinery, resulting in image misalignment and ghosting, which affects image quality and the driver's environmental perception.

Method used

By acquiring the vehicle's pose parameters in the initial calibration state and the vehicle body posture change parameters in the current working state, dynamic compensation is performed to generate the current pose parameters. Based on these parameters, the images captured by the surround-view camera are transformed and stitched together to generate the stitched target surround-view image.

Benefits of technology

It improves the imaging reliability of vehicle surround view images, provides a more realistic and reliable perception of the vehicle's surrounding environment, eliminates image misalignment and ghosting, and enhances driver safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434728A_ABST
    Figure CN122434728A_ABST
Patent Text Reader

Abstract

The application discloses a kind of adaptive pose look-around splicing method, device, equipment, storage medium and product, it is related to image processing technical field, comprising: obtaining the initial pose parameter of each look-around camera of vehicle in initial calibration state;Obtain the vehicle body attitude change parameter of vehicle in current working state;According to vehicle body attitude change parameter, the initial pose parameter is dynamically compensated, generates the current pose parameter corresponding to each look-around camera;Based on the current pose parameter, the current look-around image collected by each look-around camera is transformed and spliced, and the target look-around image after splicing is generated.The application is dynamically compensated according to vehicle body attitude change parameter to initial pose parameter, solve the misregistration and ghosting phenomenon caused by vehicle body attitude change, improve the imaging reliability of vehicle camera look-around image, provide more real and reliable vehicle surrounding environment perception image for driver.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to adaptive pose surround view stitching methods, apparatus, devices, storage media and products. Background Technology

[0002] With the expanding application of intelligent driving technology in the construction machinery field, transplanting vehicle-mounted surround view systems (i.e., panoramic imaging systems) to equipment such as excavators and loaders to eliminate blind spots and improve operational safety has become an important development trend. Existing surround view systems typically employ the following technical solution: before the vehicle leaves the factory, a one-time calibration procedure determines the fixed height and angle of each wide-angle camera around the vehicle relative to the ground, and based on this, transform parameters are generated to project multiple images onto a bird's-eye view and stitch them together. During vehicle use, the system continuously uses these preset parameters to process and fuse the real-time acquired images to generate the surround view image. However, this existing solution based on static calibration and fixed parameters has a significant drawback: it cannot adapt to the actual camera position drift caused by drastic changes in vehicle posture under dynamic, heavy-load operating conditions. When a loader is digging materials or an excavator is traveling on bumpy roads, the vehicle body experiences significant pitch, roll, and vertical displacement, causing the ground clearance and angle of the cameras mounted on the vehicle body to change in real time. If the system still uses static calibration parameters for image projection and stitching, severe image misalignment and ghosting will occur in the generated surround view image, resulting in low imaging quality of the surround view image and seriously interfering with the driver's accurate perception of the vehicle's surrounding environment.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide an adaptive pose surround view stitching method, apparatus, device, storage medium, and product, which aims to solve the technical problem of low reliability of surround view image imaging from cameras.

[0005] To achieve the above objectives, this application proposes an adaptive pose surround view stitching method, the method comprising: Acquire the initial pose parameters of each surround view camera of the vehicle in the initial calibration state; Obtain the vehicle body attitude change parameters under the current working state; Based on the vehicle body posture change parameters, the initial pose parameters are dynamically compensated to generate the current pose parameters corresponding to each of the surround view cameras; Based on the current pose parameters, the current panoramic images captured by each of the panoramic cameras are transformed and stitched together to generate a stitched target panoramic image.

[0006] In one embodiment, the initial pose parameters include an initial rotation matrix and an initial translation matrix, and the step of obtaining the initial pose parameters of each surround-view camera of the vehicle in the initial calibration state includes: The surround-view camera is calibrated to obtain the initial homography matrix corresponding to the surround-view camera; Based on the initial homography matrix, calculate the initial rotation matrix and initial translation matrix of each of the surround-view cameras relative to the vehicle coordinate system.

[0007] In one embodiment, the current pose parameters include a current rotation matrix and a current translation matrix. The step of dynamically compensating the initial pose parameters based on the vehicle body posture change parameters to generate the current pose parameters corresponding to each of the surround-view cameras includes: Based on the vehicle body attitude change parameters, generate an attitude compensation rotation matrix; The attitude compensation rotation matrix is ​​multiplied by the initial rotation matrix to obtain the current rotation matrix; Transform the initial translation matrix to the world coordinate system to obtain the initial position of the surround-view camera in the world coordinate system; Based on the vehicle body posture change parameters, offset compensation is performed on the initial position to obtain the compensated position; The compensated position is transformed back to the camera coordinate system using the current rotation matrix to obtain the current translation matrix.

[0008] In one embodiment, the step of transforming and stitching the current panoramic images captured by each of the panoramic cameras based on the current pose parameters to generate a stitched target panoramic image includes: Based on the current rotation matrix and the current translation matrix, determine the current top-down transformation matrix corresponding to each of the surround-view cameras; Based on the current top-down transformation matrix, the current panoramic images captured by each of the panoramic cameras are converted to a top-down view to obtain the corresponding single-channel top-down image; Determine the overlapping area between the single-channel top-view images of adjacent surround-view cameras; The pixels in the overlapping area are weighted and fused according to the fusion weight, and the individual top-view images are stitched together to form the target panoramic image.

[0009] In one embodiment, the step of determining the current top-down transformation matrix corresponding to each of the surround-view cameras based on the current rotation matrix and the current translation matrix includes: Based on the current rotation matrix, current translation matrix and camera internal parameters of each of the surround-view cameras, the preset world coordinate system point set is projected onto the image pixel coordinate system of each of the surround-view cameras to generate the current pixel point set. Based on the current pixel set and the preset top-view point set, the current top-view transformation matrix corresponding to the surround-view camera is calculated.

[0010] In one embodiment, the adaptive pose surround view stitching method further includes: Obtain the pixel position of the target object to be measured in the target panoramic image; Based on the current top-down transformation matrix and the current pose parameters, determine the current mapping relationship between the image coordinate system of the target surround view image and the vehicle coordinate system of the vehicle; Based on the current mapping relationship, the pixel position of the target object to be measured is transformed to the vehicle coordinate system to obtain the actual position coordinates of the target object relative to the vehicle. Based on the actual location coordinates, calculate the actual distance between the target object to be measured and the vehicle.

[0011] Furthermore, to achieve the above objectives, this application also proposes an adaptive pose surround view stitching device, which includes: The calibration pose acquisition module is used to acquire the initial pose parameters of each surround view camera of the vehicle in the initial calibration state; The camera pose acquisition module is used to acquire the vehicle body pose change parameters in the current working state; The pose dynamic compensation module is used to dynamically compensate the initial pose parameters according to the vehicle body posture change parameters, and generate the current pose parameters corresponding to each of the surround view cameras. The surround view image generation module is used to transform and stitch together the current surround view images captured by each of the surround view cameras based on the current pose parameters to generate a stitched target surround view image.

[0012] In addition, to achieve the above objectives, this application also proposes an adaptive pose surround view stitching device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the adaptive pose surround view stitching method as described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the adaptive pose surround stitching method described above.

[0014] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the adaptive pose surround view stitching method described above.

[0015] One or more technical solutions proposed in this application have at least the following technical effects: This application proposes an adaptive pose surround view stitching method, apparatus, device, storage medium, and product. The method involves acquiring the initial pose parameters of each surround view camera in an initial calibration state of a vehicle; acquiring the vehicle's body posture change parameters in the current operating state; dynamically compensating the initial pose parameters based on the body posture change parameters to generate the current pose parameters corresponding to each surround view camera; and transforming and stitching the current surround view images acquired by each surround view camera based on the current pose parameters to generate a stitched target surround view image. This application solves the stitching misalignment and ghosting phenomena caused by changes in vehicle body posture by dynamically compensating the initial pose parameters based on the body posture change parameters, improving the imaging reliability of the vehicle camera surround view images, and providing drivers with more realistic and reliable images of the vehicle's surrounding environment. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating an embodiment of the adaptive pose surround view stitching method of this application. Figure 2 A flowchart illustrating the adaptive pose surround view stitching method provided in this application embodiment; Figure 3 This is a schematic diagram of the module structure of the adaptive pose surround view stitching device according to an embodiment of this application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the adaptive pose surround view stitching method in the embodiments of this application.

[0019] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0021] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0022] The main solution of this application embodiment is: to obtain the initial pose parameters of each surround view camera in the initial calibration state of the vehicle; to obtain the vehicle body posture change parameters in the current working state of the vehicle; to dynamically compensate the initial pose parameters according to the vehicle body posture change parameters, and generate the current pose parameters corresponding to each surround view camera; and to transform and stitch the current surround view images collected by each surround view camera based on the current pose parameters, and generate the stitched target surround view image.

[0023] In this embodiment, for ease of description, the panoramic splicing system will be used as the execution subject for the following description.

[0024] With the expanding application of intelligent driving technology in the construction machinery field, transplanting vehicle-mounted surround view systems (i.e., panoramic imaging systems) to equipment such as excavators and loaders to eliminate blind spots and improve operational safety has become an important development trend. Existing surround view systems typically employ the following technical solution: before the vehicle leaves the factory, a one-time calibration procedure determines the fixed height and angle of each wide-angle camera around the vehicle relative to the ground, and based on this, transform parameters are generated to project multiple images onto a bird's-eye view and stitch them together. During vehicle use, the system continuously uses these preset parameters to process and fuse the real-time acquired images to generate the surround view image. However, this existing solution based on static calibration and fixed parameters has a significant drawback: it cannot adapt to the actual camera position drift caused by drastic changes in vehicle posture under dynamic, heavy-load operating conditions. When a loader is digging materials or an excavator is traveling on bumpy roads, the vehicle body experiences significant pitch, roll, and vertical displacement, causing the ground clearance and angle of the cameras mounted on the vehicle body to change in real time. If the system still uses static calibration parameters for image projection and stitching, severe image misalignment and ghosting will occur in the generated surround view image, resulting in low imaging quality of the surround view image and seriously interfering with the driver's accurate perception of the vehicle's surrounding environment.

[0025] This application provides a solution that dynamically compensates the initial pose parameters based on the vehicle body posture change parameters, thereby solving the stitching misalignment and ghosting phenomena caused by changes in vehicle body posture, improving the imaging reliability of the vehicle camera surround view image, and providing the driver with a more realistic and reliable image of the vehicle's surrounding environment.

[0026] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or surround view splicing system capable of performing the above functions. The following description uses a surround view splicing system as an example to illustrate this embodiment and the subsequent embodiments.

[0027] Based on this, embodiments of this application provide an adaptive pose surround view stitching method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the adaptive pose surround view stitching method of this application.

[0028] In this embodiment, the adaptive pose surround view stitching method includes steps S11 to S14: Step S11: Obtain the initial pose parameters of each surround view camera of the vehicle in the initial calibration state.

[0029] It should be noted that the initial calibration state refers to the baseline condition where the vehicle is level, stationary, and unloaded, such as parked on a flat workshop floor. The surround-view cameras are wide-angle cameras mounted around the vehicle body to capture a 360-degree field of view. The initial pose parameters are a mathematical model describing the precise spatial position and orientation of each camera relative to the vehicle coordinate system (a reference system fixed to the vehicle body, with its origin, for example, at the center of the rear axle) in the baseline state. This parameter is the mathematical basis for all subsequent image geometric transformations and is the sole reference for determining whether camera drift has occurred. The pose parameters are specifically expressed as an RT matrix consisting of a rotation matrix R and a translation matrix T.

[0030] Understandably, by calibrating under stable vehicle posture and unchanged camera position, the standard position and posture that each camera should be in can be obtained. Without this precise and repeatable initial reference, subsequent dynamic compensation will have no way to determine the relative magnitude of the posture change, and therefore cannot achieve effective correction.

[0031] Specifically, this step is executed by the processing unit in the vehicle surround view system. Its implementation includes: before the vehicle leaves the factory or during maintenance, placing the vehicle in an initial calibration state; placing a checkerboard calibration cloth of known size on the ground around the vehicle; and guiding each surround view camera to acquire multiple images containing the pattern of the calibration cloth. Using computer vision algorithms (such as the Zhang Zhengyou calibration method or its variants), the calibration software can calculate the intrinsic parameters (focal length, principal point, distortion coefficients) of each camera, as well as the initial top-view transformation matrix (homography matrix H). For further coordinate transformation, it is usually necessary to decompose the homography matrix to obtain richer three-dimensional spatial pose information, namely the rotation matrix R and the translation matrix T.

[0032] For example, a loader is parked in the calibration workshop. Technicians lay out a large checkerboard calibration cloth around the vehicle. The surround view processing unit controls four cameras (front, rear, left, and right) to take pictures, and by identifying the corner points of the checkerboard, calculates a 3x3 initial homography matrix for each camera. This matrix defines the direct projection relationship between the ground plane and the original image from the camera. This homography matrix, or its further decomposition into the RT matrix, constitutes the initial pose parameters.

[0033] Step S12: Obtain the vehicle body attitude change parameters in the current working state.

[0034] It should be noted that vehicle body attitude change parameters refer to physical quantities describing the attitude differences of the vehicle chassis or body relative to its initial calibration state. The current operating state is the dynamic condition of the vehicle during actual operation or work (such as bumpy driving or digging). These parameters are collected in real time by attitude sensors (such as high-precision IMUs) deployed on the vehicle body, and the core parameters include pitch angle, roll angle, and vertical displacement, which are used to quantify the angular and positional offset of the vehicle body in three-dimensional space.

[0035] Understandably, since the camera is rigidly mounted on the vehicle body, its spatial pose changes synchronously with the vehicle's attitude. Only by accurately capturing how the vehicle moves (i.e., attitude change parameters) can the system accurately infer what offset the camera has made, thus providing the most direct input data for the next step of dynamic compensation. This is a prerequisite for achieving adaptive perception.

[0036] Specifically, data from the high-precision inertial measurement unit (IMU) is read via a bus (such as CAN). The IMU contains a built-in gyroscope and accelerometer, which can directly measure the vehicle's angular velocity and linear acceleration. The processing unit integrates, filters, and transforms the raw data to calculate the Euler angle change relative to the initial horizontal state at the current moment. and displacement changes ,in This represents the change in pitch angle. This represents the change in roll angle. This refers to the change in yaw angle. In engineering machinery scenarios, pitch angle and vertical displacement are key parameters. These calculated changes constitute the parameters of vehicle attitude change.

[0037] For example, when the loader is scooping material, its front axle sinks due to the load. An IMU mounted on the chassis detects in real time a 5-degree forward tilt (pitch angle change Δθ = 5°) and a 3-centimeter vertical drop (ΔZ = 5°). (0.03m). The surround view processing unit reads the attitude packet calculated by the IMU at a frequency of 100Hz to obtain the current Δθ and ΔZ values.

[0038] Step S13: Based on the vehicle body posture change parameters, dynamically compensate the initial pose parameters to generate the current pose parameters corresponding to each of the surround view cameras.

[0039] It should be noted that dynamic compensation is a real-time, online mathematical calculation process. Its core is to correct the stored static initial pose parameters based on the real-time perceived changes in the vehicle's attitude. Generating the current pose parameters means calculating a new set of parameters for each camera that reflects its true spatial position and orientation under the current actual vehicle attitude. This set of parameters is dynamically updated.

[0040] Understandably, through a set of preset coordinate transformation and matrix operation rules (dynamic compensation algorithm), the system accurately synchronizes changes in the physical world (vehicle posture) to updates in the virtual model (camera pose parameters). This makes the geometric model used for image processing no longer rigid and fixed, but able to move with the vehicle body, thus ensuring that no matter how the vehicle bumps or tilts, the position of the virtual eye used for image projection always remains consistent with the actual physical camera position, providing a possibility for eliminating geometric distortion at its source.

[0041] Specifically, firstly, the acquired vehicle attitude change parameters (pitch angle change) are... Roll angle change yaw angle change Convert them into rotation matrices about the X-axis, Y-axis, and Z-axis, respectively. , , in:

[0042]

[0043]

[0044] Furthermore, the vehicle body attitude change parameters are converted into attitude compensation rotation matrices. The conversion formula is: Then, the attitude compensation rotation matrix is ​​compared with the initial rotation matrix. Multiply to obtain the current rotation matrix. Next, the initial translation matrix is... Perform coordinate transformation and displacement compensation. First, transform it to the world coordinate system to obtain... Then add displacement offset get Finally, using the new Transform it back to the camera coordinate system to obtain the current translation matrix. Obtain the current pose parameters .

[0045] For example, the system acquires the initial pose of the left front camera. According to the pitch angle changes transmitted from the IMU Generate a matrix that rotates around the Y-axis. Calculate the new rotation matrix. Simultaneously, calculation Plus vertical displacement get Finally, calculate At this point, the updated pose of the left front camera is obtained.

[0046] Step S14: Based on the current pose parameters, transform and stitch the current panoramic images captured by each of the panoramic cameras to generate a stitched target panoramic image.

[0047] It should be noted that transformation and stitching based on the current pose parameters refers to using the latest and most accurate generated camera spatial parameters to perform geometric correction and fusion on the real-time acquired raw surround view image. Transformation specifically refers to the projection transformation from the original perspective view of the camera to a unified bird's-eye view (top view). Stitching, on the other hand, involves aligning and fusing multiple sub-images from different cameras, which have been transformed to the same top view plane, according to their corresponding physical spatial positions, ultimately generating a complete target surround view image that encircles the vehicle.

[0048] Understandably, this step is the final presentation stage of the dynamic compensation effect. Its purpose is to use updated and accurate geometric model parameters to perform correct projection transformation on each frame of real-time image, thereby generating a geometrically accurate and coherent panoramic image. Because the pose parameters used for transformation are updated in real time and strictly consistent with the current physical state, even if the vehicle body shakes violently, the transformed sub-images can be precisely aligned on the top-view plane, and the overlapping areas of adjacent images can be merged according to the correct geometric relationships. This fundamentally eliminates the image misalignment, ghosting, and visual fragmentation that inevitably occur under static parameter models, achieving high-quality visual output.

[0049] Specifically, firstly, for each camera, using its current pose parameters and camera intrinsic parameters K, the projection relationship from the ground plane (Z=0) to the image plane of that camera is recalculated, thus obtaining a new homography matrix. Then, this new homography matrix is ​​used to perform a perspective transformation on the original image of the current frame, which has already undergone past distortion processing, to obtain the top view of the camera. Next, based on the updated spatial relationships of each camera, the system determines the overlapping areas between their top views and uses an adaptive fusion algorithm (such as fade-in / fade-out weighted fusion) to smooth the stitching seams. Finally, all the top views are combined according to a preset layout to form the final panoramic bird's-eye view.

[0050] For example, the system utilizes the latest left front camera and internal reference Calculate the new homography matrix based on the projection model. A single frame of image captured in real-time by the left front camera is transformed into a top-down view using a new homography matrix. The same operation is performed on the other three cameras. Since all transformation matrices are calculated based on the vehicle's current pitch of 5 degrees, these four top-down views are precisely aligned during synthesis, ultimately generating a seamless, ghost-free, complete panoramic image.

[0051] This embodiment, through the aforementioned scheme, first establishes a precise geometric benchmark by acquiring initial pose parameters. Then, it captures the root cause of camera pose drift by real-time sensing of vehicle body posture changes. Next, it dynamically compensates for the initial pose based on posture changes, enabling the camera model used for image projection to adjust in real-time and synchronously with changes in the vehicle body's physical posture, ensuring consistency between virtual geometry and the real world. Finally, based on this real-time updated accurate model, the images are transformed and stitched, ensuring that each output frame is generated based on the geometric parameters best matching the current moment. This logical closed loop of steps ensures that regardless of the vehicle body's pitch or roll, the system can instantly correct the projection parameters, thereby eliminating image misalignment and ghosting caused by parameter lag at the source. Ultimately, it provides operators with stable, coherent, and reliable panoramic surround-view images, significantly improving operational safety and confidence under complex dynamic conditions.

[0052] Based on the above implementation scheme, in one feasible implementation, the initial pose parameters include an initial rotation matrix and an initial translation matrix, and the step of obtaining the initial pose parameters of each surround-view camera of the vehicle in the initial calibration state includes S21~S22: Step S21: Calibrate the surround-view camera and obtain the initial homography matrix corresponding to the surround-view camera.

[0053] It should be noted that calibration is the process of determining the geometric model of camera imaging in computer vision. The initial homography matrix H is a 3x3 matrix. In the context of a surround-view system, it specifically refers to the matrix describing the two-dimensional projection transformation relationship between a point (X,Y,0) on the ground plane (the plane with Z=0 in the world coordinate system) and its corresponding pixel (u,v) in a specific surround-view camera image. It encapsulates a direct linear mapping relationship from the top-view (bird's-eye view) coordinates to the original image coordinates of the camera, satisfying:

[0054] in It is a scale factor.

[0055] Understandably, the purpose of this step is to obtain a mathematical tool that can directly link the coordinates of the 2D top view with the coordinates of the original images from each camera. The homography matrix H is the most efficient mathematical tool for implementing top-view transformation in surround view stitching because it directly establishes a mapping from one 2D plane to another. Obtaining the initial homography matrix not only provides the most basic data that can be directly used for image processing in subsequent steps, but also serves as the starting point and key input for further decomposing richer 3D pose information (rotation and translation matrices).

[0056] Specifically, in the initial calibration state, one or more calibration objects with known physical dimensions and patterns (such as a large checkerboard calibration cloth) are placed on the ground around the vehicle. This ensures that each surround-view camera can capture a clear image of the calibration object. After the system acquires the images, it automatically detects feature points (such as corner points) of the calibration objects in the images and uses these feature points to establish multiple pairs of corresponding points between their two-dimensional pixel coordinates (u,v) in the image and their known three-dimensional coordinates (X,Y,0) on the world coordinate system ground plane. By solving an overdetermined system of equations (such as using the Direct Linear Transform (DLT) algorithm, or a robust estimation algorithm combined with RANSAC), an optimal 3x3 homography matrix can be calculated for each camera.

[0057] For example, in the calibration workshop, a vehicle coordinate system is established with the vehicle body as the center. Four 2m*2m checkerboard calibration cloths are precisely placed around the vehicle on the ground. The system controls four cameras to take pictures simultaneously and automatically identifies the pixel coordinates of all corner points in each picture. Since the (X,Y,0) coordinates of each corner point in the vehicle coordinate system are known, the system matches more than 50 such pairs of points for the left front camera and obtains the initial homography matrix of the left front camera by solving for the coordinates.

[0058] Step S22: Based on the initial homography matrix, calculate the initial rotation matrix and initial translation matrix of each of the surround-view cameras relative to the vehicle coordinate system.

[0059] It should be noted that the initial rotation matrix is ​​a 3x3 orthogonal matrix describing the rotation (orientation) relationship of the surround-view camera's coordinate system relative to the vehicle coordinate system. The initial translation matrix is ​​a 3x1 vector describing the position coordinates of the camera's optical center in the vehicle coordinate system. Together, they constitute the camera's extrinsic parameter matrix [R|T], which fully defines its three-dimensional pose. Decomposing the rotation and translation matrices [R,T] from the homography matrix H is a process of inferring the three-dimensional pose from the two-dimensional projection relationship, which requires knowledge of the camera's intrinsic parameter matrix K.

[0060] Understandably, the purpose of this step is to elevate the two-dimensional homography matrix H, which is convenient for image transformation but limited in information, into pose parameters (rotation matrix R and translation vector T) containing complete three-dimensional spatial information. The homography matrix H only reflects the relationship between the specific plane of the ground plane and the image plane, while the rotation matrix R and translation vector T describe the complete pose of the camera in three-dimensional space. This is crucial for subsequent dynamic pose compensation based on three-dimensional spatial calculations. Because changes in vehicle body posture (such as pitch and vertical displacement) occur in three-dimensional space, compensation calculations must be based on the three-dimensional pose parameters (R, T) to perform correct and direct mathematical operations.

[0061] Specifically, the camera imaging model is as follows:

[0062] Based on the obtained initial single-valued matrix ,and ,in, and These are the first two columns of the initial rotation matrix R. Let T be the translation vector. Invert the known camera intrinsic parameter matrix K and calculate the intermediate matrix. Perform Schmidt orthogonalization and normalization on the first two columns of M to obtain Then calculate the third column using cross product. Thus, the initial rotation matrix is ​​obtained by assembly. Due to the influence of noise, the resulting initial rotation matrix... They may not be strictly orthogonal; orthogonality can be corrected through SVD decomposition. Translation vector. It can be done Estimate, or through The modulus length is normalized, where for The third column.

[0063] This embodiment employs the above-described scheme, first calibrating the camera to obtain an initial homography matrix, then calculating an initial rotation and translation matrix based on this homography matrix, thus linking the two-dimensional parameters used for image transformation in the surround-view system with the three-dimensional parameters used for three-dimensional spatial compensation. This specific step makes the acquisition of initial pose parameters clear and feasible, and provides a mathematical foundation (RT matrix) for direct three-dimensional spatial calculations for subsequent dynamic compensation, thereby ensuring the accuracy of the starting point of the entire adaptive update process and its compatibility with subsequent steps.

[0064] Based on the above implementation scheme, in one feasible implementation, the current pose parameters include a current rotation matrix and a current translation matrix, and the step of dynamically compensating the initial pose parameters according to the vehicle body posture change parameters to generate the current pose parameters corresponding to each of the surround-view cameras includes S31~S35: Step S31: Generate an attitude compensation rotation matrix based on the vehicle body attitude change parameters.

[0065] It should be noted that the attitude compensation rotation matrix It is a 3x3 rotation matrix that mathematically represents the pure rotational change of the vehicle body from its initial calibration state to its current operating state. It is calculated from the acquired vehicle body attitude change parameters, which are expressed in terms of Euler angles (pitch angle change). Roll angle change yaw angle change )express.

[0066] Understandably, the purpose of this step is to convert the easily understood physical angle changes perceived by the sensors into a mathematical form that can be used for matrix operations. Since the cameras are rigidly connected to the vehicle body, a rotation of the vehicle's coordinate system is equivalent to all camera coordinate systems fixed to the vehicle undergoing the same rotation. Therefore, generating an attitude compensation matrix Roffset representing the overall rotation of the vehicle body is a crucial input for the subsequent synchronous update of the rotation matrix for each camera. The principle is to synthesize discrete Euler angle changes into a single matrix representing the composite rotation.

[0067] Specifically, based on the Euler angle variation provided by the IMU Construct basic rotation matrices around each axis according to a predetermined rotation sequence (e.g., first heading, then pitch, then roll). , , Then, these basic rotation matrices are multiplied sequentially to obtain the final vehicle body attitude compensation rotation matrix. * * .

[0068] Step S32: Multiply the attitude compensation rotation matrix with the initial rotation matrix to obtain the current rotation matrix.

[0069] It should be noted that the multiplication process refers to matrix multiplication. The attitude compensation rotation matrix, representing the overall rotation of the vehicle body, is used... , and the initial rotation matrix representing the orientation of the camera in the initial vehicle coordinate system. Multiply to obtain the current rotation matrix of the camera under the current vehicle body posture. The calculation formula is: .

[0070] Understandably, the camera's orientation in the initial vehicle coordinate system is defined by the initial rotation matrix. When the vehicle coordinate system itself undergoes a rotation defined by the attitude compensation rotation matrix, for a camera fixed to the vehicle (i.e., rotating with the vehicle), its orientation in the new (current) vehicle coordinate system is equal to its initial orientation. The initial rotation matrix is ​​first transformed, and then the effect of the vehicle rotation is added. In matrix representation, this is... This calculation ensures that the orientation model of each camera can be updated synchronously and correctly in real time as the vehicle body rotates, reflecting the physical fact that the cameras are moving.

[0071] Specifically, for each surround-view camera, matrix multiplication is performed: ,in Representing the One camera.

[0072] Step S33: Transform the initial translation matrix to the world coordinate system to obtain the initial position of the surround-view camera in the world coordinate system.

[0073] It should be noted that the world coordinate system refers to a fixed global reference system that coincides with the vehicle coordinate system at the initial calibration time. The initial position refers to the three-dimensional coordinate vector of the camera's optical center within this fixed world coordinate system. Due to the initial translation matrix It is expressed in the camera coordinate system, representing the vector from the origin of the vehicle coordinate system to the optical center of the camera, and its coordinates in the camera coordinate system. It needs to be transformed to the world coordinate system in order to be superimposed with the offset describing the movement of the vehicle body in the world coordinate system.

[0074] Understandably, the purpose of this step is to unify the camera's position representation to a common, static reference frame (world coordinate system). Translation vector The representation depends on the orientation of the camera coordinate system. To compensate for changes in camera position caused by vehicle displacement (such as vertical sway), we need to know the camera's absolute position in the world, then apply an offset to this absolute position, and finally transform it back into a representation related to the current vehicle posture. This is a necessary preparatory step for dynamic compensation of the camera position, involving an inverse transformation from the camera coordinate system to the world coordinate system.

[0075] Specifically, the initial pose parameters are known. And the transformation from the camera coordinate system to the vehicle coordinate system (initially coinciding with the world coordinate system). For points in the world coordinate system The coordinates Pc in the camera coordinate system satisfy: The camera's optical center is at the origin (0) in the camera coordinate system. Find its position in the world coordinate system. That is, solving the equation .therefore, Since the inverse of a rotation matrix is ​​equal to its transpose, i.e. ,so .

[0076] Step S34: Based on the vehicle body posture change parameters, offset compensation is performed on the initial position to obtain the compensated position.

[0077] It should be noted that offset compensation refers to adjusting the initial position of the camera in the world coordinate system based on the displacement components described in the vehicle body attitude change parameters (mainly vertical displacement ΔZ, and possibly longitudinal ΔX and lateral ΔY). Perform vector addition. (Compensated position) This is the estimated position of the camera in the world coordinate system at the current moment after the vehicle body has moved. The calculation formula is as follows: ,in

[0078] Understandably, the purpose of this step is to compensate for changes in the spatial position of the camera caused by the overall translation of the vehicle body (such as upward, downward, forward, backward, left, and right movement). Sensors such as IMUs can provide not only angular changes but also positional changes through integration. As the camera is part of the vehicle body, its spatial position moves along with the vehicle body. This is addressed by directly applying an offset vector equal to the vehicle body's displacement to the camera's position in the world coordinate system. This method can accurately simulate this physical process, thereby updating the camera's position information. The principle is rigid body translation, meaning that all points on the vehicle body (including the camera) undergo the same displacement.

[0079] Specifically, displacement components (ΔX, ΔY, ΔZ) are extracted from the vehicle body attitude change parameters to form a displacement vector. Then, this displacement vector The calculated initial world coordinates of the camera Add: This assumes the vehicle body is a rigid body and all points have the same displacement. The calculation is performed once per camera in each processing cycle.

[0080] Step S35: Transform the compensated position back into the camera coordinate system using the current rotation matrix to obtain the current translation matrix.

[0081] It should be noted that this step is the inverse transformation of position compensation. It adjusts the compensated camera position in the world coordinate system. Transform back to the current camera coordinate system to obtain the current translation matrix. The current camera coordinate system here is determined by the current rotation matrix. Defined coordinate system. The calculation formula is: .

[0082] Understandably, the purpose of this step is to ultimately complete the dynamic update of the translation vector, making it consistent with the updated rotation matrix. These parameters are matched and together form a self-consistent and complete set of current pose parameters under the current vehicle body posture. , In the camera imaging model, the translation vector T and the rotation matrix R need to be defined in the same coordinate system (i.e., describing the transformation from the world coordinate system to the current camera coordinate system). Therefore, to obtain the new position in the world coordinate system... Then, a new rotational relationship needs to be utilized. This can be expressed as a vector from the current vehicle origin to the current camera position, with coordinates in the current camera coordinate system. This calculation completes the mapping of position parameters from the world coordinate system to the new camera coordinate system.

[0083] Specifically, the coordinate transformation relationship is as follows: for a point in the world coordinate system... Coordinates in the current camera coordinate system .when When it is the optical center of the camera itself, that is ,but 0 (camera origin). Therefore, .

[0084] This embodiment employs the above-described scheme, first generating an attitude compensation rotation matrix to quantify the vehicle body rotation, then multiplying it by the initial rotation matrix to update the camera orientation; next, updating the camera position through a three-step process: transforming the initial translation matrix to the world coordinate system, superimposing displacement offsets, and then transforming back to the current camera coordinate system. This series of steps strictly adheres to the principles of three-dimensional rigid body kinematics and coordinate transformation, accurately mapping the physical changes in the vehicle body's attitude (angles and displacements) to the mathematical updates of each camera's pose parameters. This specific algorithm ensures that regardless of the vehicle body's movement, the system can calculate an accurate camera spatial model in real time, providing a solid algorithmic foundation for generating geometrically distortion-free panoramic images.

[0085] Based on the above implementation scheme, in one feasible implementation, the step of transforming and stitching the current panoramic images captured by each of the panoramic cameras based on the current pose parameters to generate the stitched target panoramic image includes S41~S44: Step S41: Determine the current top-down transformation matrix corresponding to each of the surround-view cameras based on the current rotation matrix and the current translation matrix.

[0086] It should be noted that the current top-down transformation matrix is ​​a 3x3 homography matrix. It defines the two-dimensional projection transformation relationship between a point (X,Y,0) on the ground plane (world coordinate XY plane, Z=0) and a pixel (u,v) in the current surround-view camera image. Similar to the initial homography matrix, but it is based on dynamically updated current pose parameters. The recalculated value reflects the projected geometry of the camera in its current actual pose. Its calculation depends on the camera's intrinsic parameter K.

[0087] Understandably, the purpose of this step is to repackage the dynamic pose compensation results in 3D space into a tool for fast and efficient transformation of 2D images, namely the homography matrix H. Although we have fully described the camera pose using 3D R and T, the most effective and direct method to ultimately transform the original wide-angle image into a bird's-eye view is still to use a 2D homography matrix. This step is based on the latest and most accurate... Recalculate this transformation matrix This ensures that subsequent image transformations are based on a geometric model that best matches the current physical state, which is a key bridge to achieving high-quality dynamic stitching.

[0088] Specifically, the processing unit performs this step. The implementation process is as follows: for each camera, using its current pose parameters... Given the camera intrinsic parameter matrix K, and based on the pinhole camera model, calculate the homography matrix from the ground plane (Z=0) to the image plane. The derivation formula is as follows:

[0089] in, and It is the current rotation matrix The first two columns. Therefore, the current top-down transformation matrix. Direct extraction The first two columns, and By combining the results and then left-multiplying by the intrinsic parameter K, we can obtain the result. .

[0090] Step S42: Based on the current top-down transformation matrix, the current panoramic images captured by each of the panoramic cameras are converted to a top-down view to obtain the corresponding single-channel top-down image.

[0091] It should be noted that the top-down view refers to a virtual perspective looking vertically downwards from directly above the vehicle, and the resulting image is called a bird's-eye view or top-down view. A single-channel top-down view image refers to a rectangular image reflecting the ground plane within the field of view of a single surround-view camera after being mapped by the aforementioned current top-down transformation matrix. The current surround-view image is the raw frame captured by the camera in real time at the current moment.

[0092] Understandably, the purpose of this step is to use the latest and most accurate projection relationship, i.e., the current top-down transformation matrix, to perform geometric correction on each frame of the original panoramic image acquired in real time, unifying them onto a common two-dimensional top-down plane. Since the current top-down transformation matrix is ​​dynamically updated, it can accurately compensate for the differences in perspective distortion in the original image caused by changes in vehicle body posture. This ensures that single-path top-down images generated from different cameras at different times have consistent geometric scale and correct alignment on the common top-down plane, laying a solid geometric foundation for subsequent seamless stitching.

[0093] Specifically, the operation is performed on a single frame of raw image simultaneously captured by each camera. First, the raw image is distorted to obtain a perspective image conforming to the pinhole model. Then, each pixel of the top-view canvas of the target is processed. Using the current top-down transformation matrix inverse matrix Calculate the corresponding sampling points in the original distorted image. ,Right now:

[0094] Finally, sampling points are obtained from the original distorted image using bilinear interpolation and other resampling methods. The pixel values ​​of the position are used to fill the top view. Location. This process traverses the entire top-view area of ​​the target.

[0095] Step S43: Determine the overlapping area between the single-channel top-view images of adjacent surround-view cameras.

[0096] It should be noted that the overlapping area refers to the portion of the image corresponding to the same real ground area in the single-channel top-down image generated by adjacent surround-view cameras (such as front-left, left-rear, etc.) due to the physical overlap of their fields of view. This area is the processing area for subsequent image fusion to eliminate stitching seams and achieve a smooth transition.

[0097] Understandably, the purpose of this step is to precisely define the image boundary regions that require special processing. The goal of the surround-view system is seamless stitching, and a smooth, seamless transition comes from proper fusion in the overlapping areas. Since the camera poses are dynamically changing, theoretically, their field-of-view overlap range will also change slightly. Therefore, it is necessary to redefine or verify this overlapping area based on the current pose parameters in each processing cycle to ensure the accuracy of the fusion operation and avoid new artificial edges or information loss caused by mismatch between the fused area and the actual overlapping area in dynamic situations. However, in practical applications, as long as the dynamic compensation is accurate enough, the change in the overlapping area is very small, and a fixed overlapping area mask pre-calculated based on the initial pose can usually be used.

[0098] Specifically, based on the current top-down transformation matrix of each camera. The effective field of view (the mapped area in the top view) and its effective field of view are determined through geometric calculations (such as calculating the intersection of the polygons of the effective top view areas of two cameras) to determine the common area of ​​the top view of each pair of adjacent cameras. A more common and efficient method is to pre-calculate and store the overlapping area mask (binary image) of each pair of adjacent cameras based on the initial pose parameters during system initialization. During dynamic operation, since dynamic compensation ensures the accuracy of projection, this pre-calculated mask area still corresponds to the correct physical overlapping ground, so it can be directly reused without recalculating every frame, thus saving computational resources.

[0099] Step S44: Perform weighted fusion on the pixels in the overlapping area according to the fusion weight, and stitch the single-path top view images into the target panoramic view image.

[0100] It's important to note that the fusion weight is a value between 0 and 1, assigned to each pixel within the overlapping region. It indicates which camera's single-view top-down image should contribute how much to the pixel's color value in the final synthesized target panoramic image. The fusion weight is dynamically calculated based on the distance from the pixel to the stitching seam (or to the optical center / vanishing point of each camera), using a function (usually a linear gradient or a Gaussian / sine smoothing function). Weighted fusion is an image processing operation where, for each pixel location within the overlapping region, the pixel values ​​(such as RGB or luminance values) from the single-view top-down images of different cameras at that location are weighted and summed according to a determined normalized fusion weight combination to obtain the final synthesized image's pixel value at that location. For non-overlapping regions, the pixel value is directly taken from the top-down view of the camera that uniquely covers that region.

[0101] Understandably, this step is the final stitching step, and its purpose is to generate a visually smooth, seamless, and overall high-quality final surround view image. By using dynamically adjusted, image quality-based weights for fusion, rather than simple averaging, it ensures that at the stitching seams, the image not only transitions naturally in color and brightness, but also achieves an optimal combination of content sharpness and detail retention. This further enhances the overall image quality of the surround view system output. Even if the image edge quality of a particular camera deteriorates due to its current posture, the weight adjustment automatically reduces its contribution, ensuring the final image is clear, coherent, and reliable.

[0102] Specifically, for each pixel coordinate (x, y) in the overlapping region of the target panoramic image canvas, there are N overlapping images from cameras. Let the pixel value of the i-th camera at this position be... Its fusion weight is satisfy The final synthesized pixel value This process iterates through all overlapping pixels. For pixels in non-overlapping areas, the pixel values ​​from the corresponding unique camera are directly copied. Finally, after all areas are filled, a complete panoramic image of the target is obtained.

[0103] This embodiment, through the above-described scheme, first determines a new top-view transformation matrix based on the current pose, achieving real-time updates to the geometric correction model of the original image; then, it generates an aligned top view through image transformation; next, it identifies overlapping areas and dynamically adjusts the fusion weights; finally, it performs weighted fusion. This series of steps ensures that the stitching process can fully utilize the precise geometric alignment advantages brought by dynamic compensation, and further optimizes the visual quality of the stitched area through an adaptive fusion strategy. This specific method enables the surround-view system not only to overcome geometric misalignment under dynamic conditions, but also to produce a final image with better visual effects and smoother transitions, thereby comprehensively improving the practicality and user experience of the surround-view system.

[0104] Based on the above implementation scheme, in a feasible implementation, the step of determining the current top-down transformation matrix corresponding to each of the surround-view cameras according to the current rotation matrix and the current translation matrix includes S51~S52: Step S51: Based on the current rotation matrix, current translation matrix and camera internal parameters corresponding to each of the surround-view cameras, project the preset world coordinate system point set to the image pixel coordinate system corresponding to each of the surround-view cameras to generate the current pixel point set.

[0105] It should be noted that the preset world coordinate system point set is a group of points with known three-dimensional coordinates in the selected world coordinate system (usually coinciding with the ground plane, Z=0). For example, a two-dimensional grid of points uniformly covering the effective surface around a vehicle is represented as follows: The image pixel coordinate system refers to a two-dimensional coordinate system with the top-left corner of the image as the origin. The current pixel set is the set of predicted two-dimensional pixel coordinates obtained by projecting the aforementioned world point set onto the current camera image through the camera imaging model. Its calculation requires the current pose parameters. And camera internal parameters K.

[0106] Understandably, the purpose of this step is to use the updated, accurate camera model to perform a forward projection, simulating the precise positions where ground grid points should appear on the camera image under the current actual pose. This establishes a set of world-to-image point correspondences from the 3D physical world to the 2D image, strictly matching the current pose. Through this simulated projection, we obtain a set of accurate, dynamically generated corresponding point pairs, which is the high-quality input data necessary for the next step of recalculating the homography matrix. Because the pose parameters are dynamically updated, this projection process is also dynamic, ensuring that the generated corresponding point pairs always reflect the latest and correct geometric relationships.

[0107] Specifically, firstly, during system initialization, a point set W is defined that is uniformly sampled on the ground around the vehicle (world coordinate system XY plane). For each surround-view camera i, in each processing cycle, its current pose parameters are used... and camera intrinsic parameter matrix According to the camera imaging model formula in the document:

[0108] For each world point (X,Y,0) in W, calculate its projected pixel coordinates (u,v). The result is calculated using this formula. Get the current pixel set Since the world point Z=0, this projection is actually a homography transformation.

[0109] Step S52: Based on the current pixel set and the preset top view point set, calculate the current top view transformation matrix corresponding to the surround view camera.

[0110] It should be noted that the preset top-view point set refers to the set of pixel coordinates of the aforementioned world coordinate system point set W on the final, standard top-view (bird's-eye view) image. Typically, the goal is to generate a regular, top-down view, so this point set is formed by scaling the (X,Y) coordinates of these world points according to a preset scaling ratio. (pixels / meter), directly linearly mapped to the pixel coordinates of the top view, i.e. The current top-down transformation matrix is ​​obtained by solving for the current set of pixels. (Source point) and preset top viewpoint set The homography matrix is ​​obtained by determining the correspondence between (target points).

[0111] Understandably, the purpose of this step is to solve for the 2D projection transformation matrix from the original camera image to the target's standard top view using a set of precise, dynamically generated pairs of corresponding points. This is the standard method for calculating the homography matrix: given at least four sets of non-collinear corresponding points, a 3x3 homography matrix H can be solved such that... Due to our current set of pixels The homography matrix is ​​obtained by projecting dynamically updated 3D pose parameters, and therefore the homography matrix is ​​calculated accordingly. It contains all the information for dynamic compensation, and can accurately transform the image acquired at the current moment to a standard top-down view, and precisely align it with the physical world.

[0112] Specifically, for each camera i, we already have the current set of pixels. At the same time, there is a pre-set set of overhead viewpoints. It is achieved by multiplying the X and Y coordinates of a preset world coordinate system point set W by a preset scaling factor. Calculated as follows: (May include origin offset). Multiple one-to-one point pairs are now available: Using these point pairs (typically dozens or hundreds), the homography matrix is ​​solved through direct linear transformation algorithms or robust estimation algorithms (such as random sampling consensus algorithms). This minimizes the reprojection error. That is the required current top-down transformation matrix.

[0113] This embodiment solves the homography matrix by projecting a preset world coordinate point using the current pose parameters to obtain the current pixel set, and then establishing a correspondence between this set and a preset top-view point set. This method does not rely directly on the analytical derivation from the RT matrix to the H matrix, but instead uses numerical calculation. The advantages of this method are: firstly, it can utilize a large number of point pairs for least-squares optimization, resulting in a more robust current top-view transformation matrix that can resist minor noise in pose parameter calculations; secondly, the process is clear, consistent with general homography matrix solving procedures, and easy to implement and verify; and thirdly, it ensures that the final matrix used for image transformation is strictly consistent with the pose parameters used to generate the corresponding points, further guaranteeing the closed-loop accuracy of the entire dynamic compensation chain.

[0114] Based on the above implementation scheme, in one feasible implementation, the adaptive pose surround view stitching method further includes steps S61 to S64: Step S61: Obtain the pixel position of the target object to be measured in the target panoramic image.

[0115] It should be noted that the target object to be measured refers to the object identified in the target surround view image whose distance from the vehicle needs to be known. This is typically an obstacle, such as a pedestrian, other vehicles, or traffic cones. The pixel position refers to the representative position of the target object in the final synthesized target surround view image, usually based on the pixel coordinates of the center point of the bottom edge of the target detection bounding box. This is used to represent the point, as it typically corresponds to the point of contact between the target and the ground, and its altitude coordinate Z can be assumed to be 0 to simplify distance measurement. This location information can be provided by an integrated AI target detection module (as described in the AI ​​detection and recognition module section of the document).

[0116] Understandably, the purpose of this step is to obtain the input required for visual ranging, namely the two-dimensional position of the target in the panoramic image. Accurate ranging begins with accurate target localization. Since the panoramic image of the target is a dynamically stitched and geometrically corrected top view with a unified coordinate system, the pixel position of the target in this image has a potentially computable mapping relationship with its horizontal position (X,Y) on the real-world ground plane. This makes ranging based on a monocular top view possible. This step serves as the interface connecting visual perception and geometric measurement.

[0117] Specifically, the generated target surround view image for each frame is analyzed, outputting the detected obstacle category, confidence score, and its bounding box (x, y, width, height). Pixel positions for ranging are extracted from the detection results. A common strategy is to select the center point of the bottom edge of the detection box (x + width / 2, y + height), because this point usually corresponds to the contact point between the target and the ground, and its height (Z coordinate) mapped to the world coordinate system is 0 or known (e.g., vehicle chassis height). This greatly simplifies the back-projection calculation from 2D pixels to the 3D world. This pixel coordinate... This refers to inputting the representative position of the target object to be measured.

[0118] Step S62: Determine the current mapping relationship between the image coordinate system of the target panoramic image and the vehicle coordinate system of the vehicle based on the current top-down transformation matrix and the current pose parameters.

[0119] It should be noted that the current mapping relationship is a two-dimensional pixel coordinate from the target ring view image. This involves a mathematical transformation function or lookup table to the vehicle's three-dimensional spatial coordinates (X, Y, Z). Since the target surround view image is stitched together from multiple images, different regions may originate from different original cameras, and the stitching process may involve non-linear fusion. Therefore, this mapping relationship needs to be combined with the current top-down transformation matrix. (Defines the transformation from each camera image to a standard top view) and the current pose parameters. (The relationship between each camera and the world / vehicle coordinate system is defined) to jointly determine, thereby enabling the tracing of the original camera corresponding to any point on the image and the three-dimensional ray of that point.

[0120] Understandably, the purpose of this step is to establish an accurate coordinate back-projection model synchronized with the current vehicle body attitude. To deduce the corresponding 3D world point (especially a point on the ground plane) from a pixel in the target surround view image, we need to know which camera's original image the pixel originated from, and the current accurate pose of that camera. This is achieved through dynamically updated... and This allows for the establishment of a mapping from each pixel (or region) in the top view to the vehicle coordinate system. This mapping is dynamic and updates as the vehicle's attitude changes, ensuring that the geometric model upon which the ranging relies remains accurate. This is the core of achieving dynamic correction of ranging errors.

[0121] Specifically, for non-overlapping regions, each pixel uniquely originates from a single camera i. The mapping relationship can be determined using the parameters of that camera: First, using... Point the top view Map back to the original image coordinates of the camera (if distortion correction or other processing is needed). More importantly, directly utilize the current pose of the camera. and internal reference By combining the ground plane constraint Z=0, we can establish the pixel data from the top view. A direct calculation formula or lookup table for vehicle coordinates (X,Y). For overlapping areas, the mapping relationship of the dominant camera can be approximated, or a weighted average of the mapping results from the two cameras can be used. In each cycle, based on the latest... and For each pixel (or sparse grid point) of the entire target panoramic image, pre-calculate its corresponding (X,Y,0) coordinates and generate a pixel-world coordinate mapping lookup table.

[0122] Step S63: Based on the current mapping relationship, the pixel position of the target object to be measured is converted to the vehicle coordinate system to obtain the actual position coordinates of the target object relative to the vehicle.

[0123] It should be noted that the obtained pixel positions are determined by the current mapping relationship. As input, the two-dimensional horizontal coordinates of that point in the vehicle coordinate system are directly calculated. Since we assume the target is standing on the ground, its Z-coordinate is set to 0. This refers to the actual position coordinates of the target object relative to the vehicle. The calculation may involve interpolation (if a lookup table is used).

[0124] Understandably, this step is the core computational step in visual ranging. Its purpose is to accurately and in real-time convert the perceived results (pixel positions) in the image domain into physical domain coordinates (position in the vehicle coordinate system), which are more meaningful for vehicle control. By applying an accurate mapping relationship established based on dynamic pose parameters, the pixel positions detected in the surround view image can be restored to the horizontal position of the target object relative to the center of the vehicle in the real world. The accuracy of this coordinate directly determines the accuracy of subsequent distance calculations. Because the mapping relationship is dynamically updated, even if the vehicle is tilted, the projected position of the target on the horizontal ground can be correctly calculated, rather than incorrectly mapping along the tilted image plane, thus fundamentally correcting the ranging system error caused by changes in vehicle posture.

[0125] Specifically, if a lookup table method is used, the implementation involves simple table lookup and bilinear interpolation operations: based on pixel coordinates... The corresponding (X,Y) values ​​are read from a pre-computed lookup table (or obtained through interpolation). If real-time calculation is used, the values ​​are substituted into the mapping function F for calculation. The calculation principle is to find the intersection point of the 3D spatial ray projected from the pixel with the ground plane (Z=0). Using the current pose and intrinsic parameters, the ray equation can be established, and the coordinates are obtained by finding its intersection with the plane.

[0126] Step S64: Calculate the actual distance between the target object to be measured and the vehicle based on the actual position coordinates.

[0127] It should be noted that, after obtaining the horizontal position coordinates of the target object in the vehicle coordinate system... Next, calculate the Euclidean distance from the coordinate point to a predetermined reference point on the vehicle (usually the origin of the vehicle coordinate system, or a point on the vehicle's outline, such as the center of the bumper). This distance is the actual distance. For ground targets, calculate the horizontal distance. It is also possible to calculate the distance to specific vehicle components (such as the nearest corner point), which requires knowing the position of the reference point in the vehicle coordinate system.

[0128] Understandably, this step transforms precise three-dimensional position coordinates into a more intuitive scalar metric—distance—that directly impacts safety decisions. Precise distance information can be derived from accurate position coordinates using a simple Euclidean distance formula. Since the position coordinates are calculated based on a dynamically updated, accurate geometric model, this distance information inherits its high precision and robustness. This effectively overcomes the problem of large ranging errors in traditional surround-view systems due to the use of fixed mapping parameters when the vehicle's attitude changes, achieving dynamic correction of ranging errors under complex operating conditions.

[0129] Specifically, calculate the horizontal distance to the vehicle's origin: If the target height or vehicle reference point height is considered, the three-dimensional Euclidean distance can be calculated. According to... The sign of the target indicates which quadrant of the vehicle it is located in (e.g., front left, rear right).

[0130] This embodiment, through the above-described scheme, obtains the pixel position of the target in the surround view and, utilizing the accurate mapping relationship established by the core achievement of this invention (dynamically updated current top-down transformation matrix and current pose parameters), converts the pixel position into actual coordinates in the vehicle coordinate system, thereby calculating the precise distance. This series of steps fully utilizes the real-time, accurate geometric model provided by the dynamic adaptive system, ensuring that the accuracy of visual ranging is no longer affected by changes in vehicle posture. This not only expands the functionality of the surround view system but also elevates it from a simple visual aid to a reliable perception unit, providing reliable technical support for the application of advanced driver assistance functions (such as collision warning and automatic parking) in dynamic and complex engineering machinery conditions.

[0131] For example, to help understand the implementation process of the adaptive pose surround view stitching method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 2 , Figure 2 A simplified flowchart of an adaptive pose-based lookaround stitching method is provided, specifically: The system first initiates the initialization and monitoring branches. Initialization involves camera calibration, which acquires the initial pose parameters of each surround-view camera in the initial calibration state. This is followed by further decomposition of the camera's spatial position to extract the fundamental coordinate information for subsequent geometric transformations. Monitoring, on the other hand, captures the vehicle's posture changes in real-time, serving as the basis for dynamic compensation. Next, the system enters a core decision node for pose change, detecting whether the vehicle's posture has significantly changed compared to the initial state. If no change in pose is detected, the system skips the compensation step and directly proceeds to updating the surround-view image to maintain image refresh. If a change in pose is detected, the process enters the camera pose correction stage, where the previously decomposed initial spatial position is combined with the currently captured pose change parameters to dynamically correct the camera's virtual pose. Finally, the extrinsic parameter update step recalculates and updates the extrinsic transformation parameters of each camera based on the corrected results. Finally, the system updates the surround view image based on the latest extrinsic parameters, transforms and stitches the images captured by each camera in real time to generate the final target surround view image, and ends the cycle upon completion. This closed-loop design ensures that the surround view image can always adaptively reflect the vehicle's current real spatial attitude, effectively solving the image misalignment problem caused by vehicle body bumps or attitude changes.

[0132] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the adaptive pose surround stitching method of this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0133] This application also provides an adaptive pose surround view stitching device, please refer to... Figure 3 The adaptive pose surround view stitching device includes: The calibration pose acquisition module 301 is used to acquire the initial pose parameters of each surround view camera of the vehicle in the initial calibration state. The camera pose acquisition module 302 is used to acquire the vehicle body pose change parameters in the current working state; The pose dynamic compensation module 303 is used to dynamically compensate the initial pose parameters according to the vehicle body posture change parameters, and generate the current pose parameters corresponding to each of the surround view cameras. The surround view image generation module 304 is used to transform and stitch the current surround view images captured by each of the surround view cameras based on the current pose parameters to generate a stitched target surround view image.

[0134] The adaptive pose surround view stitching device provided in this application, employing the adaptive pose surround view stitching method in the above embodiments, can solve the technical problem of low reliability of surround view image imaging from the camera. Compared with the prior art, the beneficial effects of the adaptive pose surround view stitching device provided in this application are the same as those of the adaptive pose surround view stitching method provided in the above embodiments, and other technical features in the adaptive pose surround view stitching device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0135] This application provides an adaptive pose surround view stitching device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the adaptive pose surround view stitching method in the first embodiment described above.

[0136] The following is for reference. Figure 4This document illustrates a structural schematic diagram suitable for implementing the adaptive pose surround view splicing device in the embodiments of this application. The adaptive pose surround view splicing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The adaptive pose surround view stitching device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0137] like Figure 4 As shown, the adaptive pose surround view splicing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the adaptive pose surround view splicing device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the adaptive pose surround view stitching device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show adaptive pose surround view stitching devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0138] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0139] The adaptive pose surround view stitching device provided in this application, employing the adaptive pose surround view stitching method in the above embodiments, can solve the technical problem of low reliability of surround view image imaging from the camera. Compared with the prior art, the beneficial effects of the adaptive pose surround view stitching device provided in this application are the same as those of the adaptive pose surround view stitching method provided in the above embodiments, and other technical features in this adaptive pose surround view stitching device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0140] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0141] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0142] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the adaptive pose surround view stitching method in the above embodiments.

[0143] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0144] The aforementioned computer-readable storage medium may be included in the adaptive pose surround view splicing device; or it may exist independently and not be assembled into the adaptive pose surround view splicing device.

[0145] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the adaptive pose surround view stitching device, the adaptive pose surround view stitching device: acquires the initial pose parameters of each surround view camera in the initial calibration state of the vehicle; acquires the vehicle body posture change parameters in the current working state; dynamically compensates the initial pose parameters according to the vehicle body posture change parameters to generate the current pose parameters corresponding to each surround view camera; and transforms and stitches the current surround view images acquired by each surround view camera based on the current pose parameters to generate the stitched target surround view image.

[0146] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C++ or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0148] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0149] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described adaptive pose surround view stitching method, which can solve the technical problem of low reliability of surround view image imaging by the camera. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the adaptive pose surround view stitching method provided in the above embodiments, and will not be repeated here.

[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the adaptive pose surround view stitching method described above.

[0151] The computer program product provided in this application can solve the technical problem of low reliability of panoramic image imaging from cameras. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the adaptive pose panoramic stitching method provided in the above embodiments, and will not be repeated here.

[0152] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. An adaptive pose-based surround view stitching method, characterized in that, The adaptive pose surround view stitching method includes: Acquire the initial pose parameters of each surround view camera of the vehicle in the initial calibration state; Obtain the vehicle body attitude change parameters under the current working state; Based on the vehicle body posture change parameters, the initial pose parameters are dynamically compensated to generate the current pose parameters corresponding to each of the surround view cameras; Based on the current pose parameters, the current panoramic images captured by each of the panoramic cameras are transformed and stitched together to generate a stitched target panoramic image.

2. The adaptive pose surround view stitching method as described in claim 1, characterized in that, The initial pose parameters include an initial rotation matrix and an initial translation matrix. The step of obtaining the initial pose parameters of each surround-view camera of the vehicle in the initial calibration state includes: The surround-view camera is calibrated to obtain the initial homography matrix corresponding to the surround-view camera; Based on the initial homography matrix, calculate the initial rotation matrix and initial translation matrix of each of the surround-view cameras relative to the vehicle coordinate system.

3. The adaptive pose surround view stitching method as described in claim 2, characterized in that, The current pose parameters include the current rotation matrix and the current translation matrix. The step of dynamically compensating the initial pose parameters based on the vehicle body posture change parameters to generate the current pose parameters corresponding to each of the surround-view cameras includes: Based on the vehicle body attitude change parameters, generate an attitude compensation rotation matrix; The attitude compensation rotation matrix is ​​multiplied by the initial rotation matrix to obtain the current rotation matrix; Transform the initial translation matrix to the world coordinate system to obtain the initial position of the surround-view camera in the world coordinate system; Based on the vehicle body posture change parameters, offset compensation is performed on the initial position to obtain the compensated position; The compensated position is transformed back to the camera coordinate system using the current rotation matrix to obtain the current translation matrix.

4. The adaptive pose surround view stitching method as described in claim 3, characterized in that, The step of transforming and stitching the current panoramic images captured by each of the panoramic cameras based on the current pose parameters to generate a stitched target panoramic image includes: Based on the current rotation matrix and the current translation matrix, determine the current top-down transformation matrix corresponding to each of the surround-view cameras; Based on the current top-down transformation matrix, the current panoramic images captured by each of the panoramic cameras are converted to a top-down view to obtain the corresponding single-channel top-down image; Determine the overlapping area between the single-channel top-view images of adjacent surround-view cameras; The pixels in the overlapping area are weighted and fused according to the fusion weight, and the individual top-view images are stitched together to form the target panoramic image.

5. The adaptive pose surround view stitching method as described in claim 4, characterized in that, The step of determining the current top-down transformation matrix corresponding to each of the surround-view cameras based on the current rotation matrix and the current translation matrix includes: Based on the current rotation matrix, current translation matrix and camera internal parameters of each of the surround-view cameras, the preset world coordinate system point set is projected onto the image pixel coordinate system of each of the surround-view cameras to generate the current pixel point set. Based on the current pixel set and the preset top-view point set, the current top-view transformation matrix corresponding to the surround-view camera is calculated.

6. The adaptive pose surround view stitching method as described in claim 5, characterized in that, The adaptive pose surround view stitching method further includes: Obtain the pixel position of the target object to be measured in the target panoramic image; Based on the current top-down transformation matrix and the current pose parameters, determine the current mapping relationship between the image coordinate system of the target surround view image and the vehicle coordinate system of the vehicle; Based on the current mapping relationship, the pixel position of the target object to be measured is transformed to the vehicle coordinate system to obtain the actual position coordinates of the target object relative to the vehicle. Based on the actual location coordinates, calculate the actual distance between the target object to be measured and the vehicle.

7. An adaptive pose surround view stitching device, characterized in that, The adaptive pose surround view stitching device includes: The calibration pose acquisition module is used to acquire the initial pose parameters of each surround view camera of the vehicle in the initial calibration state; The camera pose acquisition module is used to acquire the vehicle body pose change parameters in the current working state; The pose dynamic compensation module is used to dynamically compensate the initial pose parameters according to the vehicle body posture change parameters, and generate the current pose parameters corresponding to each of the surround view cameras. The surround view image generation module is used to transform and stitch together the current surround view images captured by each of the surround view cameras based on the current pose parameters to generate a stitched target surround view image.

8. An adaptive pose surround view stitching device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the adaptive pose surround stitching method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the adaptive pose surround view stitching method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the adaptive pose surround view stitching method as described in any one of claims 1 to 6.