Vehicle augmented reality perspective method and apparatus

CN122821052APending Publication Date: 2026-09-25ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610792645.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

本发明解决了动态车载环境下乘客视角透视图像易产生位置偏移与渲染延迟的问题,实现了高精度、低延迟的增强现实透视效果

Benefits of technology

[0026]本发明实施例的一种车辆增强现实透视方法及装置,通过获取乘客在车辆坐标系中的初始六自由度位姿数据,并融合多源传感器的观测信息进行修正得到实时头部位姿;根据实时头部位姿确定双眼虚拟视线方向,计算其与车辆外表面网格模型的交点作为虚拟穿透点;基于虚拟穿透点和多路车载摄像头图像数据,经视点导向映射生成以双眼为中心的虚拟投影面上的透视图像;对透视图像进行畸变校正与多频带融合,并结合预测头部位姿执行异步时间扭曲补偿后输出至显示设备。该方法有效解决了现有技术因动态头部位姿变化与渲染延迟导致的透视图像位置偏移、视觉残留问题,实现了低延迟、高精度的车载增强现实透视效果,显著提升了乘客沉浸体验与系统鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821052A_ABST
    Figure CN122821052A_ABST
Patent Text Reader

Abstract

The application discloses a kind of vehicle augmented reality perspective method and device, belong to the intersection technical field of intelligent cabin and augmented reality.It is obtained by getting the initial six-degree-of-freedom pose of passenger in vehicle coordinate system, and real-time head pose is corrected by fusing multi-source sensor;According to real-time head pose, the intersection of the virtual line of sight direction of both eyes is calculated as a virtual penetration point with the grid model of the outer surface of the vehicle;Based on virtual penetration point and multi-path vehicle camera image, the perspective image on the virtual projection plane with both eyes as the center is generated by view point orientation mapping;Distortion correction and multi-band fusion are carried out on the perspective image, and asynchronous time warping compensation is executed by joint prediction of head pose, and output to display device.The application solves the problem that perspective image is easy to produce position deviation and rendering delay in dynamic vehicle environment, and realizes high-precision, low-latency augmented reality perspective effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cockpit and augmented reality technology, and in particular to a method and apparatus for augmented reality perspective of vehicles. Background Technology

[0002] The application of augmented reality (AR) technology in the automotive field is receiving increasing attention, especially the use of AR devices to assist occupants in perceiving the vehicle's external environment, which has become an important development direction for smart cockpits. In existing technologies, vehicle surround-view systems mostly use fisheye cameras installed around the vehicle body to generate a fixed top-down or surround view through image stitching, which is then displayed on the central control screen. These systems have significant limitations: First, the display angle is fixed and cannot dynamically adjust with occupant head movements, forcing occupants to rely on eye movement and spatial imagination to understand the external environment, resulting in a heavy cognitive burden and slow response time; second, existing surround-view systems typically only cover a limited area around the vehicle body, leaving significant blind spots above (e.g., height-restricted obstacles, entrances to multi-level parking garages) and below (e.g., road potholes, off-road conditions), resulting in insufficient field of view coverage; third, the end-to-end latency from image acquisition to display is generally over 80ms, causing noticeable lag when occupants' heads turn rapidly, easily triggering dizziness.

[0003] In addition, some studies have attempted to use AR glasses for vehicle perspective, but they rely on a single visual inertial odometry for head tracking, which is prone to drift and loss in environments with complex lighting and sparse feature textures inside the vehicle. At the same time, existing solutions usually only perform simple cropping or magnification of a single camera image, lacking light field rendering based on multi-camera fusion, resulting in problems such as stitching seams and parallax errors in perspective images.

[0004] Therefore, how to achieve high-precision, low-latency head tracking in the complex environment inside a vehicle, and synthesize a seamless perspective image that highly matches the occupant's line of sight, is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] The present invention aims to at least partially solve one of the technical problems in the related art.

[0006] To address this issue, this invention proposes a vehicle augmented reality perspective method. It acquires the initial six-DOF pose data of the passenger in the vehicle coordinate system and corrects this pose data by fusing observation information from multiple sources to obtain the real-time head pose. Based on the real-time head pose, the virtual gaze direction of the passenger's eyes is determined, and the intersection of this gaze direction with the vehicle's outer surface mesh model is calculated as the virtual penetration point. Based on the virtual penetration point and image data collected by multiple vehicle-mounted cameras, a perspective image on a virtual projection surface centered on the passenger's eyes is generated through viewpoint-guided mapping. Distortion correction and multi-band fusion are performed on the perspective image, and asynchronous time distortion compensation is executed in conjunction with the predicted head pose. The final image is then output to a display device for rendering. This invention solves the problem of positional offset and rendering delay in passenger-view perspective images under dynamic vehicle environments, achieving a high-precision, low-latency augmented reality perspective effect.

[0007] Another object of the present invention is to provide a vehicle augmented reality perspective device.

[0008] To achieve the above objectives, the present invention provides a vehicle augmented reality perspective method, comprising: The initial six-DOF pose data of the passenger in the vehicle coordinate system is obtained, and the pose data is corrected by fusing observation information from multiple sources of sensors to obtain the real-time head pose. The virtual gaze direction of the passenger's eyes is determined based on the real-time head position and posture, and the intersection of this gaze direction with the vehicle's outer surface mesh model is calculated as the virtual penetration point. Based on image data collected from virtual penetration points and multiple vehicle cameras, a perspective image on a virtual projection surface centered on the passenger's eyes is generated through viewpoint guidance mapping. The perspective image is subjected to distortion correction and multi-band fusion, and asynchronous time warp compensation is performed in combination with the predicted head pose. The final image is then output to the display device for rendering.

[0009] In one embodiment of the present invention, the step of acquiring the initial six-DOF pose data of the passenger in the vehicle coordinate system and correcting the pose data by fusing observation information from multiple sources of sensors to obtain the real-time head pose includes: The triaxial acceleration and triaxial angular velocity data collected by the built-in IMU of the AR glasses are input into the DSP core of the domain controller and fed into the constant acceleration and constant angular velocity kinematic model. The state vector and covariance matrix at time t are then recursively predicted as prior estimates. The distance observations measured by the UWB base station, the matching results of ORB feature points extracted by the fisheye camera of the AR glasses and the cockpit feature map, and the fitted center position of the head point cloud output by the depth sensor are sequentially input into the observation update module of the extended Kalman filter to form a multi-source observation vector. The predicted values ​​of each observation are calculated through the observation equation for subsequent sequential correction. Based on the extended Kalman filter algorithm, the Kalman gain is adjusted using the dynamic information matrix to sequentially update the prior estimates, eliminating absolute position drift and relative pose error, and obtaining smooth six-degree-of-freedom pose data with a frequency of 100Hz.

[0010] In one embodiment of the present invention, the update process of the extended Kalman filter includes: Construct a twelve-dimensional system state vector containing position, attitude angles, and their rates of change:

[0011] The observation noise covariance matrix is ​​adjusted in real time based on the UWB signal-to-noise ratio and the confidence level of the depth sensor data. ; Execute the Kalman update formula to output the corrected real-time head pose. Covariance Matrix :

[0012]

[0013] .

[0014] In one embodiment of the present invention, the step of determining the virtual gaze direction of the passenger's eyes based on the real-time head pose and calculating the intersection of the gaze direction with the vehicle's outer surface mesh model as the virtual penetration point includes: Real-time head pose data is input into the gaze generation module, which then parses the coordinates of the passenger's left eye. and right eye coordinates And a gaze direction vector is given for each eye. Space rays are constructed using a ray generation algorithm: .

[0015] In one embodiment of the present invention, the generation of a perspective image on a virtual projection surface centered on the passenger's eyes via viewpoint-guided mapping based on image data acquired from virtual penetration points and multiple vehicle-mounted cameras includes: The pre-constructed virtual projection sphere centered on the passenger's eyes, along with the calibrated spatial pose matrices and projection parameters of each onboard camera, are input into the viewpoint guidance mapping unit. The mapping unit then maps each discrete direction vector on the virtual projection sphere. Inverse solution for the corresponding three-dimensional points outside the vehicle in the relevant direction And using the projection equation calculate Projecting coordinates onto the pixel plane of each camera to obtain the set of multi-camera pixel coordinates for each direction:

[0016] Calculate the orientation of each camera pixel pair Mixed weights:

[0017] in, It is a pixel The corresponding optical center line of sight direction, Center of the image; The pixel color values ​​of each camera in the corresponding direction are weighted and averaged based on the hybrid weights to obtain the final color value in the corresponding direction on the virtual projection surface, thus forming a perspective image.

[0018] In one embodiment of the present invention, the step of performing distortion correction and multi-band fusion on the perspective image, and performing asynchronous time warp compensation in conjunction with the predicted head pose, and outputting the final image to a display device for rendering includes: The raw image captured by the fisheye camera is input into the radial distortion model to calculate the correction coordinates:

[0019]

[0020] in, , These are the coordinates on the normalized image plane. These are the distortion coefficients, obtained through offline calibration. The corrected left and right eye images are decomposed into Laplacian pyramids, and Gaussian weight gradient is applied to the overlapping areas on different pyramid layers to perform multi-band fusion, resulting in a seamless perspective image. The predicted head position is obtained by extrapolating the latest IMU data at the current moment. If the rendered image timestamp is delayed, the pixel orientation under the old pose is mapped to the image coordinates under the new pose using the reprojection formula, performing asynchronous time warp compensation; whereby the reprojection formula is:

[0021] The compensated image is encoded and output to AR smart glasses via a wireless link for waveguide display.

[0022] In one embodiment of the present invention, it further includes: Based on multi-frame images from a checkerboard calibration board, the intrinsic parameters and distortion coefficients of each camera are calculated using the Zhang Zhengyou calibration method. For fisheye cameras with a field of view exceeding a preset threshold, a distortion correction model suitable for fisheye models is used for fitting to generate intrinsic parameter files. Place the vehicle in an open area, place visual markers at preset positions outside the vehicle, control all cameras to synchronously acquire images containing the visual markers, use the PnP algorithm to calculate the rotation matrix and translation vector of each camera relative to the vehicle coordinate system, and generate an external parameter file. Based on the extrinsic parameter file and the vehicle CAD external surface mesh model, the optimal camera index and its pixel coordinates corresponding to each virtual line of sight are pre-calculated according to the preset angle sampling step size. The mapping relationship is stored as a lookup table and loaded into the GPU constant memory.

[0023] In one embodiment of the present invention, image data is acquired using multiple vehicle-mounted cameras, including: The front-view camera, rear-view camera, left / right-view camera, top-view camera, and bottom-view camera send video streams to the domain controller via the GMSL link. The time synchronization unit uses a precise time protocol to stamp each data stream with the same microsecond-level timestamp, outputting time-aligned image data.

[0024] In one embodiment of the present invention, after the final image is output to a display device for rendering, the method further includes: Real-time monitoring of the status of input camera signals, sensor data integrity, and algorithm calculation time; When the loss of all camera signals exceeds a preset threshold, a Level 1 fault is triggered; the last valid image is frozen and a voice prompt is output. When either the top-view or bottom-view camera loses signal, a level two fault is triggered; a low-quality alternative image is synthesized using the remaining four cameras, and a yellow warning box is superimposed in the corner of the field of view; When the UWB signal is briefly lost, a level 3 fault is triggered; the extended Kalman filter is controlled to continue outputting pose based solely on IMU and depth sensor data, while retrying the faulty device every second. If it succeeds three times in a row, it will automatically return to full-function mode.

[0025] The present invention also proposes a vehicle augmented reality perspective device, comprising: The pose fusion module is used to acquire the initial six-degree-of-freedom pose data of the passenger in the vehicle coordinate system, and to fuse the observation information of multiple source sensors to correct the pose data in order to obtain the real-time head pose. The line-of-sight penetration module is used to determine the virtual line-of-sight direction of the passenger's eyes based on the real-time head position and calculate the intersection of this line-of-sight direction with the vehicle's outer surface mesh model as the virtual penetration point. The perspective synthesis module is used to generate perspective images on a virtual projection surface centered on the passenger's eyes through viewpoint guidance mapping, based on image data collected from virtual penetration points and multiple vehicle cameras. The correction rendering module is used to perform distortion correction and multi-band fusion on the perspective image, and perform asynchronous time distortion compensation in combination with the predicted head pose, and output the final image to the display device for rendering.

[0026] This invention discloses a vehicle augmented reality perspective method and apparatus. It acquires the initial six-DOF pose data of a passenger in the vehicle coordinate system and corrects this data by fusing observation information from multiple sensors to obtain the real-time head pose. Based on the real-time head pose, it determines the virtual gaze direction of both eyes and calculates the intersection point with the vehicle's outer surface mesh model as the virtual penetration point. Based on the virtual penetration point and image data from multiple vehicle cameras, it generates a perspective image on a virtual projection surface centered on both eyes through viewpoint guidance mapping. The perspective image undergoes distortion correction and multi-band fusion, and asynchronous time distortion compensation is performed in conjunction with the predicted head pose before being output to a display device. This method effectively solves the problems of perspective image position shift and visual persistence caused by dynamic head pose changes and rendering delays in existing technologies. It achieves low-latency, high-precision vehicle augmented reality perspective effects, significantly improving passenger immersion experience and system robustness.

[0027] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0028] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a vehicle augmented reality perspective method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a vehicle augmented reality perspective device according to an embodiment of the present invention. Detailed Implementation

[0029] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] The following description, with reference to the accompanying drawings, describes a vehicle augmented reality perspective method and apparatus according to an embodiment of the present invention.

[0032] Figure 1 This is a flowchart of a vehicle augmented reality perspective method according to an embodiment of the present invention.

[0033] like Figure 1 As shown, a vehicle augmented reality perspective method includes the following steps: S1, acquire the initial six-degree-of-freedom pose data of the passenger in the vehicle coordinate system, and fuse the observation information of multiple source sensors to correct the pose data in order to obtain the real-time head pose. S2, determine the virtual gaze direction of the passenger's eyes based on the real-time head position and pose, and calculate the intersection of the gaze direction with the vehicle's outer surface mesh model as the virtual penetration point; S3, based on the image data collected by virtual penetration points and multiple vehicle cameras, generates a perspective image on a virtual projection surface centered on the passenger's eyes through viewpoint guidance mapping; S4, perform distortion correction and multi-band fusion on the perspective image, and perform asynchronous time distortion compensation in combination with the predicted head pose, and output the final image to the display device for rendering.

[0034] First, the detailed specifications and layout of the hardware system of the present invention are shown in Table 1.

[0035] Table 1

[0036] The hardware selection criteria are as follows: FOV≥120° is a necessary condition to ensure that the edge overlap area is sufficient for stitching and fusion when synthesizing panoramic images; 60fps is the key to reducing motion blur and end-to-end latency (controlled to <50ms), and matches the critical flicker frequency of the human eye.

[0037] Then, the passenger head six-degrees-of-freedom (6DoF) pose tracking and data fusion algorithm includes: Objective: To obtain the precise position of the passenger's eyes in the vehicle coordinate system. and attitude angle That is (yaw, pitch, roll).

[0038] Sensor configuration: AR glasses end: High-precision IMU (BMI160 level, sampling rate 1kHz), monocular grayscale fisheye camera (for V-SLAM, 20fps); Cabin end: 2 UWB base stations (deployed at the front and rear of the roof) to provide centimeter-level absolute position reference; 1 infrared structured light depth sensor (deployed on the roof, covering the entire cabin).

[0039] Data fusion process (based on Extended Kalman Filter (EKF)): The system state vector is defined as:

[0040] Prediction Steps (Based on IMU Angular Velocity and Acceleration Integration): Utilizing IMU Measurements and Predicting recursively using kinematic models state of time .

[0041] Update steps (multi-source observation fusion): UWB observation equations ( ): Measure the distance from the glasses tag to two base stations. Correcting absolute position drift; V-SLAM observation equations ( ): Extract feature points from the fisheye camera and match them with a pre-established feature map inside the vehicle to correct the relative pose; depth sensor observation equation ( The depth sensor in the roof directly outputs point clouds of passengers' heads, providing high-precision position observation (with higher weighting when the head is within the sensor's field of view).

[0042] Fusion formula (typical EKF update):

[0043]

[0044]

[0045] in, It is a dynamic information matrix that adjusts its weights in real time based on the signal-to-noise ratio (SNR) of the UWB signal and the confidence level of the depth sensor data. The final output is a smooth 6DoF pose at a frequency of 100Hz.

[0046] Secondly, the virtual gaze calculation and viewpoint synthesis algorithms include: Input: Coordinates of the passenger's eyes in the vehicle coordinate system and 3D mesh model of vehicle skin Spatial pose matrix of all cameras and projection parameters.

[0047] Algorithm steps: First, gaze generation: For the left eye Given the gaze direction Generates rays Based on the vehicle skin model Calculate the first intersection point between the ray and the skin. (i.e., virtual penetration point).

[0048] Second, line-of-sight extension: Keeping the direction unchanged, extend the ray to the outside of the vehicle. Calculate the intersection point of the extended ray and the camera's imaging plane, and select the camera closest to that intersection point. .

[0049] Third, Viewpoint-Oriented Mapping: This assumes a virtual projection sphere outside the vehicle, centered on the passenger's eyes. The goal is to render the external scene bounded by this sphere.

[0050] Furthermore, for any direction vector on the virtual sphere The task is to find the corresponding point on the pixel plane of the camera image; this is a light field rendering problem. We adopt a planar perspective projection warping method based on depth maps.

[0051] Mathematical mapping model: For cameras Its projection matrix is 3D points outside the vehicle The coordinates on the pixel plane are:

[0052] In the perspective system, the passenger's line of sight is known. We solve for the virtual point corresponding to this direction in reverse. (Assume Q is located on a sphere at a certain distance from the vehicle's outer surface). Then, take the image coordinates of Q projected onto all cameras, and obtain the final color value in that direction by weighted averaging.

[0053] Weight calculation (mixed weights):

[0054] in, It is a pixel The corresponding optical center line of sight. Pixels with small line of sight angles and close to the image center are preferred.

[0055] Furthermore, image stitching and distortion correction algorithms include: Because multiple fisheye cameras (FOV>150°) were used, strict distortion correction and stitching were required.

[0056] Distortion model (using radial distortion model):

[0057]

[0058] in, , These are the coordinates on the normalized image plane. It is the distortion coefficient, obtained through offline calibration.

[0059] Real-time stitching process: Calibration: When the vehicle leaves the factory, the intrinsic parameters (focal length, distortion) and extrinsic parameters (installation angle, position) of each camera are calibrated; Remapping: The sampling coordinates (lookup table LUT) of each pixel in the original camera image are pre-calculated for the eyeglass viewpoint; Multi-band Blending: In the overlapping area, in order to avoid stitching seams, Laplacian pyramid decomposition is used to perform weighted fusion of images of different frequency bands, preserving high-frequency details while smoothing low-frequency transitions.

[0060] Furthermore, end-to-end latency analysis and real-time performance assurance include: Latency Budget (Target <35ms): Acquisition Latency (5ms): Rolling shutter CMOS sensor, using global reset mechanism; Transmission Latency (3ms): Serial transmission using GMSL (Gigabit Multimedia Serial Link); Processing Latency (Main Line) (20ms): Distortion correction + remapping: 8ms (GPU accelerated); Viewpoint synthesis + frame interpolation: 10ms (depth estimation and rendering); EKF pose fusion: 2ms (DSP processing). Display Latency (7ms): AR glasses OLED screen response time + transmission.

[0061] Latency reduction techniques: Asynchronous Time Warp (ATW): In the rendering pipeline, if a new pose arrives but the new image is not yet ready, a coarse projection warp is performed based on the old image and the latest pose to fill in the image lag caused by minor delays; Sensor prediction: Utilizing the high sampling rate of the IMU to predict the future. The head pose in real time is used to render the "future" scene and offset processing delays.

[0062] Finally, the compatibility of this invention with existing 360° surround view systems includes: Typical 360° surround view camera specifications: FOV: typically 190° ~ 200° (fisheye); resolution: 1280x960 @ 30fps; dynamic range: 120dB.

[0063] Does it meet the perspective requirements? Yes: For eliminating blind spots at the A / B pillars, its 120° horizontal FOV and 30fps are sufficient (60fps can be achieved through algorithmic frame interpolation); Shortcomings: The original surround-view camera is installed too low (rearview mirror, bumper), causing severe obstruction and parallax errors when viewing the roof and chassis. Therefore, additional roof and chassis cameras are necessary; Recommended modification: Upgrade the original surround-view controller to a higher-performance SoC (such as Qualcomm Snapdragon Ride or Horizon Robotics Journey 5), embedding the aforementioned algorithm, and sharing an Ethernet transmission link.

[0064] Furthermore, as shown in Table 2, compared with the traditional 360° surround view system (displayed on the central control screen, with a fixed top-down view), the present invention: Table 2

[0065] The method of this invention effectively solves the problems of perspective image position shift and visual dizziness caused by dynamic head pose changes and rendering delays in the prior art, and achieves low-latency, high-precision in-vehicle augmented reality perspective effect, significantly improving passenger immersion experience and system robustness.

[0066] In addition, the overall system architecture of the present invention includes: First layer: Perception layer (sensors and signal input).

[0067] Visual sensor group: 1 front-view camera (location: behind the front grille / interior rearview mirror, output: GMSL serial video stream); 1 rear-view camera (location: above the license plate light, output: GMSL serial video stream); 2 left / right-view cameras (location: bottom of the left and right exterior rearview mirrors, output: GMSL serial video stream); 1 top-view camera (location: inside the shark fin antenna, output: GMSL serial video stream); 1 bottom-view camera (location: center of the chassis, output: GMSL serial video stream).

[0068] Passenger tracking sensor group: AR glasses with built-in IMU (output: 1kHz, three-axis acceleration + three-axis angular velocity); AR glasses with built-in monocular fisheye camera (output: 20fps, image feature stream); 2 UWB base stations in the carriage (location: front and rear of the roof, output: time difference of arrival TDOA, centimeter-level ranging); 1 depth sensor in the carriage ceiling (location: center of the rear ceiling, output: point cloud / depth map).

[0069] The second layer: data transmission and preprocessing layer.

[0070] Serializer / Deserializer: Deserializes each GMSL signal into a MIPI CSI-2 interface for input to the domain controller.

[0071] Time synchronization unit: PTP (Precise Time Protocol) is used to stamp all sensors with the same timestamp (with microsecond precision).

[0072] Preprocessing module: performs black level correction and lens shadow correction on the original RAW image from the camera; performs noise reduction and downsampling on the depth sensor point cloud; performs zero bias compensation and low-pass filtering on the IMU data.

[0073] The third layer: the core computing layer (domain controller - Qualcomm SA8650P or equivalent platform).

[0074] Coprocessor partitioning: DSP core: runs Extended Kalman Filter (EKF) pose fusion algorithm; GPU core: runs perspective view synthesis, multi-band image stitching, asynchronous time warp (ATW); CPU core: runs V-SLAM feature matching, UWB ranging calculation, and overall task scheduling.

[0075] Internal data bus: Shared memory (bandwidth > 25 GB / s) is used for zero-copy exchange of video frames and pose data.

[0076] Fourth layer: Communication and display layer.

[0077] Wireless transmission module: Wi-Fi 6 / 60GHz Wigig (millimeter wave), bandwidth >3Gbps, latency <5ms.

[0078] AR smart glasses: Receive compressed video streams for the left and right eyes (resolution: 1280×720 per eye, frame rate: 60~120Hz); after decompression, the streams are projected onto the retina via an optical waveguide display.

[0079] Feedback loop: The IMU data from the glasses is wirelessly transmitted back to the domain controller, forming a closed loop.

[0080] Furthermore, the core algorithm processing flow of the overall system architecture of the present invention includes: 1: System initialization and calibration process.

[0081] Step 1.1: Start self-test. Sub-step 1.1.1: Check whether the 6 cameras are connected in sequence, and read the device ID and the current image frame; Sub-step 1.1.2: Check the wireless pairing status of the UWB base station, depth sensor, and AR glasses; Sub-step 1.1.3: After all devices pass, the system enters "calibration mode"; if any fails, an error is reported and the system reverts to the factory default state.

[0082] Step 1.2: Offline Intrinsic Parameter Calibration. Sub-step 1.2.1: For each camera, play the checkerboard calibration board (10×7 corner points, 30mm side length) and acquire 20 clear images in different poses; Sub-step 1.2.2: Using Zhang Zhengyou's calibration method, calculate the focal length (fx, fy), principal point (cx, cy), and distortion coefficients (k1,k2,k3,p1,p2) for each camera. Output file "cam_intrinsics.yaml"; Sub-step 1.2.3: For fisheye cameras (FOV≥150°), additionally use the Kannala-Brandt model for fitting.

[0083] Step 1.3: Offline extrinsic parameter calibration. Sub-step 1.3.1: Place the vehicle on an open, flat surface and place Apriltag QR codes at specific marker points outside the vehicle (four corners and directly above the roof); Sub-step 1.3.2: All cameras simultaneously capture a single frame, detecting the corner points of the Apriltag in the image; Sub-step 1.3.3: Use the PNP algorithm to calculate the rotation matrix R_i and translation vector t_i of each camera relative to the vehicle coordinate system (origin: vehicle center ground projection, X front, Y left, Z top). Output file "cam_extrinsics.yaml".

[0084] Step 1.4: Import the vehicle skin 3D model. Sub-step 1.4.1: Read the vehicle body outer surface mesh (OBJ or STL format) from the vehicle CAD model, simplify it to a face count of <5000 for real-time intersection calculation; Sub-step 1.4.2: Convert the mesh vertex coordinates to the vehicle coordinate system established in step 1.3.

[0085] Step 1.5: Cockpit Feature Map Construction (for V-SLAM). Sub-step 1.5.1: Fix the depth sensor to the ceiling and scan all seat positions and passenger headroom space inside the vehicle; Sub-step 1.5.2: Extract ORB feature points and construct a 3D point cloud map "cabin_map.bin".

[0086] Step 1.6: Pre-compute the lookup table (LUT). Sub-step 1.6.1: For each virtual view direction (sampled every 1°), pre-compute which camera and pixel coordinates correspond to that direction; Sub-step 1.6.2: Store this mapping relationship as a LUT and load it into the GPU's constant memory for real-time rendering.

[0087] 2: The main loop runs in real time (processing every frame).

[0088] Define the main loop cycle: trigger once every 16.67ms (corresponding to 60fps). The following is the sequential process within a single loop.

[0089] Phase A: Parallel acquisition of data from multiple sensor sources.

[0090] A1: Camera Data Acquisition. Request the latest frames from 6 cameras in parallel (using multi-threading, with each thread corresponding to one camera); each camera returns a raw image with a timestamp (RAW10 or RAW12 format); if a camera returns a timeout (>5ms), use the previous frame instead and set an "stale flag".

[0091] A2: Passenger tracking sensor data acquisition. A2.1: Read the latest queue of the AR glasses IMU via BLE (containing all sampling points within the past 5ms, frequency 1kHz); A2.2: Read the latest image from the AR glasses fisheye camera via a dedicated wireless channel (20fps, skip if no new frame arrives); A2.3: Read the distances d1 and d2 from the glasses tag to the two base stations measured by the UWB base station via the CAN interface (update rate 1kHz); A2.4: Read the point cloud output by the depth sensor via Ethernet (30fps, skip if no new frame arrives).

[0092] A3: Timestamp alignment.

[0093] All sensor data are rearranged by timestamp (global PTP clock), and other sensor data within ±2ms are selected as the observation value of the frame, based on the camera frame timestamp.

[0094] Phase B: High-precision 6DoF pose fusion (running on DSP).

[0095] B1: State prediction (based on IMU kinematics).

[0096] Input: The previous optimal state X_{t-1} (12-dimensional vector), the acceleration a and angular velocity ω measured by the IMU; Calculation formula: X_t^- = f(X_{t-1}, a, ω, Δt), where f is a constant acceleration / constant angular velocity model; Simultaneously update the covariance matrix P_t^-.

[0097] B2: Observation update (sequential fusion).

[0098] B2.1: UWB Update.

[0099] Observation vector Z_uwb = [d1, d2]; calculate residual y_uwb = Z_uwb - h_uwb(X_t^-), where h_uwb maps pose to range; calculate Kalman gain K_uwb, and update X_t = X_t + K_uwb * y_uwb.

[0100] B2.2: V-SLAM update (only when a new fisheye image arrives).

[0101] Extract ORB feature points from the current fisheye image; perform feature matching with the cockpit map "cabin_map.bin"; if the number of successfully matched points is greater than 30, solve the current pose using PNP as the observation Z_slam; the residual y_slam = Z_slam - X_t (the first three position components), and update X_t.

[0102] B2.3: Depth sensor update (only when new point cloud arrives) Ellipse fitting is performed on the point cloud of the head region in the depth point cloud to obtain the head center position Z_depth; the residual y_depth = Z_depth - X_t (position part) is also updated through EKF.

[0103] B3: Output the final 6DoF pose (x, y, z, ψ, θ, φ) and binocular coordinates E_L, E_R.

[0104] Phase C: Virtual perspective image synthesis (runs on the GPU, executed separately for the left and right eyes).

[0105] Input: Eye coordinates E_L / E_R, current all camera frames, vehicle skin mesh M.

[0106] C1: Generates a virtual left eye image pixel by pixel.

[0107] C2: Generate the right eye image.

[0108] C3: Multi-band splicing and seamless integration.

[0109] Step C3.1: Decompose the left-eye and right-eye images into Laplacian pyramids respectively; Step C3.2: Apply Gaussian weight gradients to the overlapping regions (i.e., regions where multiple cameras contribute pixels) on different pyramid layers; Step C3.3: Reconstruct the images to obtain the final left and right-eye perspective frames.

[0110] Phase D: Delay Compensation (Asynchronous Time Warp (ATW)).

[0111] D1: Obtain the latest predicted pose.

[0112] Read the latest IMU data at the current time t_real and extrapolate it to obtain the predicted head pose Pose_pred at t_real + Δ (Δ = processing delay budget, set to 20ms).

[0113] D2: Determine if ATW is needed.

[0114] If a new rendered image is just generated in step C (with timestamp t_render) and t_render < t_real (that is, the image is slightly lagged), perform ATW; otherwise, skip this step.

[0115] D3: Perform ATW distortion.

[0116] For each pixel (u,v) in the left-eye image, calculate its corresponding world direction D_old according to the old pose Pose_render; re-project D_old to the new image coordinates (u_new, v_new) according to the newly predicted pose Pose_pred; adopt inverse mapping: for each pixel in the target image, search for its corresponding position in the old image and sample the color; the formula is: P_new= K * R_pred^T * ( R_render * (K^-1 * p_old) ), wherein K is the intrinsic parameter matrix of the virtual camera; generate the distorted left-eye image Left_ATW. Process the right-eye image in the same way.

[0117] Stage E: Encoding and wireless transmission.

[0118] E1: Video compression Splice the two images Left_ATW and Right_ATW side by side into a 1280×720 side-by-side format (or top-bottom format); use an H.265 hardware encoder, control the bit rate at 30 Mbps, and set the GoP structure to all I-frames (each frame is encoded independently to avoid dependency delay).

[0119] E2: Wireless transmission.

[0120] Package the encoded frame data and send it via a WiGig module in OFDM modulation mode; send the timestamp and pose metadata corresponding to this frame at the same time.

[0121] Stage F: Reception and display by AR glasses.

[0122] F1: Reception and decoding.

[0123] The physical layer of the AR glasses receives electromagnetic waves and demodulates the data packet; the H.265 hardware decoder decodes the data to restore the side-by-side left and right images.

[0124] F2: Splitting and optical projection.

[0125] Send the left-eye and right-eye images to the left and right light guide display chips respectively; the display refresh timing dynamically adjusts the display timing of the next frame according to the difference between the received timestamp and the local clock of the glasses to maintain smoothness.

[0126] F3: Closed-loop feedback.

[0127] The new sampling points of the glasses IMU are immediately transmitted back to the domain controller via BLE for pose prediction in the next frame.

[0128] Process 3: Exception Handling and Degradation Strategies (Parallel Monitoring Task). Triggering conditions: A camera suddenly loses signal / A sensor data is interrupted / The algorithm timeout occurs.

[0129] Step 3.1: Fault classification.

[0130] Level 1 Fault (Critical): All cameras lose all signal > the last valid image is frozen and displayed, with a voice prompt "System failure, please drive carefully"; Level 2 Fault (Partial Failure): If the top-view camera has no signal, the view of the roof is replaced by a low-quality image synthesized from the surrounding cameras, and a yellow warning box is displayed in the corner of the field of view; Level 3 Fault (Non-Critical): If the UWB signal is briefly lost, the EKF continues to work relying solely on the IMU + depth sensor, with reduced positioning accuracy but no system crash.

[0131] Step 3.2: Recovery mechanism.

[0132] The device will retry the faulty device every second. If it succeeds three times in a row, it will automatically return to full-function mode.

[0133] The system algorithm processing flow of this invention effectively solves the problems of perspective image position shift and visual dizziness caused by dynamic changes in head posture and system rendering delay in the prior art. It achieves end-to-end low-latency and high-precision in-vehicle augmented reality perspective effect, significantly improving the passenger immersion experience and the robustness of the system under different fault conditions.

[0134] To implement the method of the above embodiments, such as Figure 2 As shown, the present invention proposes a vehicle augmented reality perspective device 10, comprising: The pose fusion module 100 is used to acquire the initial six-degree-of-freedom pose data of the passenger in the vehicle coordinate system, and to fuse the observation information of multiple source sensors to correct the pose data in order to obtain the real-time head pose.

[0135] The line-of-sight penetration module 200 is used to determine the virtual line-of-sight direction of the passenger's eyes based on the real-time head position and calculate the intersection of the line-of-sight direction with the mesh model of the vehicle's outer surface as the virtual penetration point.

[0136] The perspective synthesis module 300 is used to generate a perspective image on a virtual projection surface centered on the passenger's eyes through viewpoint guidance mapping, based on image data collected from virtual penetration points and multiple vehicle-mounted cameras.

[0137] The correction rendering module 400 is used to perform distortion correction and multi-band fusion on the perspective image, and perform asynchronous time distortion compensation in combination with the predicted head pose, and output the final image to the display device for rendering.

[0138] The device of this invention effectively solves the problems of perspective image position shift and visual dizziness caused by dynamic head pose changes and rendering delays in the prior art, and achieves low-latency, high-precision in-vehicle augmented reality perspective effect, significantly improving passenger immersion experience and system robustness.

[0139] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0140] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for augmented reality perspective of vehicles, characterized in that, include: The initial six-DOF pose data of the passenger in the vehicle coordinate system is obtained, and the pose data is corrected by fusing observation information from multiple sources of sensors to obtain the real-time head pose. The virtual gaze direction of the passenger's eyes is determined based on the real-time head position and posture, and the intersection of this gaze direction with the vehicle's outer surface mesh model is calculated as the virtual penetration point. Based on image data collected from virtual penetration points and multiple vehicle cameras, a perspective image on a virtual projection surface centered on the passenger's eyes is generated through viewpoint guidance mapping. The perspective image is subjected to distortion correction and multi-band fusion, and asynchronous time warp compensation is performed in combination with the predicted head pose. The final image is then output to the display device for rendering.

2. The method as described in claim 1, characterized in that, The process of acquiring the passenger's initial six-DOF pose data in the vehicle coordinate system and then fusing observation information from multiple sources to correct the pose data to obtain real-time head pose includes: The triaxial acceleration and triaxial angular velocity data collected by the built-in IMU of the AR glasses are input into the DSP core of the domain controller and fed into the constant acceleration and constant angular velocity kinematic model. The state vector and covariance matrix at time t are then recursively predicted as prior estimates. The distance observations measured by the UWB base station, the matching results of ORB feature points extracted by the fisheye camera of the AR glasses and the cockpit feature map, and the fitted center position of the head point cloud output by the depth sensor are sequentially input into the observation update module of the extended Kalman filter to form a multi-source observation vector. The predicted values ​​of each observation are calculated through the observation equation for subsequent sequential correction. Based on the extended Kalman filter algorithm, the Kalman gain is adjusted using the dynamic information matrix to sequentially update the prior estimates, eliminating absolute position drift and relative pose error, and obtaining smooth six-degree-of-freedom pose data with a frequency of 100Hz.

3. The method as described in claim 2, characterized in that, The update process of the extended Kalman filter includes: Construct a twelve-dimensional system state vector containing position, attitude angles, and their rates of change: The observation noise covariance matrix is ​​adjusted in real time based on the UWB signal-to-noise ratio and the confidence level of the depth sensor data. ; Execute the Kalman update formula to output the corrected real-time head pose. Covariance Matrix : 。 4. The method as described in claim 1, characterized in that, The step of determining the virtual gaze direction of the passenger's eyes based on real-time head posture, and calculating the intersection of this gaze direction with the vehicle's outer surface mesh model as the virtual penetration point, includes: Real-time head pose data is input into the gaze generation module, which then parses the coordinates of the passenger's left eye. and right eye coordinates And a gaze direction vector is given for each eye. Space rays are constructed using a ray generation algorithm: 。 5. The method as described in claim 1, characterized in that, The image data collected based on virtual penetration points and multiple vehicle-mounted cameras is used to generate a perspective image on a virtual projection surface centered on the passenger's eyes through viewpoint-guided mapping, including: The pre-constructed virtual projection sphere centered on the passenger's eyes, along with the calibrated spatial pose matrices and projection parameters of each onboard camera, are input into the viewpoint guidance mapping unit. The mapping unit then maps each discrete direction vector on the virtual projection sphere. Inverse solution for the corresponding three-dimensional points outside the vehicle in the relevant direction And using the projection equation calculate Projecting coordinates onto the pixel plane of each camera to obtain the set of multi-camera pixel coordinates for each direction: Calculate the orientation of each camera pixel pair Mixed weights: in, It is a pixel The corresponding optical center line of sight direction, Center of the image; Based on the aforementioned weighted average, the pixel color values ​​of each camera in the corresponding direction are weighted to obtain the final color value in the corresponding direction on the virtual projection surface, thus forming a perspective image.

6. The method as described in claim 1, characterized in that, The process of performing distortion correction and multi-band fusion on the perspective image, and combining the predicted head pose with asynchronous time warp compensation, before outputting the final image to the display device for rendering includes: The raw image captured by the fisheye camera is input into the radial distortion model to calculate the correction coordinates: in, , These are the coordinates on the normalized image plane. These are the distortion coefficients, obtained through offline calibration. The corrected left and right eye images are decomposed into Laplacian pyramids, and Gaussian weight gradient is applied to the overlapping areas on different pyramid layers to perform multi-band fusion, resulting in a seamless perspective image. The predicted head position is obtained by extrapolating the latest IMU data at the current moment. If the rendered image timestamp is delayed, the pixel orientation under the old pose is mapped to the image coordinates under the new pose using the reprojection formula, performing asynchronous time warp compensation; whereby the reprojection formula is: The compensated image is encoded and output to AR smart glasses via a wireless link for waveguide display.

7. The method as described in claim 1, characterized in that, Also includes: Based on multi-frame images from a checkerboard calibration board, the intrinsic parameters and distortion coefficients of each camera are calculated using the Zhang Zhengyou calibration method. For fisheye cameras with a field of view exceeding a preset threshold, a distortion correction model suitable for fisheye models is used for fitting to generate intrinsic parameter files. Place the vehicle in an open area, place visual markers at preset positions outside the vehicle, control all cameras to synchronously acquire images containing the visual markers, use the PnP algorithm to calculate the rotation matrix and translation vector of each camera relative to the vehicle coordinate system, and generate an external parameter file. Based on the extrinsic parameter file and the vehicle CAD external surface mesh model, the optimal camera index and its pixel coordinates corresponding to each virtual line of sight are pre-calculated according to the preset angle sampling step size. The mapping relationship is stored as a lookup table and loaded into the GPU constant memory.

8. The method as described in claim 1, characterized in that, Image data is acquired using multiple vehicle-mounted cameras, including: The front-view camera, rear-view camera, left / right-view camera, top-view camera, and bottom-view camera send video streams to the domain controller via the GMSL link. The time synchronization unit uses a precise time protocol to stamp each data stream with the same microsecond-level timestamp, outputting time-aligned image data.

9. The method as described in claim 1, characterized in that, After outputting the final image to a display device for rendering, the method further includes: Real-time monitoring of the status of input camera signals, sensor data integrity, and algorithm calculation time; When the loss of all camera signals exceeds a preset threshold, a Level 1 fault is triggered; the last valid image is frozen and a voice prompt is output. When either the top-view or bottom-view camera loses signal, a level two fault is triggered; a low-quality alternative image is synthesized using the remaining four cameras, and a yellow warning box is superimposed in the corner of the field of view; When the UWB signal is briefly lost, a level 3 fault is triggered; the extended Kalman filter is controlled to continue outputting pose based solely on IMU and depth sensor data, while retrying the faulty device every second. If it succeeds three times in a row, it will automatically return to full-function mode.

10. A vehicle augmented reality perspective device, characterized in that, include: The pose fusion module is used to acquire the initial six-degree-of-freedom pose data of the passenger in the vehicle coordinate system, and to fuse the observation information of multiple source sensors to correct the pose data in order to obtain the real-time head pose. The line-of-sight penetration module is used to determine the virtual line-of-sight direction of the passenger's eyes based on the real-time head position and calculate the intersection of this line-of-sight direction with the vehicle's outer surface mesh model as the virtual penetration point. The perspective synthesis module is used to generate perspective images on a virtual projection surface centered on the passenger's eyes through viewpoint guidance mapping, based on image data collected from virtual penetration points and multiple vehicle cameras. The correction rendering module is used to perform distortion correction and multi-band fusion on the perspective image, and perform asynchronous time distortion compensation in combination with the predicted head pose, and output the final image to the display device for rendering.