Vehicle-mounted audio augmented reality system and method based on virtual-real fusion space anchoring

By using an in-vehicle audio augmented reality system based on virtual-real spatial anchoring, the problem of inconsistency between the audio prompt sound source and the movement trajectory of visual elements has been solved, realizing intuitive spatial directionality of vehicle status alarms and improving the driver's perception consistency and driving safety.

CN121957331APending Publication Date: 2026-05-01CHINA FAW CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA FAW CO LTD
Filing Date
2025-12-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing in-vehicle interactive systems, the location of audio prompts cannot be synchronized with the movement trajectory of augmented reality dynamic visual elements in real time, and vehicle status alarm signals lack intuitive spatial directionality for the physical location of the fault, resulting in increased cognitive load and low information transmission efficiency for drivers in emergency situations.

Method used

An in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring is adopted. Through controllers, augmented reality display devices, distributed speaker arrays and sensor groups, visual-spatial mapping and semantic-spatial mapping are realized. The coordinates of three-dimensional virtual sound sources are dynamically calculated, and the distributed speaker array is driven to synthesize virtual sound sources, ensuring the synchronization and clear spatial directionality of audio and visual elements in three-dimensional space.

Benefits of technology

It achieves synchronization between virtual sound source trajectory and dynamic visual image, allowing drivers to quickly locate the source of the fault through hearing without visual confirmation, reducing cognitive load, improving information transmission efficiency and driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121957331A_ABST
    Figure CN121957331A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted audio augmented reality system and method based on virtual-real fusion space anchoring, and relates to the technical field of intelligent cabin man-machine interaction, and the system comprises a controller, an augmented reality display device, a distributed loudspeaker array, and a sensor group. The controller receives the virtual image pixel coordinates, the eyeball position coordinates and the vehicle state semantic identifier, and stores a cabin three-dimensional digital model. The system converts a dynamic pixel coordinate into a three-dimensional virtual sound source coordinate by utilizing perspective inverse projection through vision-space mapping logic; the retrieval model anchors the semantic identifier to the physical component coordinates through the semantic-space mapping logic. The audio rendering module calculates a speaker drive gain based on the virtual sound source coordinates to synthesize a virtual sound source. By constructing a unified cabin three-dimensional digital model and a double-channel mapping mechanism, spatial alignment of an audio sound image, an AR visual track and a physical part of a vehicle body is realized, and audio-visual perception splitting is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent cockpit human-computer interaction technology, and in particular to an in-vehicle audio augmented reality system and method based on virtual-real fusion spatial anchoring. Background Technology

[0002] With the rapid development of automotive intelligent technology, the intelligent cockpit has become an important interactive space connecting people and vehicles. The widespread adoption of augmented reality head-up displays and panoramic surround sound systems has greatly enriched the driver's visual and auditory experience. Current in-vehicle human-machine interaction systems typically focus on overlaying dynamic virtual navigation guidance or driver assistance information onto the windshield or display screen, and playing voice prompts or alarm sounds through speakers, attempting to build a multimodal interactive environment.

[0003] However, in existing in-vehicle interaction architectures, the visual and auditory channels often operate independently, lacking deep spatial integration. While augmented reality (AR) displays can project navigation arrows or collision warning icons to specific physical locations in front of the driver's line of sight and dynamically shift with vehicle movement, the accompanying audio cues typically rely solely on traditional channel mixing technology for simple left-right balance adjustments, or are fixed to be emitted from the driver's side speaker. This approach prevents the perceived location of sound from moving in real-time with the pixel movement of the virtual image, creating a spatial disconnect between the driver's visual target location and the location of the sound source. This inconsistency between visual and auditory perception not only weakens the immersive experience of augmented reality but also increases the driver's cognitive load in emergency situations, leading to hesitation in judgment.

[0004] Furthermore, for vehicle status monitoring alerts (such as abnormal tire pressure, open doors, and vehicles approaching from blind spots), existing technologies primarily use flashing dashboard icons combined with a general beeping sound or synthesized voice announcement. These audio alerts typically lack clear spatial directionality; the sound is often diffuse or limited to a fixed area, failing to establish a direct acoustic connection with the actual physical component causing the malfunction. After hearing the alarm, the driver still needs to shift their gaze to confirm the text or icons on the dashboard to determine the specific location of the fault (e.g., whether it's the left rear tire or the right blind spot). This reliance on visual secondary confirmation reduces the efficiency of information transmission and can easily distract the driver in complex traffic flow.

[0005] Meanwhile, traditional in-vehicle audio rendering is mostly based on a fixed speaker channel layout, and its sound field calibration is usually targeted at a preset ideal listening position. However, during actual driving, the driver's head position and gaze direction are dynamically changing. Existing systems lack a mechanism for correcting sound source coordinates based on the driver's real-time eye-tracking data, making it difficult to maintain accurate anchoring of the virtual sound source relative to the real physical environment when the listener's head deflects or shifts. In summary, how to construct an in-vehicle audio system that can accurately spatially align dynamic visual elements, the semantics of vehicle physical components, and the three-dimensional sound field has become a technical problem to be solved in the field of intelligent cockpit technology. Summary of the Invention

[0006] The purpose of this invention is to provide an in-vehicle audio augmented reality system and method based on virtual-real fusion spatial anchoring, which at least solves the technical problems of existing in-vehicle interactive systems where the location of audio prompt sound sources cannot be synchronized in real time with the motion trajectory of augmented reality dynamic visual elements, and where vehicle status alarm signals lack intuitive spatial directionality for the physical location of faults.

[0007] This invention provides the following solution:

[0008] The first aspect of the present invention provides an in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring. The system is applied to a vehicle and includes a controller, an augmented reality display device, a distributed speaker array, and a sensor group.

[0009] The controller establishes communication connections with the augmented reality display device, the distributed speaker array, and the sensor group. Internally, the controller includes an information input module, a spatial mapping and anchoring module, and an audio rendering and output module. The information input module is configured to receive pixel coordinates of virtual images from the augmented reality display device, eye position coordinates from the sensor group, and semantic identifiers of vehicle status signals, and to read the three-dimensional digital model of the vehicle cockpit.

[0010] The three-dimensional digital model defines a cockpit spatial coordinate system as the system's reference frame, a screen coordinate system describing the position of virtual images, and spatial layout data containing the physical coordinate vectors of each speaker unit in the distributed speaker array. The origin of the cockpit spatial coordinate system is set at the center of the driver's head or the position of their eyeballs.

[0011] The spatial mapping and anchoring module is configured to perform visual-spatial mapping and semantic-spatial mapping operations, converting the pixel coordinates or semantic identifiers into three-dimensional virtual sound source coordinates in the cockpit spatial coordinate system. During the visual-spatial mapping operation, the spatial mapping and anchoring module obtains the pixel coordinates of the virtual image in the screen coordinate system, calls a pre-stored inverse projection matrix and rotation transformation matrix, and converts the pixel coordinates into direction vectors in the cockpit spatial coordinate system. The inverse projection matrix is ​​generated based on the projection parameters of the augmented reality display device, and the rotation transformation matrix is ​​generated based on the installation posture of the augmented reality display device relative to the cockpit spatial coordinate system. The system normalizes the direction vector and multiplies it by a preset sound source mapping distance coefficient to obtain a relative position vector. Then, the relative position vector is vector-superimposed with the eye position coordinates input by the sensor group to generate the three-dimensional virtual sound source coordinates. The sound source mapping distance coefficient is configured as the virtual image imaging distance value of the augmented reality display device, or as a variable that dynamically changes with vehicle speed. In addition, the spatial mapping and anchoring module can be configured with trajectory smoothing processing logic, which uses a filter to process the three-dimensional virtual sound source coordinate sequence calculated at consecutive time points to generate a continuous three-dimensional sound source movement trajectory.

[0012] During the semantic-spatial mapping operation, the spatial mapping and anchoring module retrieves the semantic identifier from the semantic-spatial mapping database contained in the three-dimensional digital model. The semantic-spatial mapping database stores the mapping relationship between the semantic identifiers of vehicle components and the geometric center coordinates or acoustic radiation center coordinates of the vehicle components in the cockpit spatial coordinate system. The system extracts the retrieved geometric center coordinates or acoustic radiation center coordinates and assigns them to the three-dimensional virtual sound source coordinates.

[0013] The audio rendering and output module is configured to receive the coordinates of the three-dimensional virtual sound source and audio content data, and encapsulate the audio content data into an audio object containing spatial location metadata. This module calculates the driving gain of each speaker unit in the distributed speaker array based on the coordinates of the three-dimensional virtual sound source, and drives the distributed speaker array to synthesize a virtual sound source at the coordinates of the three-dimensional virtual sound source. The audio rendering and output module performs beamforming operations based on amplitude translation to determine the unit direction vector from the center of the listening position to the coordinates of the three-dimensional virtual sound source. It selects a subset of speaker units from the distributed speaker array to participate in the rendering, and calculates the driving gain coefficient of each speaker unit in the subset, such that the weighted vector sum of the unit direction vector of each speaker unit relative to the center of the listening position and its respective driving gain coefficient approximates the unit direction vector from the center of the listening position to the coordinates of the three-dimensional virtual sound source. The center of the listening position can be configured as a dynamic listening reference point updated in real time based on the eye position coordinates. The audio rendering and output module can also be configured to perform crosstalk cancellation processing, generating an anti-phase cancellation signal using a pre-stored transfer function for speaker units located on the driver's headrest or near the ear.

[0014] A second aspect of the present invention provides an in-vehicle audio augmented reality method based on virtual-real fusion spatial anchoring. This method is applied to an in-vehicle domain controller or audio processing unit and includes the following steps:

[0015] The system performs parallel acquisition and type discrimination of multi-source information, receives dynamic visual signals from the augmented reality display device and static semantic signals from the vehicle sensor network, and reads the driver's eye position coordinates in real time.

[0016] The visual-spatial mapping logic is executed for the dynamic visual signal. Using the perspective inverse projection principle and depth anchoring strategy, combined with the eye position coordinates, the three-dimensional virtual sound source coordinates of the virtual image corresponding to the dynamic visual signal in the cockpit space coordinate system are calculated.

[0017] The semantic-spatial mapping logic is executed on the static semantic signal. Based on the semantic identifier contained in the static semantic signal, the three-dimensional digital model of the cockpit is retrieved, and the geometric center coordinates of the corresponding vehicle physical components are extracted as the coordinates of the three-dimensional virtual sound source.

[0018] Perform an object-based audio encapsulation step to bind the audio content to be played with the coordinates of the three-dimensional virtual sound source to generate an audio object.

[0019] The sound field rendering and physical output steps are performed. Based on the coordinates of the three-dimensional virtual sound source and the physical topology data of the distributed speaker array, the driving parameters of each speaker unit in the distributed speaker array are calculated, and the distributed speaker array is driven to synthesize a virtual sound source.

[0020] The above solution achieves the following beneficial technical effects:

[0021] This application performs a visual-spatial mapping operation based on perspective inverse projection and depth anchoring to convert the two-dimensional pixel coordinates output by the augmented reality display device into three-dimensional spatial sound source coordinates relative to the driver's eye position in real time. This achieves synchronization between the virtual sound source trajectory and the dynamic virtual image projected onto the windshield in three-dimensional space, eliminating the disconnect between the sense of direction and visual information inherent in traditional in-vehicle audio, and ensuring that the driver's perceived auditory orientation remains highly consistent with their visual gaze direction.

[0022] This application utilizes a 3D digital model of the cockpit containing the physical coordinates of vehicle components to establish a static anchoring relationship between vehicle status semantic identifiers and the geometric centers or acoustic radiation centers of physical components through semantic-spatial mapping logic. This allows warning sounds such as blind spot monitoring and tire pressure warnings to be emitted directly from the actual physical location of the fault or event, enabling drivers to quickly locate the source of the abnormality through auditory instinct without having to look at the dashboard, thereby shortening emergency response time and improving driving safety.

[0023] This application employs object-based audio encapsulation and dynamic listening reference point update technology, combined with amplitude translation beamforming algorithm to drive a distributed speaker array. This not only decouples the audio content from the physical channels, but also dynamically corrects the starting point for calculating the sound source direction vector based on the driver's eye position fed back by sensors in real time. This effectively compensates for sound image positioning deviations caused by changes in the driver's sitting posture or head movements, and constructs a stable, clear, and spatially directional virtual sound field in the complex cabin acoustic environment. Attached Figure Description

[0024] Figure 1 This is a system architecture block diagram of the present invention;

[0025] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0026] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] See attached document Figure 1 , Figure 1This is a schematic diagram of an in-vehicle audio augmented reality system architecture based on virtual-real fusion and spatial anchoring, according to an embodiment of the present invention. The present invention provides an in-vehicle audio augmented reality system based on virtual-real fusion and spatial anchoring, which is applied to a vehicle. Its hardware architecture mainly includes: a controller, an augmented reality display device, a distributed speaker array, and a sensor group. The controller establishes communication connections with the augmented reality display device, the distributed speaker array, and the sensor group. The controller is configured to execute spatially anchored audio rendering logic to coordinate visual information and sound field reproduction.

[0028] Augmented reality (AR) display devices are configured to project virtual images into the forward field of view of a vehicle. In one embodiment, the AR display device is an AR head-up display (HUD) whose projection plane covers the area of ​​the vehicle's windshield. In another embodiment, the AR display device includes a transparent side window display or a central control display with AR view functionality. The AR display device establishes a two-dimensional screen coordinate system. This coordinate system uses a predetermined point on the display plane (such as the upper left corner) as its origin, and defines the horizontal direction as... The axis, perpendicular to the direction is Axis. Augmented reality display devices output virtual images in real time on the screen coordinate system. pixel coordinates To the controller.

[0029] The distributed speaker array comprises multiple speaker units positioned at various physical locations within the vehicle's cabin. These units cover the A-pillars, doors, headliner, headrests, and center console area. The system is pre-configured with a three-dimensional cabin spatial coordinate system. The coordinate system has its origin at the center of the driver's head or the position of their eyeballs, and the vehicle's direction of travel is... Axle, vehicle left side The axis, vertically upward direction is Axis. The coordinate system of each speaker unit in the distributed speaker array within the cockpit space. Each of them has a definite physical coordinate vector. ,in Indicates the first One speaker unit.

[0030] The sensor suite includes an eye-tracking sensor and a vehicle status sensor. The eye-tracking sensor is configured to acquire the driver's eye position and gaze direction in real time and generate eye position coordinates. Vehicle status sensors are connected to the vehicle bus and configured to acquire operational status data and fault signals (such as tire pressure status and blind spot monitoring signals) of various vehicle components. Each type of status signal is associated with a unique semantic identifier. .

[0031] The controller is equipped with a spatially anchored audio rendering engine, which is logically divided into: an information input module, a spatial mapping and anchoring module, and an audio rendering and output module.

[0032] The information input module is configured to receive pixel coordinates from the virtual image of the augmented reality display device and eye position coordinates from the sensor array. and semantic identifiers of vehicle status signals In addition, the information input module stores a three-dimensional digital model of the vehicle cabin, which includes the physical coordinate vectors of each speaker unit in the distributed speaker array. And the key physical components of the vehicle (including but not limited to tires, windows, and air conditioning vents) in the cabin space coordinate system. Geometric center coordinates in .

[0033] The spatial mapping and anchoring module is configured to convert input visual or semantic information into a cockpit spatial coordinate system. 3D virtual sound source coordinates For dynamic visual elements from augmented reality display devices, the spatial mapping and anchoring module performs a visual-spatial mapping operation. This operation utilizes the principle of perspective inverse projection, based on pixel coordinates. Calculate the position of the virtual sound source in three-dimensional space. The spatial mapping and anchoring module executes the following formula to determine the coordinates of the virtual sound source. :

[0034] ;

[0035] in: This represents the calculated virtual sound source in the cockpit space coordinate system. 3D coordinate vector ; This represents the driver's eye position vector input from the sensor array; This represents the preset sound source mapping distance coefficient, used to define the depth distance of the virtual sound source relative to the driver; This represents the inverse of the projection transformation matrix of the augmented reality display device, used to convert pixel coordinates into direction vectors in the view coordinate system; This represents the coordinate system from the line-of-sight coordinate system to the cockpit space coordinate system. The rotation transformation matrix is ​​used to correct the angle between the line of sight and the vehicle body direction; Represents the virtual image in the screen coordinate system homogeneous coordinate vector in ; The Euclidean norm operation for vectors is used to normalize direction vectors.

[0036] For vehicle status signals from the sensor array, the spatial mapping and anchoring module performs semantic-spatial mapping operations. This module bases its operations on the received semantic identifiers. The system retrieves the geometric center coordinates of the corresponding vehicle physical components from the 3D digital model stored in the information input module. And directly assign this coordinate to the virtual sound source coordinate. This establishes a static anchoring relationship between semantic information and physical space.

[0037] The audio rendering and output module is configured to receive the virtual sound source coordinates output by the spatial mapping and anchoring module. And the corresponding audio content data. The audio rendering and output module encapsulates the audio content into an audio object containing spatial location metadata, and then renders it according to the virtual sound source coordinates. Calculate the drive gain of each speaker unit in the distributed speaker array. The audio rendering and output module performs amplitude-shift-based beamforming calculations, and calculates the gain coefficient of each speaker unit. The following vector superposition relationship is satisfied:

[0038] ;

[0039] in: This indicates the coordinates from the center of the listening position to the virtual sound source. The unit direction vector; This indicates the number of speaker units involved in rendering, which is a subset of the distributed speaker array; Indicates the first The driving gain coefficient of each speaker unit; Indicates the first The unit direction vector of each speaker unit relative to the center of the listening position.

[0040] Through the coordinated operation of the above modules, the controller drives the distributed speaker array to move at the virtual sound source coordinates. The synthesized virtual sound source ensures that the location of the generated sound is spatially consistent with the location of the virtual image displayed on the augmented reality display device or the actual location of the vehicle's physical components.

[0041] To achieve spatial alignment between visual elements and auditory perception, this invention constructs a unified three-dimensional digital model of the vehicle cabin and strictly defines a coordinate reference system in this model.

[0042] The three-dimensional digital model of the cockpit first defines a global cockpit spatial coordinate system. The cockpit's spatial coordinate system It is a right-handed Cartesian coordinate system, with its origin at... The driver's head center position is set at the driver's seat design reference point or calibrated by an eye-tracking system. Within the cockpit space coordinate system. In the middle, the definition The positive direction of the axis points directly forward in the direction of vehicle travel, defined as... The positive direction of the axis points to the left side of the vehicle body, defined as follows: The positive direction of the axis points vertically upwards towards the vehicle roof. The position of any point in space within the cabin is determined by a three-dimensional coordinate vector in this coordinate system. This coordinate system is uniquely determined. It serves as the system's reference system and is used for all geometric transformations, sound source localization calculations, and physical component position calibrations in subsequent steps.

[0043] The system also defines a screen coordinate system that is attached to the augmented reality display device. The screen coordinate system This is a two-dimensional planar coordinate system used to describe the position of the virtual image on the display imaging surface. The top-left corner of the display imaging surface is set as the origin, and the horizontal direction to the right is defined as... The positive axis, defined as the vertically downward direction. Positive axis. Each virtual pixel generated by an augmented reality display device corresponds to a set of coordinates. During the initial calibration phase, the system obtains the screen coordinate system using vehicle design parameters. Relative to the cockpit space coordinate system The rotation and translation matrices and projection parameters are used to establish the rigid body transformation relationship between the two-dimensional pixel plane and the three-dimensional physical space.

[0044] The 3D digital model of the cockpit contains pre-built spatial layout data for a distributed speaker array. The system abstracts each physical speaker unit as a sound emission point and, based on the vehicle's computer-aided design data, extracts the geometric center of each speaker unit in the cockpit's spatial coordinate system. The precise coordinates within. This coordinate data is stored as a location set. ,in Indicates the first Three-dimensional coordinate vector of each speaker unit This set of locations serves as the boundary condition for beamforming calculations in subsequent audio rendering algorithms.

[0045] The 3D digital model of the cockpit further includes a semantic-spatial mapping database. This database establishes semantic identifiers for vehicle components. physical space coordinates The system establishes an index relationship between components. For point-like components (such as blind spot indicators and door locks), the system stores their geometric center coordinates; for area-like components (such as windows, tires, and hoods), the system stores the coordinates of the area's geometric centroid or acoustic radiation center. For example, the identifier for the left rear tire maps to the center of the left rear wheel hub. The coordinates in the database are the center coordinates of the passenger-side window mapped to the passenger-side window plane. This database is stored in the controller's non-volatile memory, allowing the controller to directly retrieve the corresponding 3D spatial target point based on the input semantic signal without performing real-time geometric probing.

[0046] The system also dynamically maintains a listening reference point. In the default mode, Coincident at the origin of the cockpit space coordinate system In dynamic mode, the system uses the eye position coordinates input from the sensor array. Real-time updates The numerical value. This listening reference point. Used to define the optimal listening position in the audio rendering algorithm, ensuring that the direction vector calculation of the virtual sound source is always based on the driver's current actual head position, thereby correcting the sound image positioning deviation caused by changes in the driver's sitting posture.

[0047] See attached document Figure 2 , Figure 2 This is a flowchart of an in-vehicle audio augmented reality method based on virtual-real fusion spatial anchoring according to an embodiment of the present invention. The method is executed by an in-vehicle domain controller or audio processing unit, and aims to achieve consistency between audio cues and their associated visual elements or physical components in three-dimensional space. The method includes sequentially executed steps of data acquisition and classification, spatial coordinate mapping calculation, audio object encapsulation, and sound field rendering output.

[0048] First, the system performs parallel acquisition and type discrimination of multi-source information. The controller receives input signals in real time through the vehicle communication interface and divides them into two processing streams based on the data structure characteristics of the signals: the first stream is dynamic visual signals, which originate from the augmented reality display device and contain the real-time pixel coordinates of the virtual image in the screen coordinate system. The second category is static semantic signals, which originate from the vehicle sensor network or body controller. These signals contain semantic identifiers that characterize the state of vehicle components or the type of fault. At the same time, the system continuously reads the driver's current eye position coordinates through an eye-tracking interface. , which serves as the dynamic origin for subsequent spatial calculations.

[0049] Next, the system performs spatial coordinate mapping calculation steps according to the signal type. For the first type of dynamic visual signal, the system initiates the visual-spatial mapping logic. This logic combines the two-dimensional pixel coordinates in the screen coordinate system with the current eye position coordinates, and uses a perspective inverse projection geometry algorithm to calculate a spatial ray originating from the driver's eye and passing through the screen pixels. Based on a preset depth strategy (such as projecting onto a virtual image plane or a fixed-distance sphere), the system determines the coordinates of a three-dimensional virtual sound source in the cockpit spatial coordinate system along this ray. The calculation process continues as the virtual image shifts on the screen, making... In three-dimensional space, a sound source movement trajectory is formed that is synchronized with the visual trajectory.

[0050] For the second type of static semantic signal, the system initiates semantic-spatial mapping logic. This logic does not perform geometric projection calculations but instead executes a database retrieval operation. The controller then determines the target based on the received semantic identifier. The system locates the corresponding vehicle physical component index in the pre-set 3D digital model of the cockpit and extracts the geometric center coordinates or acoustic radiation center coordinates of that component. The system then directly assigns the extracted physical coordinates as the coordinates of the 3D virtual sound source. This allows us to pinpoint the location of the sound transmission as the actual physical location where the fault or event occurred.

[0051] Subsequently, the system performs an object-based audio encapsulation step. The controller compares the audio content to be played (including waveform data or synthesized speech stream) with the calculated coordinates of the three-dimensional virtual sound source. Binding is performed to generate a separate audio object. The data structure of this audio object contains the audio payload and spatial location metadata that changes over time or remains constant.

[0052] Finally, the system performs sound field rendering and physical output steps. The audio rendering engine receives the audio object and calculates the driving parameters of each speaker unit based on the physical topology data of the distributed speaker array in the cockpit. This calculation process applies beamforming or vector amplitude translation algorithms, based on the virtual sound source coordinates. Relative to the direction vector of the listening position, the gain coefficient and phase delay of each speaker unit in the array are analyzed. The controller uses these parameters to drive the speaker array to produce sound, and through the physical superposition of multiple sound waves, the target coordinates within the cockpit space are determined. It reconstructs a clear phantom sound source, completing closed-loop control from information input to spatialized sound field output.

[0053] This embodiment details how the system processes dynamic visual signals from augmented reality display devices and converts pixel motion on a two-dimensional screen into sound source trajectories in three-dimensional space in real time.

[0054] In this process, the system first establishes a geometric solution model of the viewing direction. The controller periodically acquires the pixel coordinates of the virtual object in the current frame output by the AR display device. This coordinate represents the position of virtual cue symbols (such as navigation arrows) on the imaging plane. Simultaneously, the controller synchronously acquires the real-time position of the driver's primary eye in the cockpit space coordinate system, as fed back by the eye-tracking sensor. To construct the spatial vector pointing from the observer to the virtual object, the system calls the inverse projection matrix stored in memory. and rotation transformation matrix .in Generated based on the optical intrinsic parameters (such as focal length and optical axis center) of AR display devices. The system generates the calibration based on the device's mounting posture (rotation angle) in the vehicle body coordinate system. The system will then use pixel coordinates... The coordinates are converted to homogeneous coordinates, and then inverse projection and rotation transformations are performed sequentially to restore the two-dimensional points to three-dimensional spatial direction vectors in the cockpit coordinate system.

[0055] Subsequently, the system performs depth anchoring calculations for the virtual sound source coordinates. The inverse projection operation yields a normalized direction vector, representing the line-of-sight direction of the virtual object seen by the driver, but its specific depth position in three-dimensional space remains undetermined. To ensure a natural sound experience and avoid spatial conflicts with real objects, the system introduces a depth coefficient. To determine the final virtual sound source coordinates In one configuration mode, the depth coefficient The distance is set to be equal to the virtual image imaging distance of the AR-HUD (e.g., 7.5 meters or 10 meters), so that the position of the sound emission coincides with the position of the optical virtual image of the virtual image, achieving the strongest audiovisual fusion. In another configuration mode, the depth coefficient... It is set as a variable that changes dynamically with vehicle speed; as the vehicle speed increases, Increasing the value makes the sound seem farther away, aligning with the driver's psychological expectation of focusing on distant objects while traveling at high speeds. The system then multiplies the calculated direction vector by this depth coefficient. And superimposed the driver's eye position vector This allows us to obtain the absolute coordinates of the virtual sound source in the cockpit coordinate system. .

[0056] This embodiment also features a specially designed trajectory smoothing and synchronization mechanism. Due to pixel coordinates... and eyeball position If the sampling rate differs, or if the signal itself contains high-frequency jitter, the directly calculated sound source coordinates will jump. Therefore, the system introduces a Kalman filter or low-pass filter at the coordinate output. The filter adjusts the coordinates calculated at consecutive time points. The coordinate sequence is smoothed to filter out minute displacements caused by sensor noise, generating a smooth and continuous three-dimensional sound source movement trajectory. This smoothing ensures that when the AR navigation arrow smoothly moves across the windshield from left to right, the driver's audible prompts also move continuously and smoothly in space, without any discontinuities or jumps in the auditory experience. The smoothed real-time coordinate data is instantly transmitted to the audio rendering module, ensuring that the lag time of sound position changes is below the threshold perceptible to the human ear, achieving a seamless integration of visual and auditory dynamics.

[0057] This embodiment details how the system processes vehicle status signals that lack screen coordinate information and converts them into spatial audio with clear physical directionality.

[0058] In this embodiment, the core logic of the system lies in establishing a mapping bridge between abstract semantics and physical space. The controller first listens for and parses vehicle status messages via the vehicle bus (such as CANFD or Ethernet). The system internally predefines a semantic library containing various vehicle events, and each event (such as tire pressure warning, door open, blind spot monitoring warning) is assigned a unique semantic identifier. When the system detects a specific state bit flip or a threshold trigger (for example, the left rear tire pressure sensor reading is below a safe threshold), the controller immediately generates a corresponding semantic identifier. Then start the mapping process.

[0059] Unlike processing dynamic visual signals, this embodiment employs a deterministic lookup table mapping mechanism. The controller accesses a pre-stored 3D digital model database of the cockpit. This database stores a semantic-coordinate mapping table. This table not only records simple point-to-point relationships but also includes acoustic localization strategies for complex physical components. For the semantic meaning of abnormal pressure in the left rear tire, the system does not simply point to the geometric center of the left rear wheel, but instead, based on a preset acoustic strategy, queries the coordinates of the optimal acoustic projection point of this component inside the cockpit. For example, considering the vehicle's sound insulation structure, tires pointing directly outwards cause blurred sound and image. Therefore, the system sets the target coordinates to the inner area below the left rear door panel, a location that is strongly associated with the left rear tire in auditory perception. For the right blind spot warning semantics, the coordinates retrieved by the system are set to the area behind the right B-pillar or C-pillar to realistically simulate the threat direction of vehicles approaching from behind; or set to the right rearview mirror position to guide the driver's visual observation of the rearview mirror.

[0060] Obtain target physical coordinates The system then directly locked it as the virtual sound source coordinates. The system maintains this coordinate until the alarm signal is cleared. To enhance the warning effect, the system also modulates the sound diffusion parameter based on the semantic type. For precisely located faults such as an open door, the system sets a smaller diffusion, making the sound appear as a point source, precisely pointing to the door lock. For regional threats such as a vehicle in a blind spot, the system sets a larger diffusion, making the sound appear to cover the entire side and rear area, increasing the sense of urgency and the warning range. Ultimately, these audio objects with fixed spatial coordinates and specific diffusion parameters are sent to the rendering engine, allowing the driver to instantly determine the specific location of the fault based solely on auditory instinct, without needing to look at the dashboard, achieving highly efficient diagnosis through sound localization.

[0061] This embodiment elaborates on how the system converts audio objects carrying spatial coordinates into multi-channel electrical signals that drive physical loudspeakers, thereby reconstructing a virtual sound field in real physical space.

[0062] This process begins with the encapsulation phase of the audio object. The audio rendering and output module receives the audio content data (PCM stream) from the upstream module, as well as the coordinates of the three-dimensional virtual sound source calculated in real time. The system packages these two elements, along with the sound's metadata (including the sound source's type tag, volume gain reference, diffusion coefficient, etc.), into a standardized audio object. This object-based format decouples the audio content from its playback location, meaning that the same prompt tone can be dynamically rendered to any location in the vehicle as needed, without being limited by pre-recorded channel formats.

[0063] Subsequently, the system enters the core sound field rendering and calculation stage. To synthesize a clear sound image in locations without physical speakers (such as the center of the windshield or a specific point on the left rear door panel), the system employs vector amplitude translation or beamforming algorithms. The controller first determines the coordinates of the virtual sound source, centered on a listening reference point (usually the driver's head). Target sound source direction vector Next, the system searches and selects a subset of physical loudspeakers (e.g., three loudspeakers that form the smallest triangular region surrounding the target direction) that are closest to the target direction in the preset loudspeaker array topology.

[0064] The system calculates the drive gain coefficient of each loudspeaker in the subset based on the principle of acoustic image localization. The computational logic follows the vector composition rule: the direction vectors of each participating speaker... Its gain The weighted vector sum should approximate the direction vector of the target sound source as closely as possible in terms of direction. That is, satisfying Meanwhile, to ensure a constant total sound power and avoid sudden fluctuations in sound image size during movement, the sum of the squares of all gain coefficients must be normalized to 1. For unselected loudspeakers, their gain coefficient is set to zero.

[0065] To further enhance the focus and spatial isolation of the sound image, the system incorporates crosstalk cancellation processing on top of gain calculations. For speaker units positioned on the driver's headrest or near the ears, the controller utilizes the known transfer function from the speaker to both ears to generate an inverse cancellation signal and use it to correct the output of each channel. This effectively attenuates the left speaker signal that would otherwise crosstalk to the driver's right ear, thereby enhancing the sound pressure level difference and time difference between the two ears. For far-field speakers positioned further away, the system primarily relies on the energy focusing characteristics of beamforming to enhance the sense of localization. Through these processes, the localization of the virtual sound source becomes sharper and more realistic. Finally, the gain-weighted and signal-processed multi-channel audio streams are sent to a power amplifier to drive the speaker array distributed throughout the vehicle to work together, forming a stable and precisely positioned virtual sound source in the listener's perception.

[0066] This section further clarifies the working state and interaction logic of the system of the present invention in actual operation through specific driving scenarios.

[0067] The dynamic anchoring process in augmented reality navigation assistance. Assume a vehicle is traveling on a multi-lane urban road and is about to turn right at an intersection 500 meters ahead. The AR-HUD system projects a blue virtual navigation arrow onto the corresponding location on the windshield. As the vehicle approaches the intersection, the arrow visually moves from the center of the field of vision to the right. At this moment, the system of this invention captures the pixel trajectory of the arrow in real time and drives the audio rendering engine to generate a synchronously moving virtual sound source. The driver not only sees the arrow sliding to the right but also hears a voice prompt to turn right or a specific navigation prompt tone, the sound location of which smoothly transitions precisely from directly in front to the right front. This high degree of overlap between visual and auditory trajectories strongly reinforces the intention to turn right at the sensory level, effectively preventing the driver from taking the wrong lane at complex intersections.

[0068] The static semantic anchoring process in vehicle status alarms. When the left rear tire of a vehicle experiences a rapid drop in tire pressure due to a nail puncture, traditional dashboard alarms often require the driver to look down to identify the icon. With this invention, the vehicle chassis sensors detect the anomaly and trigger a semantic signal indicating low left rear tire pressure. The system immediately queries the cockpit model and locates the physical coordinates below the left rear door as the sound source. Subsequently, a rapid alarm tone or voice announcement of abnormal left rear tire pressure sounds directly from the driver's left rear. The driver does not need to look at the dashboard; they can instantly determine the exact location of the fault simply by the direction of the sound, enabling a faster decision to slow down and pull over, thus improving emergency response efficiency and driving safety.

[0069] Spatial directionality is applied in intelligent voice interaction. When a driver asks the voice assistant via voice command, "Is the passenger-side window not closed properly?" or the voice assistant proactively prompts, "Please note that the passenger-side window is not fully closed," the voice assistant's response is no longer broadcast throughout the vehicle. Instead, it uses spatial rendering technology to position its voice specifically at the passenger-side window. This interaction method gives the AI ​​assistant a sense of presence, making it seem as if it is sitting in the passenger seat or right next to the window, greatly enhancing the naturalness, intuitiveness, and technological immersion of human-computer interaction.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring, characterized in that, include: Controller, augmented reality display device, distributed speaker array, and sensor array; The controller establishes communication connections with the augmented reality display device, the distributed speaker array, and the sensor group, respectively. The controller is internally equipped with an information input module, a spatial mapping and anchoring module, and an audio rendering and output module. The information input module is configured to receive pixel coordinates of virtual images from the augmented reality display device, eye position coordinates from the sensor group, and semantic identifiers of vehicle status signals, and to read the three-dimensional digital model of the vehicle cabin. The spatial mapping and anchoring module is configured to perform visual-spatial mapping operations and semantic-spatial mapping operations to convert the pixel coordinates or the semantic identifier into three-dimensional virtual sound source coordinates in the cockpit spatial coordinate system. The audio rendering and output module is configured to receive the coordinates of the three-dimensional virtual sound source and the audio content data, encapsulate the audio content data into an audio object containing spatial location metadata, calculate the driving gain of each speaker unit in the distributed speaker array based on the coordinates of the three-dimensional virtual sound source, and drive the distributed speaker array to synthesize a virtual sound source at the coordinates of the three-dimensional virtual sound source.

2. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 1, characterized in that, The three-dimensional digital model stored in the information input module is defined as follows: The cockpit space coordinate system, with its origin set at the center of the driver's head or eye position, serves as the reference system for the system's geometric transformation and sound source localization calculations. The screen coordinate system, which is attached to the augmented reality display device, describes the position of the virtual image on the display imaging surface; And pre-constructed spatial layout data of the distributed speaker array, the spatial layout data including the physical coordinate vector of each speaker unit in the distributed speaker array in the cockpit spatial coordinate system.

3. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 2, characterized in that, When performing the visual-spatial mapping operation, the spatial mapping and anchoring module is configured to perform the following operations: Obtain the pixel coordinates of the virtual image in the screen coordinate system; The pre-stored inverse projection matrix and rotation transformation matrix are invoked to convert the pixel coordinates into direction vectors in the cockpit space coordinate system; The inverse projection matrix is ​​generated based on the projection parameters of the augmented reality display device, and the rotation transformation matrix is ​​generated based on the installation attitude of the augmented reality display device relative to the cockpit space coordinate system. The direction vector is normalized and multiplied by a preset sound source mapping distance coefficient to obtain the relative position vector; The relative position vector is vector-superimposed with the eyeball position coordinates input by the sensor group to generate the three-dimensional virtual sound source coordinates.

4. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 3, characterized in that, The sound source mapping distance coefficient is configured as the virtual image imaging distance value of the augmented reality display device; Alternatively, the sound source mapping distance coefficient can be configured as a variable that changes dynamically with vehicle speed, increasing as vehicle speed increases.

5. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 3, characterized in that, The spatial mapping and anchoring module is also equipped with trajectory smoothing processing logic, which uses a Kalman filter or a low-pass filter to filter the three-dimensional virtual sound source coordinate sequence calculated at consecutive time points to generate a continuous three-dimensional sound source movement trajectory.

6. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 1, characterized in that, When performing the semantic-spatial mapping operation, the spatial mapping and anchoring module is configured to perform the following operations: Based on the received semantic identifier, a retrieval is performed in the semantic-spatial mapping database contained in the three-dimensional digital model; The semantic-spatial mapping database stores the mapping relationship between the semantic identifiers of vehicle components and the geometric center coordinates or acoustic radiation center coordinates of the vehicle components in the cockpit spatial coordinate system. Extract the retrieved geometric center coordinates or acoustic radiation center coordinates, and assign the geometric center coordinates or acoustic radiation center coordinates to the three-dimensional virtual sound source coordinates.

7. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 1, characterized in that, The audio rendering and output module is configured to perform beamforming operations based on amplitude translation, the beamforming operations specifically including: Determine the unit direction vector pointing from the center of the listening position to the coordinates of the three-dimensional virtual sound source; Select a subset of speaker units from the distributed speaker array to participate in the rendering; Calculate the driving gain coefficient of each loudspeaker unit in the subset, such that the weighted vector sum of the unit direction vector of each loudspeaker unit relative to the center of the listening position and its respective driving gain coefficient approximates the unit direction vector pointing from the center of the listening position to the coordinates of the three-dimensional virtual sound source.

8. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 7, characterized in that, The hearing position center is configured as a dynamic hearing reference point; the system is configured to update the value of the dynamic hearing reference point according to the eye position coordinates input in real time by the sensor group, and correct the calculation starting point of the unit direction vector.

9. The in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring according to claim 7, characterized in that, The audio rendering and output module is also configured to perform crosstalk cancellation processing; for speaker units located on the driver's headrest or near the ears, an anti-phase cancellation signal is generated using a pre-stored speaker-to-ear transfer function to correct the output signal of the corresponding channel.

10. A vehicle-mounted audio augmented reality method based on virtual-real fusion spatial anchoring, characterized in that, Based on the in-vehicle audio augmented reality system based on virtual-real fusion spatial anchoring as described in any one of claims 1 to 9, the method is applied to an in-vehicle domain controller or audio processing unit, including: Perform multi-source information parallel acquisition and type discrimination steps, receive dynamic visual signals from augmented reality display devices and static semantic signals from vehicle sensor networks, and read the driver's eye position coordinates in real time; For the dynamic visual signal, a visual-spatial mapping logic is executed. Using the perspective inverse projection principle and depth anchoring strategy, combined with the eyeball position coordinates, the three-dimensional virtual sound source coordinates of the virtual image corresponding to the dynamic visual signal in the cockpit space coordinate system are calculated. The semantic-spatial mapping logic is executed on the static semantic signal, and the three-dimensional digital model of the cockpit is retrieved based on the semantic identifier contained in the static semantic signal. The geometric center coordinates of the corresponding vehicle physical components are extracted as the coordinates of the three-dimensional virtual sound source. Perform an object-based audio encapsulation step to bind the audio content to be played with the coordinates of the three-dimensional virtual sound source to generate an audio object; The sound field rendering and physical output steps are performed. Based on the coordinates of the three-dimensional virtual sound source and the physical topology data of the distributed speaker array, the driving parameters of each speaker unit in the distributed speaker array are calculated, and the distributed speaker array is driven to synthesize a virtual sound source.

Citation Information

Cited By

  • A voice and visual interaction control method for safe driving

    CN122275933A