Realistic soundscape augmentation method, system, apparatus, and medium for a vehicle

By acquiring multi-source environmental data for environmental semantic analysis and 3D audio rendering algorithms, channel signals for the in-vehicle speaker array are generated, solving the problem of environmental data not being fused in real time in the in-vehicle audio system. This enables an immersive soundscape experience and personalized audio control, thus improving the user experience.

CN122496769APending Publication Date: 2026-07-31CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA FAW CO LTD
Filing Date
2026-04-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing in-vehicle audio systems lack the ability to deeply integrate and intelligently analyze external environmental data perceived by the vehicle as a real-time input source, resulting in low interactivity of the in-vehicle audio experience and affecting the user experience.

Method used

By acquiring multi-source environmental data, performing environmental semantic analysis, constructing a virtual driving environment, identifying the current vehicle's driving scene, generating multiple sound objects based on the soundscape rule library, and using 3D audio rendering algorithms to generate channel signals that drive the in-vehicle speaker array, thereby rendering the corresponding soundscape for passengers.

Benefits of technology

It achieves dynamic mapping between sound and the real-time physical environment, providing an immersive experience that unifies the sound field and the visual field, improving the interactivity and personalization of in-vehicle audio, reducing reliance on visual screens, and meeting the interaction needs of the era of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496769A_ABST
    Figure CN122496769A_ABST
Patent Text Reader

Abstract

This application provides a method, system, device, and medium for enhancing the realistic soundscape of a vehicle, belonging to the field of vehicle technology. The method includes: performing environmental semantic analysis on multi-source environmental data to construct a virtual driving environment and identify the current vehicle's driving scene; using real-time environmental semantic extraction technology based on multi-sensor fusion as the core input for soundscape generation to solve the problem of sound being disconnected from the real-time physical environment; based on the driving scene, querying a pre-set soundscape rule library to generate multiple sound objects, mapping the dynamic changes of the external environment onto the acoustic layer in real time, providing an immersive experience with a unified sound field and visual field; based on the driving scene, using a pre-set 3D audio rendering algorithm to generate signals for each channel of the in-vehicle speaker array, rendering corresponding soundscapes for passengers in the corresponding seating areas. Precise spatial audio rendering brings a sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other, enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a method, system, device, and medium for enhancing the real-world soundscape of a vehicle. Background Technology

[0002] With the rapid development of autonomous driving technology, the importance of in-vehicle audio systems is becoming increasingly prominent. Existing in-vehicle audio technologies mainly focus on two directions: one is improving sound quality, such as creating an immersive surround sound experience through more speaker channels and advanced sound field algorithms (such as Dolby Atmos); the other is noise reduction and communication, such as using active noise cancellation technology to offset road and wind noise, and using beamforming technology to improve the clarity of voice assistant interaction and calls.

[0003] In existing technologies, traditional in-vehicle audio architectures are closed loops: the audio source (music, navigation prompts) is output to the audio processor, and then to the speaker. This system does not deeply integrate and intelligently analyze the vehicle's perceived external environmental data as a core, real-time input source. It lacks an intelligent middleware that can understand environmental semantics and generate or modulate audio accordingly, resulting in low interactivity of the in-vehicle audio experience and affecting the user experience. Summary of the Invention

[0004] The main objective of this application is to provide a method, system, device, and medium for enhancing the real-world soundscape of a vehicle, in order to solve one or more technical problems existing in the prior art, and at least provide a beneficial option or create conditions.

[0005] To achieve the above objectives, one aspect of this application proposes a method for enhancing the real-world soundscape of a vehicle. The method includes: acquiring multi-source environmental data, performing environmental semantic analysis on the multi-source environmental data, constructing a virtual driving environment, and identifying the current driving scenario of the vehicle. Based on the driving scenario, the system queries the established soundscape rule library to determine multiple sound elements corresponding to the driving scenario, and generates multiple sound objects based on the multiple sound elements. Based on the driving scenario, the playback positions of multiple sound objects in the virtual driving environment are determined. According to the corresponding playback positions and corresponding sound objects, the 3D audio rendering algorithm is used to generate signals for each channel of the in-vehicle speaker array, so as to render the corresponding sound scene for passengers in the corresponding seating area.

[0006] Furthermore, the step of performing environmental semantic analysis on the multi-source environmental data to construct a virtual driving environment and identify the current vehicle's driving scenario includes: Based on the multi-source environmental data, dynamic and static features within the designated area are identified, and the relative positions of the dynamic and static features with the current vehicle are continuously tracked. Using the established environmental semantic model, the multi-source environmental data is analyzed to generate semantic soundscape mapping labels corresponding to the dynamic features and the static features; The virtual driving environment is constructed based on the relative position, the dynamic features, the static features, and the semantic soundscape mapping labels, and the driving scene is identified.

[0007] Furthermore, the step of querying the established soundscape rule base to determine multiple sound elements corresponding to the driving scenario, and generating multiple sound objects based on the multiple sound elements, includes: Based on the driving scenario, the semantic soundscape mapping labels of the dynamic features, and the semantic soundscape mapping labels of the static features, the set soundscape rule base is queried to determine multiple sound elements corresponding to the driving scenario, the dynamic features, and the static features; Using the provided audio synthesis software, based on the sound elements corresponding to the driving scenario, the sound elements corresponding to the dynamic features and the sound elements corresponding to the static features are synthesized to obtain the sound object of the dynamic features and the sound object of the static features.

[0008] Furthermore, the step of generating the signals for each channel of the in-vehicle speaker array based on the corresponding playback position and the corresponding sound object using the established 3D audio rendering algorithm includes: Based on the relative positions of the dynamic and static features in the driving scene, the playback position is determined according to the relative positions, and the playback volume of the corresponding sound object is rendered using the set 3D audio rendering algorithm according to the playback position. Based on the playback position, a playback area corresponding to the sound object is determined. Based on the playback area and the playback volume, the in-vehicle speaker array to be invoked is determined to generate the corresponding channel signal.

[0009] To achieve the above objectives, another aspect of this application proposes a vehicle-based realistic soundscape enhancement system, the system comprising: The scene recognition module is used to acquire multi-source environmental data, perform environmental semantic analysis on the multi-source environmental data, construct a virtual driving environment, and identify the current vehicle's driving scene. The soundscape generation module is used to query the set soundscape rule library according to the driving scenario, determine multiple sound elements corresponding to the driving scenario, and generate multiple sound objects based on the multiple sound elements; The soundscape zoning rendering module is used to determine the playback positions of multiple sound objects in the virtual driving environment based on the driving scene, and generate the signals of each channel of the in-vehicle speaker array using the set 3D audio rendering algorithm according to the corresponding playback positions and corresponding sound objects, so as to render the corresponding soundscape for passengers in the corresponding seating area.

[0010] Furthermore, the scene recognition module includes: The location recognition module is used to identify dynamic and static features within a set area based on the multi-source environmental data, and to continuously track the relative positions of the dynamic and static features with the current vehicle. The environmental semantic analysis module is used to analyze the multi-source environmental data using the established environmental semantic model, and generate semantic soundscape mapping labels corresponding to the dynamic features and the static features. The scene construction module is used to construct the virtual driving environment and identify the driving scene based on the relative position, the dynamic features, the static features and the semantic soundscape mapping label.

[0011] Furthermore, the soundscape generation module includes: The mapping module is used to query the set sound scene rule base according to the driving scene, the semantic sound scene mapping label of the dynamic feature and the semantic sound scene mapping label of the static feature, and determine multiple sound elements corresponding to the driving scene, the dynamic feature and the static feature; The object generation module is used to synthesize, using the provided audio synthesis software, the sound elements corresponding to the dynamic features and the sound elements corresponding to the static features based on the sound elements corresponding to the driving scene, to obtain the sound object of the dynamic features and the sound object of the static features.

[0012] Furthermore, the soundscape partitioning rendering module includes: The volume rendering module is used to determine the playback position based on the relative positions of the dynamic features and the static features in the driving scene, and to render the playback volume of the corresponding sound object according to the playback position using the set 3D audio rendering algorithm. The partition rendering module is used to determine the playback area corresponding to the sound object based on the playback position, and to determine the in-vehicle speaker array to be invoked based on the playback area and the playback volume, so as to generate the corresponding channel signal.

[0013] To achieve the above objectives, another aspect of the present application provides a vehicle control device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described method for enhancing the real-world soundscape of a vehicle.

[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for enhancing the real-world soundscape of a vehicle.

[0015] The embodiments of this application include at least the following beneficial effects: This application provides a method, system, device, and medium for enhancing the real-world soundscape of a vehicle. This solution acquires multi-source environmental data, performs environmental semantic analysis on the multi-source environmental data, constructs a virtual driving environment, and identifies the current driving scene of the vehicle. Based on multi-sensor fusion environmental semantic real-time extraction technology, it uses this as the core input for soundscape generation, solving the problem of sound being disconnected from the real-time physical environment. According to the driving scene, it queries the established soundscape rule base to determine multiple sound elements corresponding to the driving scene, generates multiple sound objects based on these sound elements, and maps the dynamic changes of the external environment onto the acoustic layer in real time, providing an immersive experience that unifies the sound field and the field of view, and defining the physical world's... The system establishes a set of rules and execution mechanisms for the dynamic mapping relationship between images / events and virtual sound elements. Based on driving scenarios, it determines the playback positions of multiple sound objects in the virtual driving environment. According to the corresponding playback positions and sound objects, it uses a set 3D audio rendering algorithm to generate signals for each channel of the in-vehicle speaker array, rendering corresponding soundscapes for passengers in the corresponding seating areas. Precise spatial audio rendering brings a sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other. It provides a new way of presenting information, conveying environmental information to passengers in a more natural and non-intrusive way, reducing reliance on visual screens, better meeting the interaction needs of the autonomous driving era, improving the interactivity of in-vehicle audio experience, and enhancing the user experience. Attached Figure Description

[0016] Figure 1 This is a flowchart of the vehicle real-world sound enhancement method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the framework of the vehicle real-world soundscape enhancement system provided in this application embodiment; Figure 3 This is a schematic diagram of the hardware structure framework of the vehicle control device provided in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0018] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, Ethernet signaling information may also be referred to as interface signaling information, and similarly, interface signaling information may also be referred to as Ethernet signaling information. Depending on the context, the words "if" or "when" as used herein may be interpreted as "when," "in response to a determination," or "in the event of a determination."

[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] In some embodiments of one aspect of the present invention Figure 1 This is an optional flowchart of the vehicle real-world sound enhancement method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S100 to S300.

[0022] Step S100: Acquire multi-source environmental data, perform environmental semantic analysis on the multi-source environmental data, construct a virtual driving environment, and identify the current vehicle's driving scenario.

[0023] Step S200: Based on the driving scenario, query the set soundscape rule library to determine multiple sound elements corresponding to the driving scenario, and generate multiple sound objects based on the multiple sound elements.

[0024] Step S300: Based on the driving scenario, determine the playback positions of multiple sound objects in the virtual driving environment. According to the corresponding playback positions and corresponding sound objects, use the set 3D audio rendering algorithm to generate signals for each channel of the in-vehicle speaker array, so as to render the corresponding sound scene for passengers in the corresponding seating area.

[0025] Steps S100 to S300 as illustrated in this embodiment involve acquiring multi-source environmental data, performing environmental semantic analysis on the multi-source environmental data, constructing a virtual driving environment, and identifying the current vehicle's driving scenario. Based on multi-sensor fusion environmental semantic real-time extraction technology, this is used as the core input for soundscape generation, solving the problem of sound being disconnected from the real-time physical environment. According to the driving scenario, a pre-defined soundscape rule base is queried to determine multiple sound elements corresponding to the driving scenario. Based on these multiple sound elements, multiple sound objects are generated, and the dynamic changes of the external environment are mapped in real-time onto the acoustic layer, providing an immersive experience that unifies the sound field and field of view. This defines the relationship between physical world objects / events and virtual sound elements. The rules set and execution mechanism of dynamic mapping relationships; based on the driving scenario, the playback positions of multiple sound objects in the virtual driving environment are determined. According to the corresponding playback positions and corresponding sound objects, the established 3D audio rendering algorithm is used to generate signals for each channel of the in-vehicle speaker array, rendering the corresponding soundscape for passengers in the corresponding seating area. Precise spatial audio rendering brings a sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other, providing a new way of information presentation, conveying environmental information to passengers in a more natural and non-intrusive way, reducing reliance on visual screens, better meeting the interaction needs of the autonomous driving era, improving the interactivity of in-vehicle audio experience, and enhancing the user experience.

[0026] In some embodiments of S100, multi-source environmental data is acquired.

[0027] The multi-source environmental data includes: visual data of the surrounding environment collected by the vision module, lidar data of the surrounding environment collected by the lidar module, and location information provided by the GPS positioning module. The multi-source environmental data may also include other data, which are not specifically limited in this application.

[0028] Object recognition and tracking are performed based on multi-source environmental data to determine the scene features around the current vehicle and track the relative position of these scene features with respect to the current vehicle. Scene features include static and dynamic features.

[0029] Environmental semantic analysis is performed based on multi-source environmental data, transforming sensor data into an environmental model with semantic labels. This involves converting sensor data into semantic soundscape mapping labels and assigning these labels based on scene features. This forms the foundation for dynamically associating sound with reality.

[0030] Based on scene features, relative position, and semantic soundscape mapping labels, a virtual driving environment is constructed using multi-source environmental data to determine the current driving scenario of the vehicle (e.g., highway cruising, urban congestion, passing through a tunnel, or driving in the rain).

[0031] In some embodiments of S200, the established soundscape rule library is invoked based on the driving scenario.

[0032] The established soundscape rule base defines the mapping relationship between external physical objects / events and sound elements.

[0033] The scene features identified by S100 are used as external physical objects or events. The semantic soundscape mapping labels of the scene features are used to query the mapping relationship of the set soundscape rule base to determine the sound element corresponding to each scene feature.

[0034] Since a scene feature may have multiple sound elements, a real-time audio synthesizer is used to dynamically generate sound objects, thereby improving the flexibility of audio playback. This results in a sound object corresponding to a scene feature, and multiple sound objects corresponding to multiple scene features.

[0035] Through a customizable soundscape rule library, creative mapping between sound and environment is achieved, rather than simple reproduction. This solves the problem of static sound effects being disconnected from the environment and transforms the external world into a canvas that can be given different narrative meanings, thus enabling a highly personalized and diverse experience.

[0036] In some embodiments of S300, the playback positions of multiple sound objects in the virtual driving environment are determined based on the relative positions of scene features in the driving scene.

[0037] By using the playback position and sound object of each scene feature, the 3D audio rendering algorithm is used to calculate and generate the signals of each channel of the in-vehicle speaker array, so as to render the corresponding sound scene for the passengers in the corresponding seating area.

[0038] For example, if a scene feature of a truck approaching at high speed 5 meters to the left is identified, it is mapped to a sound element of a howling wind coming from the left through the established sound scene rule library. This sound element is then synthesized with another background sound element to obtain the sound object of the scene feature. Based on the relative position of the scene feature, the established 3D audio rendering algorithm is used to calculate and generate a sound object that drives the speaker array in the rear left seating area to play the sound, making it sound like it is coming from the rear left.

[0039] Anchoring virtual soundscapes in physical space, precise spatial audio rendering enhances the sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other.

[0040] In some embodiments of this invention, S100, the identification of the driving environment specifically includes the following steps: S110 identifies dynamic and static features within the designated area based on multi-source environmental data, and continuously tracks the relative positions of the dynamic and static features with the current vehicle.

[0041] S120 utilizes the established environmental semantic model to analyze multi-source environmental data and generate semantic soundscape mapping labels corresponding to dynamic and static features.

[0042] S130 constructs a virtual driving environment and identifies driving scenarios based on relative position, dynamic features, static features, and semantic soundscape mapping labels.

[0043] In some embodiments of S110, the visual model is used to perform object recognition and tracking on the visual data in the multi-source environmental data, and to identify scene features within the set area.

[0044] Scene features include static features and dynamic features. Static features include buildings and specific landmarks; dynamic features include surrounding vehicles and pedestrians.

[0045] The defined area is a detection circle with the current vehicle as the center and a set radius. Features within this detection circle can be considered scene features, and the detection circle is the defined area, in order to identify the type of surrounding vehicles, new pedestrians, and specific landmarks.

[0046] Using the SLAM algorithm, LiDAR data and location information from multi-source environmental data are processed to initially construct a virtual driving environment centered on the current vehicle. In this scenario, the relative position and speed of static features with respect to the current vehicle, as well as the relative position and speed of dynamic features with respect to the current vehicle, are continuously tracked.

[0047] In some embodiments of S120, environmental semantic analysis is performed on visual data, radar and laser data, and location information based on the initially constructed virtual driving environment.

[0048] Using the established environmental semantic model, the sensor data contained in the current scene features are transformed into an environmental model with semantic labels, that is, the sensor data is transformed into semantic soundscape mapping labels, thereby obtaining the semantic soundscape mapping labels of the current scene features.

[0049] The environmental semantic model is a pre-trained environmental semantic model. Visual data, radar and laser data, and location information of scene features are used as input features and fed into the pre-trained environmental semantic model. The pre-trained environmental semantic model outputs the semantic soundscape mapping label of the scene features.

[0050] For example, a dynamic feature is identified: a truck. The relative position of this dynamic feature (truck) to the current vehicle is identified as: 5 meters to the left, and its speed is: high speed. Using the aforementioned visual data, radar / laser data, and location information, the data is input into a trained environmental semantic model. The trained environmental semantic model outputs a semantic soundscape mapping label for this dynamic feature as: "There is a truck approaching at high speed 5 meters to the left." This semantic meaning is the prerequisite for subsequently mapping it to "a whistling sound of wind approaching from the left."

[0051] The environmental semantic model can be an environment model set up based on a large visual model or 3D object detection using Transformer / BEV, or other environmental models, and is not limited in this application.

[0052] In some embodiments of S130, the initially constructed virtual driving environment is refined based on relative position, dynamic features, static features, and semantic soundscape mapping labels, and the current driving scenario of the vehicle is determined to obtain the identified driving scenario.

[0053] Driving scenarios include: highway cruising, urban traffic congestion, driving through tunnels, and driving in the rain. Other driving scenarios may also be included, but will not be detailed in this application.

[0054] In some embodiments of this invention, S200 specifically includes the following steps for acquiring the sound object: S210: Based on the semantic soundscape mapping labels of driving scene, dynamic features and static features, query the established soundscape rule base to determine multiple sound elements corresponding to driving scene, dynamic features and static features.

[0055] S220 uses the provided audio synthesis software to synthesize sound elements corresponding to dynamic features and sound elements corresponding to static features based on sound elements corresponding to the driving scene, thereby obtaining sound objects with dynamic features and sound objects with static features.

[0056] In some embodiments of S210, the sound elements corresponding to the dynamic features are determined by querying the set sound scene rule base through the semantic sound scene mapping tags of dynamic features.

[0057] The established soundscape rule base defines the mapping relationship between external physical objects / events and sound elements.

[0058] The identified dynamic and static features are used as external physical objects or events. The semantic soundscape mapping labels of scene features are used to query the mapping relationship of the set soundscape rule base to determine the sound element corresponding to each feature.

[0059] Based on the current driving scenario, the system queries the established soundscape rule library to determine the current background sound elements, so as to achieve a creative mapping between sound and environment, rather than a simple reproduction.

[0060] For example, when other vehicles with dynamic features are identified, they are mapped to the rustling sound of wind blowing through treetops, with the volume and rhythm of the sound proportional to the relative speed of the vehicle. When static landmarks are identified and a vehicle is approaching, they are mapped to a pre-recorded educational explanation of the landmark. In some embodiments of S220, since a scene feature may have multiple sound elements, the sound elements corresponding to the driving scene are merged into the audio synthesis software to obtain a sound object corresponding to a scene feature, thus resulting in multiple sound objects corresponding to multiple scene features. This improves the flexibility of audio playback.

[0061] The proposed audio synthesis software can be a real-time audio synthesizer, dynamically generating audio with greater flexibility.

[0062] Through a customizable soundscape rule library, creative mapping between sound and environment is achieved, rather than simple reproduction. This solves the problem of static sound effects being disconnected from the environment and transforms the external world into a canvas that can be given different narrative meanings, thus enabling a highly personalized and diverse experience.

[0063] In some embodiments of this invention, S300, the identification of the driving environment specifically includes the following steps: S310 determines the playback position based on the relative positions of dynamic and static features in the driving scene, and then uses the set 3D audio rendering algorithm to render the playback volume of the corresponding sound object based on the playback position.

[0064] S320 determines the playback area corresponding to the sound object based on the playback position, and determines the in-vehicle speaker array to be invoked based on the playback area and playback volume, in order to generate the corresponding channel signal.

[0065] In some embodiments of S310, based on the current driving scenario, the relative position of the dynamic feature and the current vehicle is determined. Using this relative position, the playback position of the current dynamic feature in the virtual driving environment is determined. The established 3D audio rendering algorithm is used to calculate and render the sound object at this playback position, and the playback volume is adjusted.

[0066] For example, based on the current driving scenario of high-speed cruising, the relative position of the dynamic feature truck and the current vehicle is determined: 5 meters to the left. Based on this relative position, the playback position of the dynamic feature truck in the virtual driving environment is determined to achieve precise positioning. The 3D audio rendering algorithm is used to calculate and render the sound object with the sound of the car engine and the howling wind at the playback position, so that the playback volume is presented as if it is played from far to near.

[0067] Similarly, based on the current driving scenario, the relative position of the static feature and the current vehicle is determined. This relative position is then used to determine the playback position of the static feature within the virtual driving environment. The established 3D audio rendering algorithm is then used to calculate and render the sound object at this playback position, adjusting the playback volume.

[0068] The proposed 3D audio rendering algorithm can be either high-order Ambisonics or object-based audio rendering.

[0069] In some embodiments of S320, the playback area corresponding to the sound object is determined again by the playback position, so as to determine the speaker array that the current vehicle needs to call. Based on the playback area, the speaker array that the current vehicle needs to call is determined. Based on the called speaker array, the corresponding channel signal is generated so as to be able to play the sound object.

[0070] Reference Figure 2 Another embodiment of this application also provides a vehicle real-world soundscape enhancement system, including: a scene recognition module, a soundscape generation module, and a soundscape partition rendering module.

[0071] The scene recognition module can acquire multi-source environmental data, perform environmental semantic analysis on the multi-source environmental data, construct a virtual driving environment, and identify the current vehicle's driving scene.

[0072] The soundscape generation module can query the set soundscape rule library according to the driving scenario, determine multiple sound elements corresponding to the driving scenario, and generate multiple sound objects based on the multiple sound elements.

[0073] The soundscape zoning rendering module can determine the playback positions of multiple sound objects in the virtual driving environment based on the driving scene. According to the corresponding playback positions and corresponding sound objects, it uses the set 3D audio rendering algorithm to generate the signals of each channel of the in-vehicle speaker array, so as to render the corresponding soundscape for passengers in the corresponding seating area.

[0074] In this embodiment, the scene recognition module acquires multi-source environmental data.

[0075] The multi-source environmental data includes: visual data of the surrounding environment collected by the vision module, lidar data of the surrounding environment collected by the lidar module, and location information provided by the GPS positioning module. The multi-source environmental data may also include other data, which are not specifically limited in this application.

[0076] The scene recognition module can identify and track objects based on multi-source environmental data, determine the scene features around the current vehicle, and track the relative position of these scene features with respect to the current vehicle. Scene features include static features and dynamic features.

[0077] The scene recognition module can perform environmental semantic analysis based on multi-source environmental data, transforming sensor data into an environmental model with semantic labels. This involves converting sensor data into semantic sound-scene mapping labels and assigning these labels based on scene features. This is the foundation for dynamically associating sound with reality.

[0078] The scene recognition module can construct a virtual driving environment based on scene features, relative position, and semantic sound and scene mapping labels, using multi-source environmental data, and determine the current driving scenario of the vehicle (e.g., highway cruising, urban congestion, passing through a tunnel, driving in the rain).

[0079] The soundscape generation module can call the set soundscape rule library based on the driving scenario.

[0080] The established soundscape rule base defines the mapping relationship between external physical objects / events and sound elements.

[0081] The soundscape generation module can use the identified scene features as external physical objects or events, utilize the semantic soundscape mapping tags of the scene features, query the mapping relationship of the set soundscape rule base, and determine the sound element corresponding to each scene feature.

[0082] Since a scene feature may have multiple sound elements, the sound scene generation module can dynamically generate sound objects through a real-time audio synthesizer, improving the flexibility of audio playback. This results in a sound object corresponding to one scene feature, and multiple sound objects corresponding to multiple scene features.

[0083] Through a customizable soundscape rule library, creative mapping between sound and environment is achieved, rather than simple reproduction. This solves the problem of static sound effects being disconnected from the environment and transforms the external world into a canvas that can be given different narrative meanings, thus enabling a highly personalized and diverse experience.

[0084] The soundscape partitioning rendering module can determine the playback position of multiple sound objects in the virtual driving environment based on the relative positions of scene features in the driving scene.

[0085] The soundscape zoning rendering module can calculate the playback position and sound object of each scene feature using the set 3D audio rendering algorithm to generate the signals of each channel driving the in-vehicle speaker array, so as to render the corresponding soundscape for passengers in the corresponding seating area.

[0086] For example, if a scene feature of a truck approaching at high speed 5 meters to the left is identified, it is mapped to a sound element of a howling wind coming from the left through the established sound scene rule library. This sound element is then synthesized with another background sound element to obtain the sound object of the scene feature. Based on the relative position of the scene feature, the established 3D audio rendering algorithm is used to calculate and generate a sound object that drives the speaker array in the rear left seating area to play the sound, making it sound like it is coming from the rear left.

[0087] Anchoring virtual soundscapes in physical space, precise spatial audio rendering enhances the sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other.

[0088] In another embodiment of this application, the scene recognition module includes: a location recognition module, an environmental semantic analysis module, and a scene construction module.

[0089] The location recognition module can identify dynamic and static features within a set area based on multi-source environmental data, and continuously track the relative position of the dynamic and static features with the current vehicle.

[0090] The environmental semantic analysis module can use the established environmental semantic model to analyze multi-source environmental data and generate semantic soundscape mapping labels corresponding to dynamic and static features.

[0091] The scene construction module can build a virtual driving environment and identify driving scenarios based on relative position, dynamic features, static features and semantic soundscape mapping labels.

[0092] In this embodiment, the location recognition module can use the set visual model to perform object recognition and tracking on visual data in multi-source environmental data, and identify scene features within the set area.

[0093] Scene features include static features and dynamic features. Static features include buildings and specific landmarks; dynamic features include surrounding vehicles and pedestrians.

[0094] The defined area is a detection circle with the current vehicle as the center and a set radius. Features within this detection circle can be considered scene features, and the detection circle is the defined area, in order to identify the type of surrounding vehicles, new pedestrians, and specific landmarks.

[0095] The location recognition module can use the SLAM algorithm to process LiDAR data and location information from multi-source environmental data. Centered on the current vehicle, it initially constructs a virtual driving environment. In this scenario, it continuously tracks the relative position and speed of static features with respect to the current vehicle, as well as the relative position and speed of dynamic features with respect to the current vehicle.

[0096] The environmental semantic analysis module can perform environmental semantic analysis on visual data, radar and laser data, and location information based on the initially constructed virtual driving environment.

[0097] The environmental semantic analysis module can use the established environmental semantic model to transform the sensor data contained in the current scene features into an environmental model with semantic labels, that is, to transform the sensor data into semantic soundscape mapping labels, thereby obtaining the semantic soundscape mapping labels of the current scene features.

[0098] The environmental semantic model is a pre-trained environmental semantic model. Visual data, radar and laser data, and location information of scene features are used as input features and fed into the pre-trained environmental semantic model. The pre-trained environmental semantic model outputs the semantic soundscape mapping label of the scene features.

[0099] For example, a dynamic feature is identified: a truck. The relative position of this dynamic feature (truck) to the current vehicle is identified as: 5 meters to the left, and its speed is: high speed. Using the aforementioned visual data, radar / laser data, and location information, the data is input into a trained environmental semantic model. The trained environmental semantic model outputs a semantic soundscape mapping label for this dynamic feature as: "There is a truck approaching at high speed 5 meters to the left." This semantic meaning is the prerequisite for subsequently mapping it to "a whistling sound of wind approaching from the left."

[0100] The environmental semantic model can be an environment model set up based on a large visual model or 3D object detection using Transformer / BEV, or other environmental models, and is not limited in this application.

[0101] The scene construction module can refine the initially constructed virtual driving environment based on relative position, dynamic features, static features, and semantic soundscape mapping labels, and determine the current driving scene of the vehicle, thus identifying the driving scene.

[0102] Driving scenarios include: highway cruising, urban traffic congestion, driving through tunnels, and driving in the rain. Other driving scenarios may also be included, but will not be detailed in this application.

[0103] In another embodiment of this application, the soundscape generation module includes a mapping module and an object generation module.

[0104] The mapping module can query the set sound scene rule base based on the semantic sound scene mapping tags of driving scene, dynamic features and static features, and determine multiple sound elements corresponding to driving scene, dynamic features and static features. The object generation module can use the set audio synthesis software to synthesize sound elements corresponding to dynamic features and sound elements corresponding to static features based on sound elements corresponding to driving scenarios, thereby obtaining sound objects with dynamic features and sound objects with static features.

[0105] In this embodiment, the mapping module can query the set sound scene rule base through the semantic sound scene mapping tags of dynamic features to determine the sound elements corresponding to the dynamic features.

[0106] The established soundscape rule base defines the mapping relationship between external physical objects / events and sound elements.

[0107] The identified dynamic and static features are used as external physical objects or events. The semantic soundscape mapping labels of scene features are used to query the mapping relationship of the set soundscape rule base to determine the sound element corresponding to each feature.

[0108] Based on the current driving scenario, the system queries the established soundscape rule library to determine the current background sound elements, so as to achieve a creative mapping between sound and environment, rather than a simple reproduction.

[0109] For example, when other vehicles with dynamic features are identified, they are mapped to the rustling sound of wind blowing through treetops, with the volume and rhythm of the sound proportional to the relative speed of the vehicle. When static landmarks are identified and a vehicle is approaching, they are mapped to a pre-recorded educational explanation of the landmark. Since a scene feature may have multiple sound elements, the object generation module can use the configured audio synthesis software to integrate the sound elements corresponding to the driving scene, resulting in a sound object corresponding to one scene feature, or multiple sound objects corresponding to multiple scene features. This improves the flexibility of audio playback.

[0110] The proposed audio synthesis software can be a real-time audio synthesizer, dynamically generating audio with greater flexibility.

[0111] Through a customizable soundscape rule library, creative mapping between sound and environment is achieved, rather than simple reproduction. This solves the problem of static sound effects being disconnected from the environment and transforms the external world into a canvas that can be given different narrative meanings, thus enabling a highly personalized and diverse experience.

[0112] In another embodiment of this application, the soundscape partition rendering module includes: a volume rendering module and a partition rendering module.

[0113] The volume rendering module can determine the playback position based on the relative positions of dynamic and static features in the driving scene, and then render the playback volume of the corresponding sound object using the set 3D audio rendering algorithm based on the playback position. The partition rendering module can determine the playback area corresponding to the sound object based on the playback position, and determine the in-vehicle speaker array to be invoked based on the playback area and playback volume, so as to generate the corresponding channel signal.

[0114] The volume rendering module can determine the relative position of dynamic features and the current vehicle based on the current driving scene. Using this relative position, it determines the playback position of the current dynamic feature within the virtual driving environment. The module then uses a pre-designed 3D audio rendering algorithm to calculate and render the sound object at that playback position, adjusting the playback volume accordingly.

[0115] For example, based on the current driving scenario of high-speed cruising, the relative position of the dynamic feature truck and the current vehicle is determined: 5 meters to the left. Based on this relative position, the playback position of the dynamic feature truck in the virtual driving environment is determined to achieve precise positioning. The 3D audio rendering algorithm is used to calculate and render the sound object with the sound of the car engine and the howling wind at the playback position, so that the playback volume is presented as if it is played from far to near.

[0116] Similarly, the volume rendering module can determine the relative position of static features and the current vehicle based on the current driving scene. Using this relative position, it determines the playback position of the static feature within the virtual driving environment. The module then uses the established 3D audio rendering algorithm to calculate and render the sound object at that playback position, adjusting the playback volume accordingly.

[0117] The proposed 3D audio rendering algorithm can be either high-order Ambisonics or object-based audio rendering.

[0118] The partition rendering module can determine the playback area corresponding to the sound object by the playback position, thereby determining the speaker array that the current vehicle needs to call. Based on the playback area, it determines the speaker array that the current vehicle needs to call, and generates the corresponding channel signal based on the called speaker array to play the sound object.

[0119] This solution acquires multi-source environmental data, performs environmental semantic analysis on the data, constructs a virtual driving environment, and identifies the current vehicle's driving scenario. It utilizes real-time environmental semantic extraction technology based on multi-sensor fusion, using this as the core input for soundscape generation to address the disconnect between sound and the real-time physical environment. Based on the driving scenario, it queries a pre-defined soundscape rule base to determine multiple sound elements corresponding to the scenario. These sound elements are then used to generate multiple sound objects, mapping the dynamic changes of the external environment onto the acoustic layer in real time. This provides an immersive experience that unifies the sound field and field of view, defining a rule set for the dynamic mapping relationship between physical world objects / events and virtual sound elements. The system integrates and executes audio; based on the driving scenario, it determines the playback positions of multiple sound objects in the virtual driving environment. According to the corresponding playback positions and sound objects, it uses the established 3D audio rendering algorithm to generate signals for each channel of the in-vehicle speaker array, rendering the corresponding soundscape for passengers in the corresponding seating area. Precise spatial audio rendering brings a sense of presence, while multi-zone control ensures that personalized experiences do not interfere with each other, providing a new way of presenting information. It conveys environmental information to passengers in a more natural and non-intrusive way, reduces reliance on visual screens, better meets the interaction needs of the autonomous driving era, improves the interactivity of in-vehicle audio experience, and enhances the user experience.

[0120] The following is a detailed description and explanation of the solution of this invention, using a specific scenario of a car carrying tourists traveling through the ancient city of Rome as an example: The vehicle's GPS located that it was approaching the Roman Forum, the camera identified the outline of the Senate building, and the lidar accurately measured the relative distance and angle to the building, obtaining multi-source environmental data.

[0121] Environmental semantic analysis is performed on the multi-source environmental data to construct a virtual driving environment and identify the historical landmark: the Senate. Based on the driving scenario, the established soundscape rule base is queried to determine multiple sound elements corresponding to the ancient city of Rome. Based on the sound elements of the Senate, the Senate's narration is played overlaid with the ambient sound of the ancient Roman marketplace.

[0122] The narration about the Senate is rendered using 3D audio technology, generating signals to drive the various channels of the vehicle's speaker array. This makes the sound appear to come directly from the direction of the Senate ruins, and the sound image remains stable as the vehicle moves and turns. Simultaneously, the system generates corresponding ambient sounds of an ancient Roman marketplace based on the current vehicle speed—faint conversations among people and vendors' calls—with the volume increasing as the vehicle speed decreases. When the vehicle turns onto a modern street, it can intelligently incorporate some spatiotemporally integrated sound effects.

[0123] Another embodiment of this application provides a vehicle control device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method for enhancing the real-world soundscape of a vehicle. This vehicle control device can be any smart terminal, including a tablet computer, an in-vehicle computer, or similar device.

[0124] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0125] Please see Figure 3 , Figure 3 The hardware structure of a vehicle control device according to another embodiment is illustrated. The vehicle control device includes: The processor can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to achieve the technical solutions provided in the embodiments of this application. The memory can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory and called by the processor to execute the vehicle real-world sound enhancement method of the embodiments of this application. Input / output interfaces are used to implement information input and output; The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). A bus is used to transfer information between various components of a device, such as processors, memory, input / output interfaces, and communication interfaces. The processor, memory, input / output interfaces, and communication interfaces communicate with each other within the device via a bus.

[0126] This invention also provides a vehicle, including the real-world sound enhancement method for the vehicle described above.

[0127] The vehicle can be a private car, such as a sedan, SUV, MPV, or pickup truck. It can also be a commercial vehicle, such as a van, bus, small truck, or large semi-trailer. The vehicle must have an electric motor capable of outputting power or acting as a generator to store mechanical energy. When the vehicle is a new energy vehicle, it can be a hybrid or a pure electric vehicle.

[0128] Since the vehicle applies all the technical solutions of the above-described vehicle control device, it has at least all the beneficial effects brought about by the technical solutions of the above embodiments, which will not be repeated here.

[0129] Another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for enhancing the real-world soundscape of a vehicle.

[0130] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0131] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0132] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0133] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0135] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0136] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0137] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0138] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0140] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method of reality sound scene enhancement of a vehicle, characterized in that, The method includes: Acquire multi-source environmental data, perform environmental semantic analysis on the multi-source environmental data, construct a virtual driving environment, and identify the current vehicle's driving scenario; Based on the driving scenario, the system queries the established soundscape rule library to determine multiple sound elements corresponding to the driving scenario, and generates multiple sound objects based on the multiple sound elements. Based on the driving scenario, the playback positions of multiple sound objects in the virtual driving environment are determined. According to the corresponding playback positions and corresponding sound objects, the 3D audio rendering algorithm is used to generate signals for each channel of the in-vehicle speaker array, so as to render the corresponding sound scene for passengers in the corresponding seating area.

2. The method of reality sound scene enhancement of a vehicle according to claim 1, characterized by, The step of performing environmental semantic analysis on the multi-source environmental data, constructing a virtual driving environment, and identifying the current vehicle's driving scenario includes: Based on the multi-source environmental data, dynamic and static features within the designated area are identified, and the relative positions of the dynamic and static features with the current vehicle are continuously tracked. Using the established environmental semantic model, the multi-source environmental data is analyzed to generate semantic soundscape mapping labels corresponding to the dynamic features and the static features; The virtual driving environment is constructed based on the relative position, the dynamic features, the static features, and the semantic soundscape mapping labels, and the driving scene is identified.

3. The method of reality sound scene enhancement of a vehicle according to claim 2, characterized in that, The query establishes a soundscape rule base to determine multiple sound elements corresponding to the driving scenario, and generates multiple sound objects based on these sound elements, including: Based on the driving scenario, the semantic soundscape mapping label of the dynamic feature, and the semantic soundscape mapping label of the static feature, the set soundscape rule base is queried to determine multiple sound elements corresponding to the driving scenario, the dynamic feature, and the static feature; Using the provided audio synthesis software, based on the sound elements corresponding to the driving scenario, the sound elements corresponding to the dynamic features and the sound elements corresponding to the static features are synthesized to obtain the sound object of the dynamic features and the sound object of the static features.

4. The method of reality sound scene enhancement of a vehicle according to claim 2, characterized by, The step of generating signals for each channel of the in-vehicle speaker array based on the corresponding playback position and the corresponding sound object using the established 3D audio rendering algorithm includes: Based on the relative positions of the dynamic and static features in the driving scenario, the playback position is determined according to the relative positions, and the playback volume of the corresponding sound object is rendered using the set 3D audio rendering algorithm according to the playback position. Based on the playback position, a playback area corresponding to the sound object is determined. Based on the playback area and the playback volume, the in-vehicle speaker array to be invoked is determined to generate the corresponding channel signal.

5. A realistic soundscape augmentation system for a vehicle, characterized by The system includes: The scene recognition module is used to acquire multi-source environmental data, perform environmental semantic analysis on the multi-source environmental data, construct a virtual driving environment, and identify the current vehicle's driving scene. The soundscape generation module is used to query the set soundscape rule library according to the driving scenario, determine multiple sound elements corresponding to the driving scenario, and generate multiple sound objects based on the multiple sound elements; The soundscape zoning rendering module is used to determine the playback positions of multiple sound objects in the virtual driving environment based on the driving scene, and generate the signals of each channel of the in-vehicle speaker array using the set 3D audio rendering algorithm according to the corresponding playback positions and corresponding sound objects, so as to render the corresponding soundscape for passengers in the corresponding seating area.

6. The real sound scene augmentation system of a vehicle according to claim 5, characterized in that, The scene recognition module includes: The location recognition module is used to identify dynamic and static features within a set area based on the multi-source environmental data, and to continuously track the relative positions of the dynamic and static features with the current vehicle. The environmental semantic analysis module is used to analyze the multi-source environmental data using the established environmental semantic model, and generate semantic soundscape mapping labels corresponding to the dynamic features and the static features. The scene construction module is used to construct the virtual driving environment and identify the driving scene based on the relative position, the dynamic features, the static features and the semantic soundscape mapping label.

7. The method of reality sound scene enhancement of a vehicle according to claim 2, wherein, The soundscape generation module includes: The mapping module is used to query the set sound scene rule base according to the driving scene, the semantic sound scene mapping label of the dynamic feature and the semantic sound scene mapping label of the static feature, and determine multiple sound elements corresponding to the driving scene, the dynamic feature and the static feature; The object generation module is used to synthesize, using the provided audio synthesis software, the sound elements corresponding to the dynamic features and the sound elements corresponding to the static features based on the sound elements corresponding to the driving scene, to obtain the sound object of the dynamic features and the sound object of the static features.

8. The method of reality sound scene enhancement of a vehicle according to claim 2, wherein, The soundscape partition rendering module includes: The volume rendering module is used to determine the playback position based on the relative positions of the dynamic features and the static features in the driving scene, and to render the playback volume of the corresponding sound object according to the playback position using the set 3D audio rendering algorithm. The partition rendering module is used to determine the playback area corresponding to the sound object based on the playback position, and to determine the in-vehicle speaker array to be invoked based on the playback area and the playback volume, so as to generate the corresponding channel signal.

9. A vehicle control device characterized by comprising: The system includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the method for enhancing the realistic soundscape of a vehicle as described in any one of claims 1 to 4.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. When the computer program is executed by the processor, it implements the method for enhancing the realistic soundscape of the vehicle as described in any one of claims 1 to 4.