Vehicle sound transmission method, vehicle, vehicle controller, and storage medium

By using a sound source separation model to separate the target sound signal from mixed external sounds, the problem of critical sounds not being perceived in a timely manner due to vehicle sound insulation design is solved. This enables the transmission of critical sounds and noise shielding, thereby improving driving safety and user experience.

CN122337243APending Publication Date: 2026-07-03ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-04-03
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing vehicle sound insulation designs, while blocking out external noise, prevent occupants from timely perceiving key external sound information, thus affecting driving safety.

Method used

By acquiring the occupants' voice transmission requirements, the target sound signal is separated from the mixed sound outside the vehicle using a sound source separation model and played inside the vehicle. This involves the combined use of a speech audio modality alignment module, a sound source separation module, a time-frequency conversion layer, an encoding unit, a decoding unit, and an inverse time-frequency conversion layer to accurately extract key safety sounds.

Benefits of technology

It enables occupants to promptly perceive critical safety sounds outside the vehicle, such as emergency vehicle alarms and traffic warnings, while simultaneously blocking out irrelevant noise, thereby improving driving safety and in-vehicle quietness, adapting to complex environments, and enhancing functional reliability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337243A_ABST
    Figure CN122337243A_ABST
Patent Text Reader

Abstract

This application relates to the field of vehicle control, specifically to a vehicle sound transmission method, a vehicle, a vehicle controller, and a storage medium. The vehicle includes a microphone located outside the vehicle and a speaker located inside the vehicle. The vehicle sound transmission method includes: acquiring the occupant's sound transmission request, which indicates the target external sound source the occupant expects the vehicle to transmit; acquiring mixed external sound; separating the target sound signal from the mixed external sound based on a sound source separation model and the sound transmission request, the target sound signal being the sound signal emitted by the target external sound source; and playing the target sound signal inside the vehicle. This application can shield irrelevant external noise and enable occupants to perceive key safety sounds from outside the vehicle, improving driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle control, specifically to a method for transmitting vehicle sound, a vehicle, a vehicle controller, and a storage medium. Background Technology

[0002] In the process of upgrading automobiles to be more intelligent and comfortable, the vehicle's body sealing and sound insulation systems are constantly being optimized, which can effectively isolate external road noise, mechanical noise and other interference, creating a quiet and comfortable driving environment for the occupants.

[0003] However, while this high sound insulation design isolates external noise, it may also prevent occupants from timely perceiving key external sound information, such as traffic warnings, emergency vehicle alarms, and pedestrian warnings, posing a potential risk to driving safety. Summary of the Invention

[0004] In view of the above, embodiments of this application provide a vehicle sound transmission method, a vehicle, a vehicle controller, and a storage medium, which aim to shield irrelevant external noise and enable occupants to perceive key safety sounds from outside the vehicle, thereby improving driving safety.

[0005] In a first aspect, embodiments of this application provide a vehicle sound transmission method, the vehicle sound transmission method comprising: Obtain the occupant's voice transmission requirement, wherein the voice transmission requirement indicates the target external sound source that the occupant expects the vehicle to transmit; Collect mixed sounds from outside the vehicle; Based on the sound source separation model and the sound transmission requirements, the target sound signal is separated from the mixed sound outside the vehicle. The target sound signal is the sound signal emitted by the target external sound source. Play the target sound signal inside the vehicle.

[0006] The embodiments of this application can accurately capture the passenger's transmission needs and accurately extract the target sound signal from the mixed sound outside the vehicle by combining the sound source separation model. This is beneficial for adapting to the personalized needs of passengers, enabling passengers to perceive key safety sounds from outside the vehicle in a timely manner (such as emergency vehicle alarms and traffic warning sounds), while also shielding irrelevant noise to maintain the quietness inside the vehicle, improving the transmission accuracy and adaptability to complex environments, and enhancing functional reliability and user experience.

[0007] In some embodiments, the sound source separation model includes a language audio modality alignment module and a sound source separation module. Separating the target sound signal from the mixed external sound based on the sound source separation model includes: The sound transmission requirement is converted into a text describing the target external sound source. Based on the language audio modality alignment module, the target vehicle exterior sound source description text is converted into an embedded representation of the audio modality; The embedded representation and the mixed external sound are input into the sound source separation module to obtain the target sound signal.

[0008] This application embodiment can support the separation of any text-describable external sound source (such as pedestrian shouts, ambulance sirens, non-motorized vehicle bells, etc.), no longer limited by the separation of fixed categories of sound sources.

[0009] In some embodiments, the audio source separation module includes: A time-frequency conversion layer is used to convert the mixed external sound into a frequency domain signal; An encoding unit, cascaded with the time-frequency conversion layer, comprises an alternately cascaded first fusion layer and an encoding layer; the encoding unit is used to fuse the frequency domain signal with the embedded representation through the first fusion layer to obtain a first fused feature; to encode the first fused feature through the encoding layer to obtain an encoded feature; and to determine the output feature of the encoding unit based on the encoded feature; A decoding unit, cascaded with an encoding unit, comprising alternating cascaded decoding layers and a second fusion layer; the decoding unit is used to decode the output features of the encoding unit through the decoding layers to obtain decoded features; to fuse the decoded features with the embedded representation through the second fusion layer to obtain a second fused feature; and to determine the output features of the decoding unit based on the second fused feature. A reverse time-frequency conversion layer, cascaded with the decoding unit, is used to perform reverse time-frequency conversion on the output features of the decoding unit to obtain the target sound signal.

[0010] This technical solution incorporates the embedded representation of the target sound source in both the encoding and decoding stages. This avoids the attenuation of the target external sound source information in the embedded representation, significantly improves the separation accuracy and signal-to-noise ratio, ensures the temporal continuity and frequency integrity of the target sound signal, and enhances the model's ability to adapt to different vehicle speeds and road conditions.

[0011] In some embodiments, the sound source separation model includes multiple candidate single sound source separation models; The method for separating the target sound signal from the mixed sound outside the vehicle based on the sound source separation model includes: Based on the target external sound source, a target single sound source separation model is selected from the multiple candidate single sound source separation models; The target sound signal is separated from the mixed sound outside the vehicle based on the target single sound source separation model.

[0012] By adopting this technical solution, multiple single-source separation models can be used and selected as needed, which can significantly reduce the computing load of the vehicle terminal, improve the real-time performance of source separation, and enable more thorough learning of the spectral characteristics of specific sources. It also facilitates the subsequent expansion of single-source separation models as needed to adapt to different regional traffic environments.

[0013] In some embodiments, the audio source separation model includes a multi-audio source separation model; The method for separating the target sound signal from the mixed sound outside the vehicle based on the sound source separation model includes: Based on the multi-source separation model, multiple candidate sound signals are separated from the mixed sound outside the vehicle; The target sound signal is obtained by filtering the sound signals emitted by the target external sound source from the plurality of candidate sound signals.

[0014] In some embodiments, the vehicle includes multiple audio playback devices, and playing the target sound signal inside the vehicle includes: The relative position of the target external sound source and the vehicle is determined based on the target sound signal; The volume distribution ratio of each broadcasting device is determined based on the relative position and the position of the multiple broadcasting devices inside the vehicle; The target sound signal is played by controlling the multiple audio devices according to the volume distribution ratio.

[0015] This technical solution maps the layout of the audio equipment and allocates the volume ratio based on the relative position of the target sound source and the vehicle, enabling passengers to intuitively perceive the location of the sound source outside the vehicle through hearing, thereby improving driving safety. At the same time, it conforms to the spatial hearing characteristics of the human ear, optimizes the auditory experience, and reduces fatigue.

[0016] In some embodiments, obtaining the occupant's voice pass-through requirement includes: In response to the occupant's configuration request for the target external audio source, candidate configuration modes are displayed on the display interface. The candidate configuration modes include at least one of the following: language command configuration mode, audio source type configuration mode, and scene configuration mode. In response to the occupant's selection of a candidate configuration mode, the target configuration mode selected by the occupant is determined; When the target configuration mode includes the language instruction configuration mode, listen to the occupant's language instructions and determine the sound transmission requirement based on the occupant's language instructions; When the target configuration mode includes the sound source type configuration mode, the types of external sound sources are displayed on the display interface, and the sound transmission requirement is determined based on the types of external sound sources selected by the occupant. When the target configuration mode includes a scenario configuration mode, the sound transmission requirement is determined based on the vehicle's driving scenario.

[0017] This technical solution provides three modes for acquiring sound transmission requirements: language commands, sound source type selection, and scene configuration, to meet the occupants' flexible configuration needs for target external sound sources.

[0018] Secondly, embodiments of this application also provide a vehicle, wherein the vehicle sound transmission device includes: The audio transmission requirement acquisition module is used to acquire the occupant's audio transmission requirement, which indicates the target external sound source that the occupant expects the vehicle to transmit. The signal acquisition module is used to acquire mixed sounds from outside the vehicle through the microphone; The vehicle exterior sound separation module is used to separate the target sound signal from the mixed sound outside the vehicle based on the sound source separation model and the sound transmission requirements. The target sound signal is the sound signal emitted by the target external sound source. The playback control module is used to play the target sound signal inside the vehicle.

[0019] Thirdly, embodiments of this application also provide a vehicle controller, the vehicle controller including a processor and a memory, the memory being used to store instructions, and the processor being used to call the instructions in the memory, causing the vehicle controller to execute the vehicle sound transmission method as described in the first aspect.

[0020] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer instructions that, when executed on a vehicle controller, cause the vehicle controller to perform the vehicle sound transmission method as described in the first aspect. Attached Figure Description

[0021] Figure 1 This is a flowchart of the steps of a vehicle sound transmission method according to an embodiment of this application.

[0022] Figure 2 This is a schematic diagram of a scenario for audio separation based on a general configurable audio separation model provided in an embodiment of this application.

[0023] Figure 3 This is a schematic diagram of the model structure of an audio source separation module provided according to an embodiment of this application.

[0024] Figure 4 This is a schematic block diagram illustrating the training of a general configurable audio separation model according to an embodiment of this application.

[0025] Figure 5This is a schematic diagram of an audio separation scenario based on a candidate single-source separation model provided in an embodiment of this application.

[0026] Figure 6 This is a schematic block diagram illustrating the training of a candidate single-source separation model according to an embodiment of this application.

[0027] Figure 7 This is a schematic diagram of an audio separation scenario based on a multi-source separation model provided in an embodiment of this application.

[0028] Figure 8 This is a schematic block diagram illustrating the training of a multi-source separation model according to an embodiment of this application.

[0029] Figure 9 This is a schematic diagram of a vehicle sound transmission method according to an embodiment of this application.

[0030] Figure 10 This is a structural block diagram of a vehicle provided according to an embodiment of this application.

[0031] Figure 11 This is a schematic diagram of the structure of a vehicle controller provided according to an embodiment of this application. Detailed Implementation

[0032] To better understand the above-mentioned objectives, features, and advantages of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0033] The following description sets forth many specific details to provide a full understanding of this application. The described embodiments are only some, not all, of the embodiments of this application.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0035] It should be further noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0036] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.

[0037] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0038] This application provides a vehicle sound transmission method, a vehicle, a vehicle controller, and a computer-readable storage medium.

[0039] The vehicle includes an external sound-collecting device for capturing sounds from outside the vehicle. This device can be a microphone, a pickup, or similar equipment, but is not limited to these.

[0040] Specifically, multiple microphones can be installed on the outer side of the vehicle body, covering key areas around the vehicle to form a microphone array, ensuring all-around pickup of external sounds.

[0041] For example, a radio array may include a front radio, a left radio, a right radio, and a rear radio.

[0042] The front microphone can be positioned on the front bumper, grille, or edge of the windshield. It is used to pick up ambient sounds directly in front of the vehicle, such as vehicle horns, traffic lights, and pedestrian shouts.

[0043] The left-side microphone can be installed near the left-side rearview mirror or on the door frame to collect sounds from the left side of the vehicle, such as horns from vehicles approaching from the left or warning sounds from non-motorized vehicles.

[0044] The installation position of the right-side radio can be symmetrical to that of the left-side radio, and it can be placed on the right-side rearview mirror or door frame to pick up ambient sounds from the right side of the vehicle.

[0045] The rear sound receiver can be installed on the rear bumper, the edge of the trunk, or the rear windshield area to capture sounds from behind the vehicle, such as warning sounds from vehicles behind and reversing alert sounds.

[0046] The aforementioned four-way sound pickup layout of "front-left-right-rear" enables 360° ambient sound pickup without blind spots, providing multi-dimensional acoustic data support for subsequent sound source separation and localization.

[0047] The layout of the above-described radio array is merely an example. In practical applications, the number of radio receivers can be flexibly set according to the cost and positioning accuracy requirements. This application does not limit this. For example, with a higher cost budget, the number of radio receivers can be increased to achieve more accurate sound source positioning.

[0048] The vehicle may also include an in-vehicle audio system for playing sound into the vehicle. The audio system may be a loudspeaker, but is not limited to this. The number and location of the audio system can also be configured according to actual application requirements, and this application embodiment does not limit this.

[0049] Both the aforementioned radio and broadcast devices can communicate with the vehicle controller. The vehicle controller can be used to execute a vehicle sound transmission method, which may include: the vehicle controller acquiring the occupant's sound transmission request, the sound transmission request indicating the occupant's desired target external sound source for transmission; acquiring mixed external sound through the radio, then separating the target sound signal from the mixed external sound based on a sound source separation model and the sound transmission request, the target sound signal being the sound signal emitted by the target external sound source; and playing the target sound signal inside the vehicle through the broadcast device.

[0050] The embodiments of this application can accurately capture the passenger's transmission needs and accurately extract the target sound signal from the mixed sound outside the vehicle by combining the sound source separation model. This is beneficial for adapting to the personalized needs of passengers, enabling passengers to perceive key safety sounds from outside the vehicle in a timely manner (such as emergency vehicle alarms and traffic warning sounds), while also shielding irrelevant noise to maintain the quietness inside the vehicle, improving the transmission accuracy and adaptability to complex environments, and enhancing functional reliability and user experience.

[0051] This vehicle controller is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, processors, microprogrammed control units (MCUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. This vehicle controller can be an in-vehicle intelligent cockpit domain controller, an in-vehicle edge computing unit, etc., but is not limited to these.

[0052] refer to Figure 1 As shown, Figure 1 This is a flowchart illustrating the steps of a vehicle sound transmission method according to an embodiment of this application. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements. The vehicle sound transmission method may include the following steps: Step 101: Obtain the occupant's voice transmission requirement. The voice transmission requirement indicates the target external sound source that the occupant expects the vehicle to transmit.

[0053] Passengers refer to users such as drivers or passengers in a vehicle.

[0054] Sound transmission refers to the process of transmitting external sounds into the vehicle.

[0055] External sound sources refer to various sound sources generated in the external environment of a vehicle. For example, external sound sources may include sound sources corresponding to the sirens of emergency vehicles (such as ambulances, fire trucks, and police cars), traffic signal prompts, pedestrian sounds, non-motorized vehicle sounds, construction machinery sounds, etc., but are not limited to these.

[0056] The vehicle controller can obtain the occupant's voice transmission request from the human-machine interface. This human-machine interface can be a voice interface or a user interface. The human-machine interface can be provided by the vehicle controller itself or by other interactive devices. For example, the interactive device can communicate with the vehicle controller to transmit the voice transmission request to the vehicle controller; this application embodiment does not limit this.

[0057] For example, an occupant triggers a configuration request for a target external audio source based on this human-machine interface, such as by clicking the configuration program icon for the target external audio source on the in-vehicle display screen.

[0058] The vehicle controller can respond to the occupant's configuration request for a target external audio source and display candidate configuration modes on the display interface.

[0059] Candidate configuration modes may include one or more of the following, such as language command configuration mode, sound source type configuration mode, and scene configuration mode, but are not limited to these: For example, candidate configuration modes may also include pass-through volume configuration mode, which allows occupants to adjust the pass-through volume independently, avoiding conflicts with the volume adjustments of other sounds (such as music) in the vehicle.

[0060] Passengers can select at least one candidate configuration mode as the target configuration mode from the candidate configuration modes provided in the human-machine interface.

[0061] The vehicle controller can determine the target configuration mode selected by the occupant in response to the occupant's selection of a candidate configuration mode.

[0062] When the target configuration mode includes the language command configuration mode, the vehicle controller listens to the occupant's language commands and determines the sound transmission requirement based on the occupant's language commands.

[0063] The voice command mode allows passengers to freely describe the sounds they need to transmit. For example, if a passenger issues the voice command "the sound of wind and birdsong in the forest" while driving, the vehicle controller can use the sound of wind and birdsong in the forest as the target external sound source.

[0064] When the target configuration mode includes the sound source type configuration mode, the types of external sound sources are displayed on the display interface, and the sound transmission requirement is determined based on the types of external sound sources selected by the occupant.

[0065] The audio source type configuration mode allows passengers to select the audio source to be transmitted from a number of preset types of external audio sources (such as emergency vehicle alarms, human voices, traffic prompts, etc.) and use it as the target external audio source.

[0066] When the target configuration mode includes a scenario configuration mode, the sound transmission requirement is determined based on the vehicle's driving scenario.

[0067] The driving scenarios may include, but are not limited to, urban commuting, highway driving, and parking modes.

[0068] The vehicle controller can pre-store the mapping relationship between driving scenarios and target external sound sources. This mapping relationship can be configured by the occupants or the default mapping relationship of the vehicle controller can be used.

[0069] When the target configuration mode includes a scenario configuration mode, the vehicle controller can determine the vehicle's driving scenario and find the sound transmission requirements that match the driving scenario based on the mapping relationship.

[0070] In some embodiments, the step of the vehicle controller determining the driving scenario of the vehicle may include: providing a human-machine interface for configuring the driving scenario; and obtaining the driving scenario of the vehicle from the human-machine interface in response to a occupant's driving scenario configuration operation.

[0071] In other embodiments, the step of the vehicle controller determining the driving scenario of the vehicle may include: acquiring sensor data from on-board sensors and vehicle control signals; and determining the driving scenario of the vehicle based on the vehicle control signals and sensor data.

[0072] The sensor data can be images captured by an image sensor, vehicle speed captured by a vehicle speed sensor, etc., but is not limited to these. The vehicle control signals can be parking signals, stationary signals, etc., but are not limited to these.

[0073] The above-described method for determining vehicle driving scenarios is merely an example. In actual applications, the method for determining vehicle driving scenarios can be set according to requirements, and this application embodiment does not limit this.

[0074] Step 102: Collect mixed sounds from outside the vehicle.

[0075] For example, a radio can collect mixed sounds from outside the vehicle, convert the mixed sounds into a signal (such as an electrical signal) that the vehicle controller can interpret, and transmit the signal to the vehicle controller so that the vehicle controller can collect the mixed sounds from outside the vehicle through the radio.

[0076] Step 103: Based on the sound source separation model and sound transmission requirements, separate the target sound signal from the mixed sound outside the vehicle.

[0077] The target sound signal is the sound signal emitted by the external sound source of the target vehicle.

[0078] The sound source separation model is used to separate the sound signals belonging to the target external sound source from the mixed sound outside the vehicle. For example, if the target external sound sources include sound source A, sound source B, and sound source C, the sound source separation model can separate the sound signals belonging to sound source A, sound source B, and sound source C from the mixed sound outside the vehicle.

[0079] The sound source separation model can be built upon a basic artificial intelligence model based on deep learning. This basic artificial intelligence model can be a convolutional neural network (CNN), a gated recurrent unit (GRU), or a Transformer, but is not limited to these.

[0080] For example, training equipment can train and optimize a basic artificial intelligence model using a massive amount of external sound samples, enabling the basic artificial intelligence model to have real-time classification and noise reduction capabilities, thereby allowing the sound source separation model to accurately separate the target sound signal with low noise.

[0081] The external sound sample can include external sounds from different scenarios and those with different levels of interference noise.

[0082] The sound source separation model can perform Fourier transform on the mixed sound outside the vehicle to extract frequency and time domain features; then analyze the frequency and time domain features to accurately distinguish the target sound signal from the interference noise and generate a noise mask; finally, filter the mixed sound outside the vehicle through the mask to output a clean target sound signal to ensure the integrity and clarity of the target sound signal.

[0083] Compared to traditional signal processing-based sound source separation methods, the embodiments of this application can effectively separate the target sound signal from interference noise in complex multi-sound source aliasing scenarios, significantly improving the accuracy and reliability of sound source separation.

[0084] In some embodiments, the audio source separation model may include at least one of a general configurable audio separation model, a candidate single audio source separation model, and a multi-audio source separation model.

[0085] The following sections introduce the model structure, inference steps, and training steps of the general configurable audio separation model, the candidate single-source audio separation model, and the multi-source audio separation model.

[0086] 1. General configurable audio separation model: 1.1 Model Structure and Reasoning Steps: A general configurable audio separation model is used to separate a target sound signal from mixed external sound based on a target external sound source description text. This target external sound source description text describes the target external sound source that the occupant expects the vehicle to transmit. The target external sound source description text includes at least the type of the target external sound source, and may also include descriptive text of the sound characteristics of the target external sound source, such as amplitude, frequency, and timbre; however, this application embodiment does not limit this.

[0087] For example, an occupant can issue a voice command carrying the occupant's voice transmission requirements; the vehicle controller can convert the voice command into text to obtain a description text of the target external sound source; then, a general configurable audio separation model can separate the target sound signal from the mixed external sound based on the description text of the target external sound source.

[0088] refer to Figure 2 As shown, the general configurable audio separation model may include a language audio modality alignment module and a sound source separation module.

[0089] The steps by which the vehicle controller separates the target sound signal from the mixed sound outside the vehicle based on a general configurable audio separation model and sound pass-through requirements may include the following steps 2a, 2b to 2c: Step 2a: The vehicle controller converts the sound transmission requirement into a text description of the target external sound source.

[0090] For example, the vehicle controller can convert the occupant's voice commands, the type of target external sound source selected by the occupant, or the type of target external sound source mapped by the vehicle's driving scenario into text description information to obtain the target external sound source description text.

[0091] Step 2b: The vehicle controller converts the target external sound source description text into an audio embedding based on the language audio modality alignment module.

[0092] The language audio modality alignment module can employ a multimodal semantic pre-trained model.

[0093] The multimodal semantic pre-training model can be an Audio Language Model (ALM), a Contrastive Language-Audio Pre-training (CLAP) model, or an Omni-Modal Large Language Model (Omni-MLLM), such as the QwenOmni model, but is not limited to these.

[0094] The aforementioned multimodal semantic pre-trained models can all be trained on large-scale language audio pairing databases, including AudioSet, WavCaps, and Auto-ACD datasets, but are not limited to these.

[0095] Different multimodal semantic pre-trained models have different pre-training tasks. For example, the audio language model ALM uses a mask prediction task similar to that of the large language model LLM to complete pre-training.

[0096] The contrastive language-audio pre-trained model is trained using audio-language contrastive learning, while the Tongyi Thousand Questions full-modal model QwenOmni is pre-trained through multi-task parallel training, including Automatic Speech Recognition (ASR) and Audio Caption.

[0097] These models can all serve as language audio modality alignment modules, transforming target external sound source description text, such as a user's language description, into an embedded representation of the audio modality.

[0098] Step 2c: The vehicle controller inputs the embedded representation and the mixed external sound into the sound source separation module to obtain the target sound signal.

[0099] For example, the audio source separation module may adopt an autoregressive audio separation model, such as a Residual Network (ResNet) or a Deep Multi-scale U-net Convolutional Separation (DMUCS), or it may adopt a generative model, such as a Diffusion Model (DM). This application embodiment does not limit this.

[0100] This application embodiment can add a fusion layer to the above-described autoregressive audio separation model to fuse the embedded representation and the mixed external sound, thereby obtaining the model structure of the sound source separation module. For example, this sound source separation module can fuse the embedded representation and the mixed external sound to obtain fused features; and separate the target sound signal based on these fused features.

[0101] The following describes the model structure of the audio source separation module. In actual application, the model structure adopted by the audio source separation module can be determined according to the in-vehicle computing power configuration and effect requirements.

[0102] refer to Figure 3 As shown, Figure 3 This is a schematic diagram of the model structure of the audio source separation module provided in one embodiment of this application.

[0103] The audio source separation module includes: a time-frequency conversion layer, an encoding unit, a decoding unit, and an inverse time-frequency conversion layer cascaded in sequence. The encoding unit includes an alternately cascaded first fusion layer and an encoder layer, and the decoding unit includes an alternately cascaded decoder layer and a second fusion layer.

[0104] The time-frequency conversion layer is used to convert the mixed external sound into a frequency domain signal.

[0105] The coding unit is configured to fuse the frequency domain signal with the embedded representation through the first fusion layer to obtain a first fused feature; encode the first fused feature through the coding layer to obtain a coded feature; and determine the output feature of the coding unit based on the coded feature.

[0106] In some embodiments, determining the output feature of the coding unit based on the coding feature includes: when the coding unit includes a first fusion layer and a decoding layer, the coding feature is the output feature of the coding unit; when the coding unit includes multiple alternately cascaded first fusion layers and coding layers; the coding unit can perform feature fusion of the coding feature with the embedded representation through the next first fusion layer cascaded with the coding layer to obtain the fused feature output by the next first fusion layer, and continue to encode the fused feature through the next coding layer cascaded with the next first fusion layer until the last network layer of the coding unit is reached to obtain the output feature of the coding unit.

[0107] The decoding unit is used to decode the output features of the encoding unit through the decoding layer to obtain decoded features; to fuse the decoded features with the embedded representation through the second fusion layer to obtain second fused features; and to determine the output features of the decoding unit based on the second fused features.

[0108] In some embodiments, determining the output feature of the decoding unit based on the second fusion feature may include: when the decoding unit includes a second fusion layer and a decoding layer, the second fusion feature may be the output feature of the decoding unit; when the decoding unit includes multiple alternately cascaded second fusion layers and decoding layers; the decoding unit may decode the second fusion feature through the next decoding layer cascaded with the second fusion layer to obtain the decoding feature output by the next decoding layer, and perform feature fusion of the decoding feature and the embedded representation through the next second fusion layer cascaded with the next decoding layer, until the last network layer of the decoding unit is reached to obtain the output feature of the decoding unit.

[0109] The inverse time-frequency conversion layer is used to perform inverse time-frequency conversion on the output features of the decoding unit to obtain the target sound signal.

[0110] The encoding and decoding layers can adopt the ResNet block structure, while the first and second fusion layers can adopt feature-wise linear modulation (FILM), or cross attention mechanism (Cross-Attn) and other model structures.

[0111] The number of the first fusion layer, encoding layer, decoding layer and second fusion layer can be configured according to actual application requirements, and the embodiments of this application do not limit this.

[0112] Continue to refer to Figure 3As shown, for example, the time-frequency conversion layer can employ Short-Time Fourier Transform (STFT) technology. The time-frequency conversion layer is cascaded with the coding unit.

[0113] The coding unit may include three first fusion layers (denoted as fusion layer1, fusion layer2 and fusion layer3 respectively) and three coding layers (denoted as encoder layer1, encoder layer2 and encoder layer3 respectively), and the three first fusion layers and the three coding layers are alternately cascaded to form the coding unit.

[0114] Encoding units can be cascaded with decoding units.

[0115] The decoding unit may include three decoding layers (denoted as decoder layer1, decoder layer2 and decoder layer3 respectively) and two second fusion layers (denoted as fusion layer4 and fusion layer5 respectively). The three decoding layers and the two second fusion layers are alternately cascaded to form the decoding unit.

[0116] The decoding unit can be cascaded with the inverse time-frequency conversion layer.

[0117] The inverse time-frequency conversion layer can employ the Inverse Short-Time Fourier Transform (ISTFT) technique.

[0118] The following is combined Figure 3 Explain the role of each network layer.

[0119] The time-frequency conversion layer is used to convert the mixed sound outside the vehicle into a frequency domain signal. For example, the mixed sound outside the vehicle is subjected to a short-time Fourier transform to obtain a frequency domain signal, and the frequency domain signal is input to the next level first fusion layer cascaded with the time-frequency conversion layer.

[0120] The first fusion layer is used to fuse the output features of the previous network layer with the embedded representation to obtain a first fused feature, and to input the first fused feature into the next level coding layer concatenated with the first fusion layer.

[0121] In the case where the current first fusion layer is the first fusion layer of the coding unit, such as the current first fusion layer being the first fusion layer 1 (denoted as fusion layer 1), the output characteristics of the above-mentioned upper-level network layer are the frequency domain signals output by the time-frequency conversion layer.

[0122] When the current first fusion layer is not the first fusion layer of the coding unit, such as the current first fusion layer being the first fusion layer 2 (denoted as fusion layer 2) or the first fusion layer 3 (fusion layer 3), the output features of the above-mentioned previous network layer are the coding features output by the previous coding layer concatenated with the current first fusion layer.

[0123] The encoding layer is used to encode the first fused feature to obtain the encoded feature, and to input the encoded feature into the next level network layer.

[0124] In the case where the current coding layer is not the last coding layer of the coding unit, such as the current coding layer being the first coding layer 1 (denoted as encoder layer 1) or the first coding layer 2 (denoted as encoder layer 2), the next level network layer is the next first fusion layer concatenated with the current coding layer.

[0125] When the current coding layer is the last coding layer of a coding unit, for example, when the current coding layer is the first coding layer 3 (denoted as encoder layer 3), the next level network layer is the next decoding layer concatenated with the current coding layer, such as decoder layer 1.

[0126] The decoding layer is used to decode the output features of the previous network layer to obtain decoded features, and to input the decoded features into the next network layer.

[0127] For example, the output features of the network layer above the decoding layer 1 are the output features of the coding layer 3, and the next network layer is the second fusion layer 4; the output features of the network layer above the decoding layer 2 are the output features of the second fusion layer 4, and the next network layer is the second fusion layer 5; the output features of the network layer above the decoding layer 3 are the output features of the second fusion layer 5, and the next network layer is the inverse time-frequency conversion layer.

[0128] The second fusion layer is used to fuse the output features of the previous decoding layer with the embedded representation to obtain a second fused feature, and to input the second fused feature into the next decoding layer cascaded with the second fusion layer.

[0129] The inverse time-frequency conversion layer is used to perform inverse time-frequency conversion on the output features of the previous network layer (such as decoder layer 3) to obtain the target sound signal. For example, a short-time Fourier transform is performed on the mixed sound outside the vehicle to obtain a frequency domain signal.

[0130] The embodiments of this application employ a general configurable audio separation model that can support the separation of various types of target external sound sources. In addition, it can also support language interaction with occupants. For example, if a user inputs the language "ambulance siren and truck horn", the general configurable audio separation model can separate the target sound signal that matches the language description.

[0131] The inference process of the general configurable audio separation model has been introduced above. The training process of the configurable audio separation model is described below.

[0132] refer to Figure 4 As shown, the training steps for a general configurable audio separation model may include: Before training, a general configurable audio separation model architecture is built, integrating the language audio modality alignment module and the sound source separation module to ensure that the model has the ability to separate sound sources guided by text description.

[0133] The target audio source types to be separated are determined, and a paired database labeled with text descriptions and corresponding audio data is selected as the training set.

[0134] The audio source mixer randomly selects multiple target audio sources and mixes them according to a random volume ratio that ensures each audio source is within a reasonable signal-to-noise ratio range (e.g., -10dB to 20dB), generating mixed audio samples. At the same time, it generates text description information samples, which are description texts of the audio source types to be separated.

[0135] The training device inputs the mixed audio samples and text description information samples into a general configurable audio separation model to obtain the target separated audio source signals (such as the separated audio A) corresponding to the text description information samples.

[0136] Then, the target separated audio source signals output by the general configurable audio separation model are compared with the original audio source signals before mixing, and the error loss value between the two is calculated. Based on the loss value, the parameters of the language audio modality alignment module and the audio source separation module are updated synchronously through the backpropagation mechanism. After multiple rounds of iterative training until the loss value converges to the preset threshold, the training of the general configurable audio separation model is completed.

[0137] 2. Candidate Single Sound Source Separation Model: Candidate single-source sound separation models are used to separate the sound signal of a single source from mixed sounds outside the vehicle. For example, candidate single-source sound separation model 1 can separate the alarm sound from mixed sounds outside the vehicle; candidate single-source sound separation model 2 can separate the wind sound from mixed sounds outside the vehicle.

[0138] In some embodiments, the sound source separation model may include multiple candidate single sound source separation models; step 103 may include: selecting a target single sound source separation model from the multiple candidate single sound source separation models based on the target external sound source; and separating the target sound signal from the mixed external sound based on the target single sound source separation model.

[0139] For example, refer to Figure 5 As shown, the vehicle controller can determine the target external sound source based on the type of external sound source selected by the occupant in the sound source type configuration mode, or the sound transmission requirement matching the driving scenario in the scenario mode. Then, the model selector selects the target single sound source separation model (such as separation model A and separation model B) from multiple candidate single sound source separation models. Then, each target single sound source separation model separates the signal collected by the sound receiving device (i.e., the mixed external sound source) to obtain the sound source to be transmitted (i.e., the target sound signal).

[0140] The above describes the reasoning process of the candidate single-source separation model. The training process of the candidate single-source separation model will be introduced below.

[0141] refer to Figure 6 As shown, Figure 6 This is a schematic block diagram illustrating the training process of a candidate single-source separation model provided in an embodiment of this application.

[0142] The training dataset is an audio dataset labeled with different sound source types. For example, open-source audio datasets such as Audioset and VGGsound can be used.

[0143] Before model training, the types of sound source signals to be separated need to be determined, such as alarm sounds, wind sounds, and human voices. Then, for each determined sound source type, a corresponding candidate single-source sound separation model (such as sound source separation model A, sound source separation model B, sound source separation model C, etc.) is constructed and trained separately. The model structure of the candidate single-source sound separation model can adopt mature sound source separation model architectures such as ResNet and Demucs.

[0144] During training, the audio mixer randomly selects multiple audio sources and mixes them according to random volume ratios. The volume ratios of the multiple audio sources must ensure that each source signal is within a reasonable signal-to-noise ratio range, such as -0dB to 20dB.

[0145] During training, each candidate single-source sound separation model outputs its corresponding target separated sound source signal (i.e., separated sound source A, separated sound source B, separated sound source C, etc.). The separated sound source signal is compared with the original sound source signal before mixing, and the difference between the two is calculated by a loss function. The loss functions that can be selected include time-domain L1 loss, frequency-domain mean square error loss, multi-resolution short-time Fourier transform loss, etc., but are not limited to these.

[0146] After multiple rounds of iterative training and model parameter updates, several candidate single-source separation models with single-source output capability can be obtained.

[0147] The embodiments of this application adopt a candidate single audio source separation model, which has outstanding audio source separation effect for a small number of predetermined audio sources and is suitable for audio source separation under preset scene modes (such as "city commuting mode" and "high-speed driving mode").

[0148] 3. Multi-source separation model: Multi-source sound separation models are used to separate sound signals from multiple sources in a mixed sound environment outside the vehicle. For example, a multi-source sound separation model can separate alarm sounds, wind sounds, and human voices from a mixed sound environment outside the vehicle.

[0149] When the sound source separation model includes a multi-sound source separation model, step 103 may include: separating multiple candidate sound signals from the mixed sound outside the vehicle based on the multi-sound source separation model; filtering the sound signals emitted by the target external sound source from the multiple candidate sound signals to obtain the target sound signal.

[0150] For example, refer to Figure 7 As shown, the mixed sound outside the vehicle collected by the sound receiving device is processed by a multi-source separation model to obtain a fixed number of candidate sound signals after separation. Then, the sound signals emitted by the target external sound source are filtered according to the type of external sound source selected by the occupant in the sound source type configuration mode, or the sound transmission requirements that match the driving scenario in the scene mode, to obtain the target sound signal.

[0151] The above describes the reasoning process of the multi-source separation model. The training process of the multi-source separation model will be introduced below.

[0152] refer to Figure 8 As shown, Figure 8 This is a schematic block diagram illustrating the training process of a multi-source separation model provided in an embodiment of this application.

[0153] Before model training, determine the types of sound sources to be separated (such as alarm sounds, wind sounds, human voices, etc.), and select open-source audio datasets (such as Audioset, VGGsound) labeled with the corresponding sound source types as training data. The model structure can adopt mature sound source separation model architectures such as ResNet and Demucs, and the output layer of the model is adapted to the required dimensions based on the need for simultaneous output of multiple sound sources.

[0154] The audio source mixer randomly selects multiple target audio sources and mixes them according to a random volume ratio. The volume ratio must ensure that the signals of each audio source are within a reasonable signal-to-noise ratio range, generating a mixed audio sample containing multiple audio sources.

[0155] Then, the mixed audio samples are input into the multi-source separation model. The model can output the separation signals of all preset target audio sources in a single forward propagation. Each separation signal is compared with the corresponding original audio source signal before mixing, and the separation error is calculated by a preset loss function (such as time domain L1 Loss, frequency domain MSE Loss, Multi-Resolution STFT Loss).

[0156] Based on the error calculated by the loss function, the model parameters of the multi-source separation model are updated through backpropagation. After multiple rounds of iterative training, the error between the separated source signals and the original source signals output by the model converges to a preset threshold, thus completing the training of the multi-source separation model.

[0157] It is understandable that, because the task of separating multiple sound sources needs to be completed simultaneously, the number of parameters and the amount of computation of the multi-output model are higher than those of the single-output model. For example, the number of parameters and the amount of computation of the multi-source separation model that can simultaneously separate four types of sound sources usually need to be about twice that of the single-output model in order to achieve a separation effect comparable to that of the multi-source separation model.

[0158] Candidate single-source separation models and multi-source separation models can complete the separation of several pre-designed source categories. The general configurable audio separation model can support the separation capability of various types of sources. In practical applications, different source separation models can be configured in vehicles according to the vehicle's computing power, cost, and passenger needs.

[0159] It should be noted that the model structure, inference steps, and training steps of the above-mentioned general configurable audio separation model, candidate single-source separation model, and multi-source separation model are merely examples, intended to more clearly illustrate the technical solutions of the embodiments of this application. Those skilled in the art will understand that as the model structure evolves, the above-mentioned inference steps and training steps can also evolve accordingly.

[0160] Step 104: Play the target sound signal inside the vehicle.

[0161] In some embodiments, step 104 may include: determining the relative position of the target external sound source and the vehicle based on the target sound signal; determining the volume distribution ratio of each of the plurality of broadcasting devices based on the relative position and the position of the plurality of broadcasting devices inside the vehicle; and controlling the plurality of broadcasting devices to play the target sound signal according to the volume distribution ratio.

[0162] For example, refer to Figure 9 As shown, for each target sound signal obtained after separation, combined with the sound pickup data from the multi-receiver array, the spatial position of each target external sound source in the external environment is calculated using the Direction of Arrival (DOA) algorithm (e.g., "30° to the right front, 5 meters away"). This positioning information can serve as the spatial coordinate reference for subsequent 3D sound rendering, providing data support for accurately reconstructing the sound source position.

[0163] The 3D sound rendering engine is the core functional module for achieving spatial sound transmission. Through distance-azimuth vector decomposition and volume mapping technologies, it accurately renders the separated sound sources in the 3D space inside the vehicle, thereby restoring the actual spatial location and auditory perception effect of the external sound sources. The specific implementation process is as follows: Spatial rendering and audio device allocation employ the distance-azimuth vector decomposition method, which converts the external spatial position parameters of the sound source into driving weight parameters (i.e., volume allocation ratios) for each audio device inside the vehicle. The distance-azimuth vector decomposition method can employ a vector-based amplitude panning (VBAP) method to spatialize the target sound signal and obtain the volume allocation ratios; however, this application embodiment does not limit this approach.

[0164] For example, the vehicle controller calculates the output weights of in-vehicle audio devices in different locations, such as the front sound field, rear sound field, and side sound field, based on the relative distance and orientation angle between the sound source and the vehicle, so as to achieve a spatial positioning effect that prioritizes the enhancement of the playback of the external sound source and the corresponding audio device inside the vehicle.

[0165] Taking the sound source on the right front of the vehicle as an example, the system will enhance the output power of the right front speaker and the right speaker in the center console, while controlling the left speaker to maintain a weak output state. Through the coordinated sound output of multiple speakers, a spatial orientation auditory effect is created.

[0166] To ensure auditory comfort and output consistency of sound transmitted within the vehicle, the system introduces a mapping standard between external sound pressure level and in-vehicle playback volume. Referencing industry standards such as the "Automotive Acoustic Comfort Design Specification" and "Performance Requirements for In-Vehicle Audio Systems," the system accurately maps the sound pressure level of external sound sources (such as a 60dB horn sound from outside the vehicle) to a standard in-vehicle volume range (such as 45-55dB) that meets the requirements of driving safety and auditory comfort.

[0167] When the sound pressure level of an external sound source changes with distance (such as a distant sound source gradually approaching the vehicle), the system can dynamically adjust the output volume of each in-vehicle audio device to ensure that the volume at the user's hearing location remains stable within the standard range, avoiding the discomfort of fluctuating sound levels. Simultaneously, users can manually adjust the overall sound transmission volume according to their needs.

[0168] The embodiments of this application can accurately capture the passenger's transmission needs and accurately extract the target sound signal from the mixed sound outside the vehicle by combining the sound source separation model. This is beneficial for adapting to the personalized needs of passengers, enabling passengers to perceive key safety sounds from outside the vehicle in a timely manner (such as emergency vehicle alarms and traffic warning sounds), while also shielding irrelevant noise to maintain the quietness inside the vehicle, improving the transmission accuracy and adaptability to complex environments, and enhancing functional reliability and user experience.

[0169] Based on the same idea as the vehicle sound transmission method in the above embodiments, this application also provides a vehicle including a sound transmission device, which can be used to execute the above-described vehicle sound transmission method. For ease of explanation, the structural diagrams of the vehicle embodiments only show the parts related to the embodiments of this application. Those skilled in the art will understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0170] like Figure 10 As shown, the vehicle sound transmission device includes a transmission demand acquisition module 1001, a signal acquisition module 1002, an external sound separation module 1003, and a playback control module 1004. In some embodiments, the above modules can be programmable software instructions stored in a memory and executable by a processor. It is understood that in other embodiments, the above modules can also be program instructions or firmware embedded in a processor.

[0171] The transmission requirement acquisition module 1001 is used to acquire the occupant's voice transmission requirement, wherein the voice transmission requirement indicates the target external sound source that the occupant expects the vehicle to transmit. Signal acquisition module 1002 is used to acquire mixed sounds from outside the vehicle; The vehicle exterior sound separation module 1003 is used to separate the target sound signal from the mixed sound outside the vehicle based on the sound source separation model and the sound transmission requirements. The target sound signal is the sound signal emitted by the target external sound source. The playback control module 1004 is used to play the target sound signal inside the vehicle.

[0172] This application also provides a vehicle that includes a vehicle sound transmission device.

[0173] Figure 11 This is a schematic diagram of an embodiment of the vehicle controller of this application.

[0174] The vehicle controller 100 includes a memory 20, a processor 30, and a computer program 40 stored in the memory 20 and executable on the processor 30. When the processor 30 executes the computer program 40, it implements the steps described in the above-described vehicle sound transmission method embodiment, for example... Figure 1 The method for transmitting vehicle sound is shown.

[0175] For example, the computer program 40 can also be divided into one or more modules / units, which are stored in the memory 20 and executed by the processor 30. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 40 in the vehicle controller 100.

[0176] Those skilled in the art will understand that the schematic diagram is merely an example of the vehicle controller 100 and does not constitute a limitation on the vehicle controller 100. It may include more or fewer components than shown, or combine certain components, or different components. For example, the vehicle controller 100 may also include input / output devices, network access devices, buses, etc.

[0177] Processor 30 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors, single-chip microcomputers, or any conventional processor.

[0178] The memory 20 can be used to store computer programs 40 and / or modules / units. The processor 30 implements various functions of the vehicle controller 100 by running or executing the computer programs and / or modules / units stored in the memory 20 and by calling the data stored in the memory 20. The memory 20 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the vehicle controller 100 (such as audio data), etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0179] If the modules / units integrated into the vehicle controller 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately added to or subtracted from the content as required by the legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium may not include electrical carrier signals and telecommunication signals.

[0180] In the several embodiments provided in this application, it should be understood that the disclosed vehicle controller and method can be implemented in other ways. For example, the vehicle controller embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and other division methods may be used in actual implementation.

[0181] Furthermore, the functional units in the various embodiments of this application can be integrated into the same processing unit, or each unit can exist physically separately, or two or more units can be integrated into the same unit. The integrated units described above can be implemented in hardware or in the form of hardware plus software functional modules.

[0182] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and not restrictive in all respects. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or vehicle controllers recited in the vehicle controller claims may also be implemented by the same unit or vehicle controller through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.

Claims

1. A vehicle sound transmission method, characterized by, The vehicle sound transmission method includes: Obtain the occupant's voice transmission requirement, wherein the voice transmission requirement indicates the target external sound source that the occupant expects the vehicle to transmit; Collect mixed sounds from outside the vehicle; Based on the sound source separation model and the sound transmission requirements, the target sound signal is separated from the mixed sound outside the vehicle. The target sound signal is the sound signal emitted by the target external sound source. Play the target sound signal inside the vehicle.

2. The vehicle sound transmission method of claim 1, wherein, The sound source separation model includes a language audio modality alignment module and a sound source separation module. The separation of the target sound signal from the mixed external sound based on the sound source separation model includes: The sound transmission requirement is converted into a text describing the target external sound source. Based on the language audio modality alignment module, the target vehicle exterior sound source description text is converted into an embedded representation of the audio modality; The embedded representation and the mixed external sound are input into the sound source separation module to obtain the target sound signal.

3. The vehicle sound transmission method of claim 2, wherein, The sound source separation module includes: A time-frequency conversion layer is used to convert the mixed external sound into a frequency domain signal; An encoding unit, cascaded with the time-frequency conversion layer, comprises an alternately cascaded first fusion layer and an encoding layer; the encoding unit is used to fuse the frequency domain signal with the embedded representation through the first fusion layer to obtain a first fused feature; to encode the first fused feature through the encoding layer to obtain an encoded feature; and to determine the output feature of the encoding unit based on the encoded feature; A decoding unit, cascaded with an encoding unit, comprising alternating cascaded decoding layers and a second fusion layer; the decoding unit is used to decode the output features of the encoding unit through the decoding layers to obtain decoded features; to fuse the decoded features with the embedded representation through the second fusion layer to obtain a second fused feature; and to determine the output features of the decoding unit based on the second fused feature. A reverse time-frequency conversion layer, cascaded with the decoding unit, is used to perform reverse time-frequency conversion on the output features of the decoding unit to obtain the target sound signal.

4. The vehicle sound transmission method of claim 1, wherein, The sound source separation model includes multiple candidate single sound source separation models; The method for separating the target sound signal from the mixed sound outside the vehicle based on the sound source separation model includes: Based on the target external sound source, a target single sound source separation model is selected from the multiple candidate single sound source separation models; The target sound signal is separated from the mixed sound outside the vehicle based on the target single sound source separation model.

5. The vehicle sound transmission method of claim 1, wherein, The sound source separation model includes a multi-sound source separation model; The method for separating the target sound signal from the mixed sound outside the vehicle based on the sound source separation model includes: Based on the multi-source separation model, multiple candidate sound signals are separated from the mixed sound outside the vehicle; The target sound signal is obtained by filtering the sound signals emitted by the target external sound source from the plurality of candidate sound signals.

6. The vehicle sound transmission method of claim 1, wherein, The vehicle includes multiple audio playback devices, and playing the target sound signal inside the vehicle includes: The relative position of the target external sound source and the vehicle is determined based on the target sound signal; The volume distribution ratio of each broadcasting device is determined based on the relative position and the position of the multiple broadcasting devices inside the vehicle; The target sound signal is played by controlling the multiple audio devices according to the volume distribution ratio.

7. The vehicle sound transmission method of claim 1, wherein, The requirement to obtain the occupant's voice transmission includes: In response to the occupant's configuration request for the target external audio source, candidate configuration modes are displayed on the display interface. The candidate configuration modes include at least one of the following: language command configuration mode, audio source type configuration mode, and scene configuration mode. In response to the occupant's selection of a candidate configuration mode, the target configuration mode selected by the occupant is determined; When the target configuration mode includes the language instruction configuration mode, listen to the occupant's language instructions and determine the sound transmission requirement based on the occupant's language instructions; When the target configuration mode includes the sound source type configuration mode, the types of external sound sources are displayed on the display interface, and the sound transmission requirement is determined based on the types of external sound sources selected by the occupant. When the target configuration mode includes a scenario configuration mode, the sound transmission requirement is determined based on the vehicle's driving scenario.

8. A vehicle characterized by comprising: The vehicle includes a vehicle sound transmission device, which comprises: The audio transmission requirement acquisition module is used to acquire the occupant's audio transmission requirement, which indicates the target external sound source that the occupant expects the vehicle to transmit. The signal acquisition module is used to acquire mixed sounds from outside the vehicle. The vehicle exterior sound separation module is used to separate the target sound signal from the mixed sound outside the vehicle based on the sound source separation model and the sound transmission requirements. The target sound signal is the sound signal emitted by the target external sound source. The playback control module is used to play the target sound signal inside the vehicle.

9. A vehicle controller comprising a processor and a memory, wherein, The memory is used to store instructions, and the processor is used to call the instructions in the memory to cause the vehicle controller to execute the vehicle sound transmission method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a vehicle controller, cause the vehicle controller to perform the vehicle sound pass-through method as described in any one of claims 1 to 7.