Method and system for ensuring smooth communication in large private car
By collecting and processing audio data in large private cars, removing noise interference and separating significant sound sources, the problem of voice communication barriers in the car is solved, and a clearer and more reliable passenger communication experience is achieved.
Patent Information
- Application Number
- CN202510526581.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Voice communication barriers are caused by the cockpit structure and complex acoustic environment in large private cars. Especially in high-speed driving, entertainment system is turned on or noise-free scenarios, rear passengers find it difficult to clearly receive voice information from front passengers, affecting driving comfort and safety.
By collecting audio data in the car, extracting reference sound data from the central control device, and performing denoising processing to generate communication voice data. The Conv-TasNet model is used to separate the sound source, extract significant sound sources, and calculate the speaker output sound pressure level according to the microphone position to perform voice gain compensation.
Effectively remove noise interference in the car, accurately separate voice data between passengers, improve the clarity and intelligibility of voice communication, and improve the passenger communication experience.
Smart Images

Figure CN120091256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle central control, and particularly to a method and system for ensuring smooth communication in a large private car. Background Art
[0002] Due to the cockpit structure characteristics of large private cars and the complex acoustic environment, there are significant obstacles to in-vehicle voice communication. In scenarios such as high-speed driving, turning on the in-vehicle entertainment system, or when there is a lot of external noise, rear passengers often have difficulty clearly receiving the voice information of front-row passengers due to the long distance from the front-row sound source and the enhanced environmental noise masking effect. They need to increase the speaking volume or communicate repeatedly, which seriously affects the driving comfort and safety.
[0003] Existing technologies improve the external noise intrusion by improving the body sound insulation materials or sealing design, but such methods can only limitedly improve the signal-to-noise ratio and have no regulatory effect on the existing voice interference in the vehicle; some vehicle models rely on fixed thresholds to adjust the volume or noise reduction intensity, and cannot dynamically optimize the audio output according to the real-time sound source position and in-vehicle noise level.
[0004] Therefore, there is an urgent need for a method to ensure smooth communication in a large private car to break through the limitations of existing technologies. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for ensuring smooth communication in a large private car to solve the problems raised in the above background art.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: A method for ensuring smooth communication in a large private car, the method comprising the following steps: Step S100: Collect in-vehicle audio data and extract reference audio data from the central control device; Step S200: Based on the reference audio data of the central control device, perform noise reduction processing on the audio data to remove the reference audio data in the audio data and generate the communication voice data between passengers; Step S300: Output a sound source signal according to the Conv-TasNet model to construct a sound source set; calculate the time-domain energy of all sound source signals in the sound source set at the current frame, sort the sound source signals in descending order according to the time-domain energy to generate a sound source sequence, and extract the first sound source signal as the significant sound source, and separate the communication voice data according to the significant sound source; Step S400: Calculate the output sound pressure level of each speaker according to the microphone position corresponding to the separated communication voice data, and perform voice gain compensation according to the sound pressure level difference.
[0007] As a preferred technical solution, the collection of in-vehicle audio data includes: Each seat in a large private car is equipped with a microphone and a speaker, and each microphone forms a one-to-one audio unit binding relationship with its corresponding speaker; The voice information inside the car is captured in real time by the microphones on the seats and converted into audio signals, and the central control device converts the audio signals into corresponding audio data.
[0008] It should be noted that in large private cars (such as luxury MPVs, commercial vehicles, etc.), due to the large compartment space and scattered seats, the front and rear row passengers need to raise their voices to communicate. Therefore, each seat is equipped with a high-performance microphone and a high-fidelity speaker, forming a one-to-one audio collection and playback; the microphones on the seats are used to capture the audio information inside the car and convert it into clear audio signals; at the same time, the microphones adopt advanced digital signal processing technology to ensure stable and high-quality audio input in various in-car environments; the speakers on each seat adopt high-quality materials and advanced acoustic technologies, which can accurately restore the audio signals; the central control device, as the processing core of the whole method, is responsible for digitizing and processing the audio signals captured by the microphones. It uses a high-end digital signal processor (DSP) to analyze, optimize and distribute the audio data in real time, and through complex algorithms, it identifies the audio signals of different seats and performs intelligent noise reduction, echo cancellation and audio enhancement processing, so as to ensure that the output of each speaker is clear and pure.
[0009] As a preferred technical solution, the reference audio data includes music data played by the vehicle system, voice prompt information data and environmental noise data.
[0010] It should be noted that inside large private cars, in order to ensure unobstructed communication between passengers, some reference audio data needs to be removed, including music data played by the vehicle system, voice prompt information data and environmental noise data. During the communication between passengers, music playback may interfere with the conversation and affect the clarity and intelligibility of the voice. In order to ensure smooth communication, the currently playing music data needs to be removed to reduce the interference of background music on passenger communication; although voice prompts provide convenient information feedback in the intelligent system, during the communication between passengers, voice prompts may interrupt the conversation and affect the coherence of communication. In order to ensure smooth conversation, unnecessary voice prompt data needs to be removed; there may be various background noises in the car environment, such as engine noise, tire noise, etc. These noises will reduce the communication quality and affect the comfort and satisfaction of passengers. In order to ensure clear and smooth communication, these reference audio data need to be removed.
[0011] As a preferred technical solution, the denoising process of the audio data to generate communication voice data includes: Based on the timestamp of the audio data, extract the reference audio data of the central control device at the timestamp; Remove the reference audio data from the audio data through the normalized least mean square algorithm, and denote the audio data after removing the reference audio data as the alternating current voice data.
[0012] It should be noted that the timestamp of the audio data is used for algorithm positioning and processing of the audio data in a specific time period to ensure that the audio data processed by the algorithm corresponds to the alternating current voice data at the current moment. The reference audio data is used as the input of the algorithm to remove the interference in the audio data. Through the adaptive filtering process of the data normalized least mean square algorithm (NLMS algorithm), the reference audio data is effectively removed, thereby extracting clear alternating current voice data. The NLMS algorithm includes initializing the filter coefficients and parameters, setting the initial value of the adaptive filter coefficient vector, the filter order (i.e., the length of the filter coefficient vector), and the correction step constant of the NLMS algorithm; sampling the input signal and constructing the input vector, including sampling the input signal according to the timestamp of the audio data and constructing the input vector based on the filter order; calculating the filter output and the error signal and updating the filter coefficients, including calculating the filter output based on the filter coefficient vector at the current moment; calculating the error signal based on the alternating current voice data after removing the reference audio data; updating the filter coefficients, including calculating the normalized step factor and updating the filter coefficient vector; repeating the above steps iteratively, so as to effectively remove the interference (reference audio data) in the audio data, extract clear alternating current voice data (the voice of the passenger conversation), and enhance the communication experience between passengers.
[0013] As a preferred technical solution, the source separation processing of the alternating current voice data includes: Construct an alternating current voice data set , where represents the alternating current voice data corresponding to the j-th microphone, and J represents the total number of microphones; Segment each alternating current voice data into short frames, and input the short frames into the Conv-TasNet model to obtain multiple source signals, and construct a source set , where represents the i-th source signal, and I represents the total number of source signals; Based on each source signal , calculate the time-domain energy of the current frame, and the calculation formula is: ; where N represents the total number of sampling points of the current frame, represents the waveform signal value of the i-th source signal at the sampling point t; Calculate the time-domain energy of all sound source signals in the sound source concentration at the current frame. Sort the sound source signals in descending order according to the time-domain energy to generate a sound source sequence, and extract the first sound source signal as the significant sound source; Extract the significant sound sources corresponding to all the AC voice data in the AC voice dataset, and calculate the similarity of the significant sound sources. The calculation formula is: ; Among them, represents the similarity between the a-th and b-th significant sound sources, and , and respectively represent the waveform signal values of the a-th and b-th significant sound sources at the sampling point t, and respectively represent the means of the waveform signal values of the a-th and b-th significant sound sources; When the similarity between any two significant sound sources is greater than or equal to the sound source threshold, compare the time-domain energy of the current frame; Extract the sound source sequence of the AC voice data with small time-domain energy, delete the first sound source signal, optimize the sound source sequence, and extract the first sound source signal to update the significant sound source; When the similarity between any two significant sound sources is less than the sound source threshold, output the separated AC voice data according to the voiceprint characteristics of the significant sound sources and using the sound separation technology.
[0014] It should be noted that the Conv-TasNet model mainly consists of three parts: an encoder (Encoder), a separator (Separator), and a decoder (Decoder); among them, the encoder is used to convert the mixed waveform into a high-resolution feature representation; the encoder extracts audio features through one-dimensional convolution operations to generate a feature matrix, and each element of this matrix represents the features of the audio signal at a specific time and frequency; the separator estimates the mask (Mask) of each sound source based on the Time Convolution Network (TCN). The TCN consists of multiple one-dimensional convolutional blocks, and each convolutional block contains dilated convolution and depthwise separable convolution operations to model the long-term dependence of the speech signal. By applying these masks to the feature matrix, the features of different sound sources can be separated; the decoder converts the separated feature matrix back into a waveform; the decoder reconstructs the source waveform through one-dimensional transposed convolution operations (ConvTrans1D) to generate the separated audio signal, and then performs similarity calculations on the separated audio signals (significant sound sources) to exclude the possibility of extracting the same audio signal as the significant sound source from different seats, thereby improving the accuracy of separating AC voice data.
[0015] As a preferred technical solution, calculating the output sound pressure level of each speaker according to the microphone position corresponding to the separated AC voice data includes: Obtaining the microphone position corresponding to the separated AC voice data and calculating the acoustic attenuation of the AC voice data. The calculation formula is: ; wherein, represents the acoustic attenuation from the speaker at seat q to the microphone at seat m; represents the distance from the speaker at seat q to the microphone at seat m; represents a preset reference distance; represents the reverberation time in the vehicle; represents the environmental reflection coefficient; Based on the acoustic attenuation of the AC voice data, calculating the effective sound pressure level received by the microphone at seat m: , wherein, represents the effective sound pressure level received by the microphone at seat m from the speaker at seat q, represents the sound pressure level output by the speaker at seat q.
[0016] As a preferred technical solution, the voice gain compensation includes: When the effective sound pressure level received by the microphone at seat m from the speaker at seat q is less than the sound pressure threshold, outputting the sound pressure level of the AC voice data corresponding to seat m output by the speaker at seat q.
[0017] It should be noted that acoustic attenuation, that is, the attenuation of sound waves, refers to the phenomenon that when sound waves propagate in a medium, their intensity (or sound energy) gradually weakens as the propagation distance increases. In the present invention, by calculating the acoustic attenuation between seats, the echo heard by passengers can be avoided, and the communication experience between passengers can be improved.
[0018] A system for ensuring smooth communication in a large private vehicle includes an acquisition module, a denoising module, a separation module, and an output module: The acquisition module includes a sound pickup unit and a speaker unit configured for multiple seats, forming a physical position binding, and is used to capture the audio information in the vehicle and convert it into an audio signal; It should be noted that the sound pickup unit can capture the audio information in the vehicle in all directions without dead angles, including the communication between passengers, the sound of the in-vehicle entertainment system, and possible environmental noises, etc. The speaker unit is used to convert the processed audio signal into the output AC voice data to achieve high-quality audio output; the acquisition module converts the captured audio information into a digital audio signal in real time through advanced audio coding technology and directly transmits it to the audio bus of the central control device to ensure the real-time performance and integrity of the audio data.
[0019] The denoising module includes a data generation unit and a processing unit. The data generation unit is used to convert the audio signal collected by the acquisition module into audio data; the processing unit is used to denoise the audio data converted by the data generation unit to generate AC voice data; It should be noted that the data generation unit receives the audio signal from the acquisition module, and uses an efficient audio decoding algorithm to convert it into audio data that can be processed, ensuring the accuracy and reliability of the audio data, and providing a high-quality data source for subsequent audio processing; the processing unit uses an advanced denoising algorithm (Normalized Least Mean Square algorithm) to perform real-time denoising processing on the audio data; this algorithm can intelligently identify and remove reference audio data in the audio data, such as music playback data, voice prompt data, road noise, wind noise, engine noise, etc., thus significantly improving the clarity of the audio data.
[0020] The separation module includes a significant sound source generation unit and a separation unit. The significant sound source generation unit is used to update the significant sound source, and the separation unit separates the AC voice data corresponding to the significant sound source through sound separation technology; It needs to be explained that the significant sound source generation unit updates the significant sound source information in the vehicle in real time through the Conv-TasNet model, intelligently identifies and tracks the main sound sources in the vehicle, such as the voices of passengers, providing accurate sound source positioning for sound separation; the separation unit, based on the sound source positioning information provided by the significant sound source generation unit, uses sound separation technology to separate the AC voice data corresponding to the significant sound source from the mixed audio data, ensuring that the AC voice data between passengers is accurately separated, reducing interference, and improving the clarity and accuracy of communication.
[0021] The output module is used to calculate the output sound pressure level of each speaker according to the microphone position and acoustic parameters; It needs to be explained that the output module calculates the output sound pressure level of each speaker according to the separated AC voice data and the acoustic environment in the vehicle, ensuring that passengers can obtain balanced and clear audio output at different seat positions; at the same time, the output module can also dynamically adjust the output sound pressure level of the speaker according to the noise level in the vehicle and the auditory needs of passengers (such as avoiding echo) to provide a personalized audio experience.
[0022] The acquisition module is directly connected to the audio bus of the central control device; The denoising module, separation module and output module are executed by a multi-modal audio processing chip.
[0023] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: In a method and system for ensuring smooth communication in a large private car provided by the present invention, it includes collecting audio data inside the car and extracting reference audio data of the central control device; performing denoising processing on the audio data to generate communication voice data; separating the communication voice data according to significant sound sources; calculating the output sound pressure levels of each speaker based on the microphone positions corresponding to the separated communication voice data, and performing voice gain compensation; by removing the reference audio data from the audio data, precise selection of the communication voice data among passengers is achieved; through precise extraction of significant sound sources, precise matching between the passenger's voice and the seat is achieved, and the sound source position is updated in real time; through the cooperation of the output sound pressure level and voice gain compensation, not only the audio output is optimized, but also the echo caused by speaker amplification is avoided, improving the communication experience of passengers. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them: Figure 1 It is a schematic diagram of the steps in the embodiments of the present invention; Figure 2 It is a schematic flowchart in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present invention in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the described embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.
[0026] As Figure 1 , Figure 2 shown, it is an embodiment of the present invention. This embodiment provides a method for ensuring smooth communication in a large private car, and the method includes the following steps: Step S100: Collect audio data inside the car and extract reference audio data from the central control device; Step S200: Based on the reference audio data of the central control device, perform denoising processing on the audio data to remove the reference audio data from the audio data and generate communication voice data between passengers; Step S300: Output the sound source signals according to the Conv-TasNet model, construct a sound source set; calculate the time-domain energy of all sound source signals in the sound source set at the current frame, sort the sound source signals in descending order according to the time-domain energy to generate a sound source sequence, extract the first sound source signal as the significant sound source, and separate the AC voice data according to the significant sound source; Step S400: Calculate the output sound pressure level of each speaker according to the microphone positions corresponding to the separated AC voice data, and perform voice gain compensation according to the sound pressure level difference.
[0027] Specifically, the audio data collected in the vehicle includes: On each seat in a large private car, a high-sensitivity microphone and a high-quality speaker are configured; each microphone is physically bound to its corresponding speaker to ensure the accurate capture and output of audio signals.
[0028] The microphone adopts omnidirectional or directional pickup technology, captures the audio information in the vehicle in real time through the microphone on the seat and converts it into an audio signal, and the central control device converts the audio signal into audio data through the built-in audio interface for subsequent processing.
[0029] Specifically, the reference audio data includes music data, voice prompt information data, and environmental noise data played by the vehicle-mounted system; the reference audio data is used for subsequent noise reduction processing to improve the clarity and accuracy of the audio data.
[0030] Specifically, the noise reduction processing of the audio data to generate AC voice data includes: Based on the timestamp information of the audio data, extract the reference audio data of the central control device at the same timestamp. The timestamp matching ensures that the noise reduction processing is consistent with the real-time nature of the audio data; Remove the reference audio data from the audio data through the normalized least mean square algorithm, and record the audio data after removing the reference audio data as AC voice data; the normalized least mean square algorithm removes the reference audio data component in the audio data through adaptive filtering technology, and retains the AC voice data between passengers for subsequent sound separation processing.
[0031] Specifically, the sound source separation processing of the AC voice data includes: Construct an AC voice data set , where represents the AC voice data corresponding to the jth microphone, and J represents the total number of microphones; For each AC voice data Segment it into short frames, and input the short frames into the Conv-TasNet model to obtain multiple sound source signals, and construct a sound source set , where denotes the i-th sound source signal, and I denotes the total number of sound source signals; Based on each sound source signal , calculate the time-domain energy of the current frame. The calculation formula is: ; where N represents the total number of sampling points in the current frame, denotes the waveform signal value of the i-th sound source signal at the sampling point t; Calculate the time-domain energy of all sound source signals in the sound source set. Sort the sound source signals in descending order according to the time-domain energy to generate a sound source sequence, and extract the first sound source signal as the prominent sound source; Extract the prominent sound sources corresponding to all the communication voice data in the communication voice dataset, and calculate the similarity of the prominent sound sources. The calculation formula is: ; where, denotes the similarity between the a-th and b-th prominent sound sources, and , and respectively denote the waveform signal values of the a-th and b-th prominent sound sources at the sampling point t, and respectively denote the means of the waveform signal values of the a-th and b-th prominent sound sources; When the similarity between any two prominent sound sources is greater than or equal to the sound source threshold, compare the time-domain energy of the current frame; Extract the sound source sequence of the communication voice data with small time-domain energy, delete the first sound source signal, optimize the sound source sequence, and extract the first sound source signal to update the prominent sound source; When the similarity between any two prominent sound sources is less than the sound source threshold, output the separated communication voice data according to the voiceprint characteristics of the prominent sound sources and using the sound separation technology.
[0032] Specifically, the calculation of the output sound pressure level of each speaker according to the microphone position corresponding to the separated communication voice data includes: Obtain the microphone position corresponding to the separated communication voice data, and calculate the acoustic attenuation of the communication voice data. The calculation formula is: ; where, denotes the acoustic attenuation from the speaker at seat q to the microphone at seat m; denotes the distance from the speaker at seat q to the microphone at seat m; denotes the preset reference distance; denotes the reverberation time in the vehicle; denotes the environmental reflection coefficient; Based on the acoustic attenuation of the AC voice data, calculate the effective sound pressure level received by the microphone at seat m: , where represents the effective sound pressure level received by the microphone at seat m from the speaker at seat q, represents the sound pressure level output by the speaker at seat q.
[0033] Specifically, the voice gain compensation includes: When the effective sound pressure level received by the microphone at seat m from the speaker at seat q is less than the sound pressure threshold, output the sound pressure level of the AC voice data corresponding to seat m output by the speaker at seat q.
[0034] Specifically, it further includes a system for ensuring smooth communication in a large private car. The system includes an acquisition module, a denoising module, a separation module, and an output module: The acquisition module includes a pickup unit and a speaker unit configured for multiple seats, forming a physical position binding, and is used to capture the audio information in the vehicle and convert it into an audio signal; The denoising module includes a data generation unit and a processing unit. The data generation unit is used to convert the audio signal collected by the acquisition module into audio data; the processing unit is used to denoise the audio data converted by the data generation unit to generate AC voice data; The separation module includes a significant sound source generation unit and a separation unit. The significant sound source generation unit is used to update the significant sound source, and the separation unit separates the AC voice data corresponding to the significant sound source through sound separation technology; The output module is used to calculate the output sound pressure level of each speaker according to the microphone position and acoustic parameters; The acquisition module is directly connected to the audio bus of the central control device to ensure real-time transmission and processing of audio data; The denoising module, the separation module, and the output module are executed by a multi-modal audio processing chip to improve the processing efficiency and stability of the system; at the same time, the multi-modal processing chip supports a variety of audio processing algorithms and models, and can be flexibly configured and optimized according to actual needs.
[0035] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, it can be implemented in whole or in part in the form of a computer program product, which includes one or more computer instructions. When loading and executing the computer program instructions on a computer, the processes or functions according to the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0036] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0037] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various changes or substitutions thereof, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for ensuring smooth communication in a large private car, characterized in that: include: Collect audio data in the car and extract reference sound data from the central control device; De-noising the audio data to generate voice data for communication between passengers; According to the significant sound sources, the communication voice data is processed for sound source separation; According to the microphone positions corresponding to the separated AC voice data, the output sound pressure level of each speaker is calculated, and voice gain compensation is performed according to the sound pressure level difference.
2. A method for ensuring smooth communication in a large private car according to claim 1, characterized in that: The audio data collected in the vehicle includes: Each seat in a large private car is equipped with a microphone and a speaker, and each microphone forms a one-to-one audio unit binding relationship with its corresponding speaker; The microphone on the seat captures the voice information in the car in real time and converts it into audio signals, and the central control device converts the audio signals into corresponding audio data.
3. A method for ensuring smooth communication in a large private car according to claim 2, characterized in that: The reference sound data includes music data played by the vehicle-mounted system, voice prompt information data and environmental noise data.
4. A method for ensuring smooth communication in a large private car according to claim 3, characterized in that: Performing denoising on the audio data to generate communication language data includes: Based on the timestamp of the audio data, extracting the reference audio data of the central control device at the timestamp; The reference sound data in the audio data is removed by a normalized least mean square algorithm, and the audio data without the reference sound data is recorded as communication speech data.
5. A method for ensuring smooth communication in a large private car according to claim 4, characterized in that: The sound source separation process of the communication voice data comprises: Constructing a communication speech dataset ,in, represents the communication speech data corresponding to the jth microphone, and J represents the total number of microphones; Each communication language data Split into short frames and input the short frames into the Conv-TasNet model to obtain multiple sound source signals and construct a sound source set ,in, represents the i-th sound source signal, and I represents the total number of sound source signals; Based on each sound source signal , calculate the time domain energy of the current frame, the calculation formula is: ; Where N represents the total number of sampling points in the current frame. Represents the waveform signal value of the i-th sound source signal at sampling point t; Calculate the time domain energy of all sound source signals in the sound source set in the current frame, sort the sound source signals in descending order according to the time domain energy to generate a sound source sequence, and extract the first sound source signal as a significant sound source; Extract the significant sound sources corresponding to all the communication speech data in the communication speech dataset, and calculate the similarity of the significant sound sources. The calculation formula is: ; in, represents the similarity between the ath and bth significant sound sources, and , and Respectively represent the waveform signal values of the ath and bth significant sound sources at sampling point t, and Respectively represent the mean of the waveform signal values of the ath and bth significant sound sources; When the similarity between any two significant sound sources is greater than or equal to the sound source threshold, compare the time domain energy of the current frame; Extract the sound source sequence of the communication speech data with small time domain energy, delete the first sound source signal, optimize the sound source sequence, and extract the first sound source signal to update the significant sound source; When the similarity between any two significant sound sources is less than the sound source threshold, the separated communication voice data is output according to the voiceprint features of the significant sound sources and using sound separation technology.
6. A method for ensuring smooth communication in a large private car according to claim 5, characterized in that: Calculating the output sound pressure level of each speaker according to the microphone position corresponding to the separated communication voice data includes: The microphone position corresponding to the separated AC voice data is obtained, and the acoustic attenuation of the AC voice data is calculated using the following formula: ; in, represents the acoustic attenuation from the loudspeaker at seat q to the microphone at seat m; represents the distance from the speaker at seat q to the microphone at seat m; Indicates the preset reference distance; Indicates the reverberation time in the car; represents the environmental reflection coefficient; Based on the acoustic attenuation of the communication speech data, calculate the effective sound pressure level received by the microphone at seat m: ,in, It represents the effective sound pressure level of the speaker at seat q received by the microphone at seat m. Indicates the sound pressure level output by the speaker at seat q.
7. A method for ensuring smooth communication in a large private car according to claim 6, characterized in that: The speech gain compensation comprises: When the effective sound pressure level of the speaker at seat q received by the microphone at seat m is less than the sound pressure threshold, the speaker at seat q outputs the sound pressure level of the corresponding communication voice data at seat m.
8. A communication system for ensuring smooth communication in a large private car according to any one of claims 1 to 7, the system comprising a collection module, a denoising module, a separation module and an output module, and its functions include: The acquisition module includes a plurality of sound pickup units and a speaker unit configured at the seats, forming a physical position binding, and is used to capture the audio information in the car and convert it into an audio signal; The denoising module includes a data generating unit and a processing unit, wherein the data generating unit is used to convert the audio signal collected by the collecting module into audio data; the processing unit is used to denoise the audio data converted by the data generating unit to generate communication voice data; The separation module includes a significant sound source generation unit and a separation unit, wherein the significant sound source generation unit is used to update the significant sound source, and the separation unit separates the communication voice data corresponding to the significant sound source through a sound separation technology; The output module is used to calculate the output sound pressure level of each speaker according to the microphone position and acoustic parameters; The acquisition module is directly connected to the audio bus of the central control device; The denoising module, separation module and output module are executed by a multimodal audio processing chip.
Citation Information
Patent Citations
Distributed microphone pickup system and method in complex scene
CN111161751A
In-vehicle audio transmission method and device, cabin host of vehicle and storage medium
CN116866784A
Anti-interference voice recognition method and device, electronic equipment and storage medium
CN118430537A
Intelligent human-shaped accompanying robot and voice interaction system thereof
CN119495288A
Method of enhancing voice quality and apparatus thereof
KR101649710B1
Cited By
Digital audio noise reduction method based on time-frequency mask separation
CN120636429A