A vehicle-mounted speech processing method, device, equipment and storage medium

By using beamforming algorithms and noise residual estimation techniques in the in-vehicle environment, the interference problem in the speech recognition of the driver and passenger was solved, achieving optimal separation and accurate service of speech in the driver and passenger areas.

CN114360529BActive Publication Date: 2026-01-27MOBILITY ASIA SMART TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011057299.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-29
Publication Date
2026-01-27
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

In a vehicle environment, when the driver and passenger speak at the same time, existing technology cannot accurately recognize the voice of either the driver or passenger, resulting in voice interference, inability to provide accurate services, and a poor user experience.

Method used

By acquiring the speech from the driver and passenger directions, beamforming algorithm is used to extract the speech from the driver and passenger, and noise residual estimation and noise and residual suppression are performed respectively to obtain accurate speech from the driver and passenger.

Benefits of technology

It achieves optimal separation of voice commands in the driver and passenger areas, avoids voice interference, provides accurate voice services, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360529B_ABST
    Figure CN114360529B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of vehicle-mounted voice processing method, device, equipment and storage medium.The method comprises: obtaining main co-pilot direction voice, and extracting main driver direction first voice and co-pilot direction first voice from main co-pilot direction voice;Residual noise estimation is carried out on main driver direction first voice and co-pilot direction first voice respectively, to determine first estimated noise residue and second estimated noise residue;According to first estimated noise residue and second estimated noise residue, noise and residual suppression is carried out on main driver direction first voice and co-pilot direction first voice respectively, to obtain main driver direction second voice and co-pilot direction second voice.The method can separate the voice of main and co-pilot area optimally, avoid the interference of main to vice or vice to main, to provide accurate main and co-pilot area voice, facilitate accurate voice recognition and provide accurate service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice processing technology, and in particular to an in-vehicle voice processing method, apparatus, device and storage medium. Background Technology

[0002] With the development of communication and technological advancements, voice interaction systems in in-vehicle environments are receiving increasing attention. In these environments, voice interaction systems can meet drivers' needs, such as checking the weather, navigation, or adjusting the air conditioning temperature. These systems can also help prevent driver distraction and improve driver focus while driving.

[0003] However, in an in-vehicle environment, there are often scenarios where the driver and passenger speak simultaneously. Treating the voices of the driver and passenger as a whole for speech recognition cannot accurately identify either the driver's or passenger's voice, resulting in residual voices from either driver or passenger, causing speech interference, failing to provide accurate service, and leading to a poor user experience. Summary of the Invention

[0004] This invention provides an in-vehicle voice processing method, apparatus, device, and storage medium that can optimally separate voice signals in the driver and passenger areas, avoiding interference between the driver and passenger or vice versa, to provide accurate voice signals for both driver and passenger areas.

[0005] In a first aspect, embodiments of the present invention provide an in-vehicle voice processing method, the method comprising:

[0006] Acquire the voice from the driver and passenger directions, and extract the first voice from the driver's direction and the first voice from the passenger's direction from the voice from the driver and passenger directions;

[0007] Noise residual estimation is performed on the first voice in the driver's direction and the first voice in the passenger's direction respectively to determine the first estimated noise residual and the second estimated noise residual.

[0008] Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

[0009] Secondly, embodiments of the present invention also provide an in-vehicle voice processing device, the device comprising:

[0010] The voice acquisition module is used to acquire voice from the driver and passenger directions, and extract the first voice from the driver direction and the first voice from the passenger direction from the driver and passenger directions.

[0011] The noise residual estimation module is used to estimate the noise residual of the first voice in the driver's direction and the first voice in the passenger's direction, respectively, and to determine the first estimated noise residual and the second estimated noise residual.

[0012] The noise and residual suppression module is used to perform noise and residual suppression on the first voice in the driver's direction and the first voice in the passenger's direction based on the first estimated noise residual and the second estimated noise residual, respectively, to obtain the second voice in the driver's direction and the second voice in the passenger's direction.

[0013] Thirdly, embodiments of the present invention also provide an electronic device, the device comprising:

[0014] One or more processors;

[0015] Storage device for storing one or more programs.

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement an in-vehicle voice processing method as described in any embodiment of the present invention.

[0017] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an in-vehicle voice processing method as described in any embodiment of the present invention.

[0018] The technical solution of this invention acquires speech in the driver and passenger directions and extracts first speech in the driver's direction and first speech in the passenger's direction from the speech in the driver's direction; performs noise residual estimation on the first speech in the driver's direction and the first speech in the passenger's direction respectively to determine the first estimated noise residual and the second estimated noise residual; and performs noise and residual suppression on the first speech in the driver's direction and the first speech in the passenger's direction respectively based on the first estimated noise residual and the second estimated noise residual to obtain the second speech in the driver's direction and the second speech in the passenger's direction. This solves the problem of speech separation in the driver and passenger directions, achieves optimal separation of speech in the driver and passenger areas, avoids interference between driver and passenger or between passenger and driver, and provides accurate speech in the driver and passenger areas. Attached Figure Description

[0019] Figure 1 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 1 of the present invention;

[0020] Figure 2 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 2 of the present invention;

[0021] Figure 3 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 2 of the present invention;

[0022] Figure 4 This is a schematic diagram of the structure of an in-vehicle voice processing device provided in Embodiment 3 of the present invention;

[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0024] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0025] Example 1

[0026] Figure 1 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where the voice in the driver and passenger directions is separated and residual is suppressed in an in-vehicle voice interaction system. This method can be executed by an in-vehicle voice processing device, which can be implemented through software and / or hardware. The device can be integrated into an electronic device such as an in-vehicle terminal. Figure 1 As shown, the method specifically includes:

[0027] Step 110: Obtain the voice from the driver and passenger sides, and extract the first voice from the driver's side and the first voice from the passenger side.

[0028] In this context, the voice input from the driver and passenger sides can be a mixture of voice signals received by the microphone in the in-vehicle voice interaction system, originating from both the driver's and passenger's directions. In a vehicle environment, it's common for the driver and passenger to communicate or issue voice commands to in-vehicle devices. Because both driver and passenger voices are voice signals, not uncommon interference signals like noise, they are difficult to separate, leading to mutual interference. For example, when the driver uses voice control to operate in-vehicle devices, the passenger's voice may become interference, or vice versa. Separating the driver and passenger voice input is crucial for accurately determining voice commands to in-vehicle devices and providing precise service to users.

[0029] To accurately separate the driver's and passenger's voice, it is first necessary to perform segmentation recognition of the voice in both directions. Existing segmentation recognition algorithms can be used to extract the first voice from the driver's direction and the first voice from the passenger's direction. The first voice from the driver's direction refers to the voice obtained by enhancing the driver's voice and suppressing the passenger's voice. The first voice from the passenger's direction refers to the voice obtained by enhancing the passenger's voice and suppressing the driver's voice.

[0030] To achieve good performance in regional recognition, that is, when the speech recognition system recognizes the first speech from the driver's direction or the first speech from the passenger's direction, it should not easily recognize uninteresting sounds. In one embodiment of the present invention, optionally, extracting the first speech from the driver's direction and the first speech from the passenger's direction from the speech may include: extracting the first speech from the driver's direction and the first speech from the passenger's direction from the speech using a beamforming algorithm.

[0031] Beamforming algorithms combine antenna technology with digital signal processing technology and can be used for the transmission and reception of directional signals. Beamforming algorithms can be categorized into adaptive beamforming algorithms, fixed beamforming algorithms, and switching beamforming algorithms. In this embodiment, any beamforming algorithm can be used to perform segmented recognition of the driver and passenger-side voice signals collected by the microphone array. It should be noted that while beamforming algorithms can suppress unwanted sounds, the suppression effect is not ideal, and residual interference to the sounds of interest remains.

[0032] For example, by using an adaptive beamforming algorithm, different antenna gains can be applied to different directions of arrival based on the different propagation paths of the driver's and co-driver's voices in space. This allows for the real-time formation of a narrow beam aimed at the driver's or co-driver's direction, while minimizing sidelobes in other directions and employing directional reception to improve the voice signal-to-noise ratio in the driver's or co-driver's direction.

[0033] To illustrate the driver's direction first voice and passenger's direction first voice in this embodiment, in one embodiment of the present invention, optionally, the driver's direction first voice includes driver's direction voice, passenger's direction voice residue, and first noise; the passenger's direction first voice includes passenger's direction voice, driver's direction voice residue, and second noise.

[0034] The "driver's direction speech" refers to the pure speech of the driver. The "passenger's direction speech residue" refers to the interference caused by the passenger's speech on the driver's direction speech. "First noise" refers to the interference signal caused by environmental noise on the driver's direction speech. The driver's direction first speech is a mixture of the driver's direction speech, the passenger's direction speech residue, and the first noise, where the passenger's direction speech residue is the residual signal after beamforming. The passenger's direction speech refers to the pure speech of the passenger. The driver's direction speech residue refers to the interference caused by the driver's speech on the passenger's direction speech. "Second noise" refers to the interference signal caused by environmental noise on the passenger's direction speech. The passenger's direction first speech is a mixture of the passenger's direction speech, the driver's direction speech residue, and the second noise, where the driver's direction speech residue is the residual signal after beamforming.

[0035] It should be noted that the same beamforming algorithm logic can be used when extracting the first speech from the driver's side and the first speech from the passenger side, but the parameters can be different. For example, when extracting the first speech from the driver's side, the speech from the driver's side can be enhanced while the speech from the passenger side can be suppressed. Conversely, when extracting the first speech from the passenger side, the speech from the passenger side can be enhanced while the speech from the driver's side can be suppressed.

[0036] Step 120: Perform noise residual estimation on the first voice in the driver's direction and the first voice in the passenger's direction respectively, and determine the first estimated noise residual and the second estimated noise residual.

[0037] The noise residual estimation can involve noise estimation and / or residual estimation of the speech. Noise estimation can be performed using algorithms such as recursive averaging noise estimation, minimum tracking, or histogram noise estimation. The residual estimation can be determined by considering either the influence coefficient of the passenger's speech on the driver's speech or the influence coefficient of the driver's speech on the passenger's speech. The influence coefficient can be determined based on ambient noise; the higher the ambient noise, the smaller the influence coefficient; conversely, the lower the ambient noise, the larger the influence coefficient. The influence coefficient can also be related to volume; for example, the louder the passenger's voice, the greater its influence on the driver.

[0038] Step 130: Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

[0039] Noise and residue suppression refers to denoising the speech to obtain clean speech free from noise and residue. Noise reduction can employ algorithms such as Wiener filtering or spectral subtraction. In this embodiment, the second speech from the driver's direction and the second speech from the passenger's direction can be obtained by denoising the first speech from the driver's direction and the first speech from the passenger's direction, respectively. The second speech from the driver's direction and the second speech from the passenger's direction can approximate the speech from the driver's direction and the passenger's direction, respectively, thus suppressing noise and residue, achieving optimal segmentation recognition without compromising normal speech.

[0040] The technical solution of this embodiment acquires the speech in the driver and passenger directions and extracts the first speech in the driver direction and the first speech in the passenger direction from the speech in the driver and passenger directions; performs noise residual estimation on the first speech in the driver direction and the first speech in the passenger direction respectively to determine the first estimated noise residual and the second estimated noise residual; and performs noise and residual suppression on the first speech in the driver direction and the first speech in the passenger direction respectively based on the first estimated noise residual and the second estimated noise residual to obtain the second speech in the driver direction and the second speech in the passenger direction. This solves the problem of speech separation in the driver and passenger directions, achieves optimal separation of speech in the driver and passenger areas, and avoids residual interference and noise interference between driver and passenger or between passenger and driver.

[0041] Example 2

[0042] Figure 2 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes:

[0043] Step 210: Obtain the voice from the driver and passenger directions, and extract the first voice from the driver direction and the first voice from the passenger direction using a beamforming algorithm.

[0044] The first voice in the driver's direction includes the voice in the driver's direction, the voice residue in the passenger's direction, and the first noise; the first voice in the passenger's direction includes the voice in the passenger's direction, the voice residue in the driver's direction, and the second noise.

[0045] Step 220: Perform noise estimation on the first voice from the driver's side and the first voice from the passenger side, respectively, and determine the first noise and the second noise.

[0046] For example, after beamforming, the first speech in the driver's direction is y1(t, f), and the first speech in the passenger's direction is y2(t, f). Here, t is time, f is frequency, and y1(t, f) represents the amplitude of y1 in the f frequency band at time t. y1(t, f) and y2(t, f) can be specifically represented as: y1(t, f) = s1(t, f) + r1(t, f) + n1(t, f) and y2(t, f) = s2(t, f) + r2(t, f) + n2(t, f). Here, s1(t, f) and s2(t, f) represent the speech in the driver's direction and the passenger's direction, respectively. r1(t, f) and r2(t, f) represent the residual speech in the passenger's direction and the residual speech in the driver's direction, respectively. n1(t, f) and n2(t, f) represent the first noise and the second noise, respectively.

[0047] Noise estimation for y1(t, f) and y2(t, f) can be performed using noise estimation algorithms such as recursive averaging, minimum tracking, or histogram noise estimation to estimate n1(t, f) and n2(t, f) respectively. In this case, the speech in the driver's direction includes the speech in the driver's direction and the speech residue in the passenger's direction, i.e., y1(t, f) - n1(t, f). The speech in the passenger's direction includes the speech in the passenger's direction and the speech residue in the driver's direction, i.e., y2(t, f) - n2(t, f).

[0048] Step 230: Perform residual estimation on the first voice in the driver's direction and the first voice in the passenger's direction respectively, and determine the first residual and the second residual.

[0049] The residual estimation can be a specific representation of uninterested speech to interested speech. In this embodiment, the first residual of the co-driver's speech to the driver's speech can be obtained by multiplying the influence coefficient by the co-driver's speech; the second residual of the driver's speech to the co-driver's speech can be obtained by multiplying the influence coefficient by the driver's speech.

[0050] In practice, neither the driver's voice nor the passenger's voice can be obtained as pure speech; both will contain remnants of the other. To specifically describe how the first and second remnants are determined, in one embodiment of this invention, optionally, the first remnant is the difference between the first speech from the passenger's direction and the second noise; the second remnant is the difference between the first speech from the driver's direction and the second noise.

[0051] Wherein, y2(t,f)-n2(t,f) can be the passenger-side voice residual with the largest first voice from the driver's direction, and y2(t,f)-n2(t,f) can be considered as the first residual. y1(t,f)-n1(t,f) can be the driver-side voice residual with the largest first voice from the passenger direction, and y1(t,f)-n1(t,f) can be considered as the second residual.

[0052] Step 240: Determine the first estimated noise residue based on the first noise and the first residue.

[0053] The first estimated noise residue can be determined by logical operations between the first noise and the first residue. For example, it can be in the form of a piecewise function or a linear expression. An exemplary first noise residue can be expressed as N1(t,f)=a×n1(t,f)+b×[y2(t,f)-n2(t,f)], where a and b are constants representing weights, and a+b=1.

[0054] To more accurately determine the first estimated noise residue, the phenomenon of sound masking can be utilized when estimating the noise residue in speech. For example, the human ear can distinguish faint sounds in a quiet environment, but in a noisy environment, these faint sounds are drowned out by background noise. This phenomenon, where the presence of one sound raises the hearing threshold of another, is called sound masking.

[0055] In one embodiment of the present invention, optionally, determining the first estimated noise residue based on the first noise and the first residue includes: determining the maximum value of each frequency band of the first noise and the first residue as the first estimated noise residue.

[0056] When considering the masking effect of sound, removing the interference from larger sounds also removes the interference from smaller sounds. Therefore, the maximum value N1(t,f) of each frequency band can be selected from all the data of n1(t,f) and y2(t,f)-n2(t,f) as the first estimated noise residue, that is, N1(t,f)=MAX([n1(t,f)],[y2(t,f)-n2(t,f)]), where MAX(a,b) means taking the maximum value of a and b.

[0057] Step 250: Determine the second estimated noise residue based on the second noise and the second residue.

[0058] The second estimated noise residue can be determined by logical operations between the second noise and the second residue. For example, it can be in the form of a piecewise function or a linear expression. An exemplary second noise residue can be expressed as N2(t,f)=a×n2(t,f)+b×[y1(t,f)-n1(t,f)], where a and b are constants representing weights, and a+b=1.

[0059] To more accurately determine the first estimated noise residue, sound masking can be utilized when estimating the noise residue of speech. In one embodiment of the present invention, optionally, determining the second estimated noise residue based on the second noise and the second residue includes: determining the maximum value of each frequency band of the second noise and the second residue as the first estimated noise residue.

[0060] When considering the masking effect of sound, removing the interference from larger sounds also removes the interference from smaller sounds. Therefore, the maximum value N2(t,f) of each frequency band can be selected from all the data of n2(t,f) and y1(t,f)-n1(t,f) as the first estimated noise residue, i.e., N2(t,f)=MAX([n2(t,f)],[y1(t,f)-n1(t,f)]).

[0061] Step 260: Take the difference between the first voice in the driver's direction and the first estimated noise residue as the second voice in the driver's direction.

[0062] The process involves denoising the first speech in the driver's direction, with the estimated residual noise used as noise during denoising. For example, y1(t,f)-N1(t,f) can be used as the second speech in the driver's direction. The suppression of the passenger's speech is strong, allowing for further processing of the first speech in the driver's direction after beamforming, leaving almost no residual passenger speech. Furthermore, the maximum denoising level needs to be limited to avoid speech distortion. For example, y1(t,f)-k×N1(t,f) can be used as the second speech in the driver's direction, where k is a constant, such as 0.8.

[0063] Step 270: Take the difference between the first voice in the co-driver's direction and the second estimated noise residue as the second voice in the co-driver's direction.

[0064] The process involves denoising the first speech from the passenger's direction, with the estimated residual noise used as noise during denoising. For example, y2(t,f)-N2(t,f) can be used as the second speech from the passenger's direction. The suppression of the driver's speech is strong, allowing for further processing of the first speech from the passenger's direction after beamforming, leaving almost no residual driver's speech. Furthermore, the maximum denoising level needs to be limited to avoid speech distortion. For example, y2(t,f)-k×N2(t,f) can be used as the second speech from the passenger's direction, where k is a constant, such as 0.8.

[0065] It should be noted that this embodiment is not limited to the above execution order. For example, step 250 can be executed first and then step 240; step 270 can be executed first and then step 260.

[0066] The technical solution of this embodiment acquires the speech in the driver and passenger directions and extracts the first speech in the driver's direction and the first speech in the passenger's direction using a beamforming algorithm. Noise estimation is performed on the first speech in the driver's direction and the first speech in the passenger's direction to determine the first noise and the second noise. Residual estimation is performed on the first speech in the driver's direction and the first speech in the passenger's direction to determine the first residual and the second residual. The first estimated noise residual is determined based on the first noise and the first residual. The second estimated noise residual is determined based on the second noise and the second residual. The difference between the first speech in the driver's direction and the first estimated noise residual is used as the second speech in the driver's direction. The difference between the first speech in the passenger's direction and the second estimated noise residual is also used as the second speech in the passenger's direction. This solves the problem of speech separation in the driver and passenger directions, achieves optimal separation of speech in the driver and passenger areas, considers the concealment of sound, has strong ability to suppress speech residuals, and has almost no speech residuals or noise interference.

[0067] Figure 3 This is a flowchart of an in-vehicle voice processing method provided in Embodiment 2 of the present invention, as follows: Figure 3 As shown, in an in-vehicle environment, the voice in the driver and passenger directions can be collected using a microphone array, and the first voice in the driver direction and the first voice in the passenger direction can be extracted using a beamforming algorithm. Specifically, the first voice y1(t, f) in the driver direction includes the voice s1(t, f) in the driver direction, the residual voice r1(t, f) in the passenger direction, and the first noise n1(t, f); the first voice y2(t, f) in the passenger direction includes the voice s2(t, f) in the passenger direction, the residual voice r2(t, f) in the driver direction, and the second noise n2(t, f).

[0068] Noise estimation can be performed on y1(t,f) and y2(t,f) obtained after beamforming to determine n1(t,f) and n2(t,f). At this point, sound masking can be considered to suppress noise and residual noise; specifically, noise reduction processing can be implemented. The maximum value N1(t,f) determined by n1(t,f) and y2(t,f)-n2(t,f) can be used as the overall noise and residual suppression during noise reduction processing. Similarly, the maximum value N2(t,f) determined by n2(t,f) and y1(t,f)-n1(t,f) can be used as the overall noise and residual suppression during noise reduction processing. For example, the second voice in the driver's direction and the second voice in the passenger's direction can be obtained through noise reduction processing using y1(t,f)-N1(t,f) and y2(t,f)-N2(t,f), respectively. The solution in this embodiment can effectively suppress the residual passenger voice from the driver's side and the residual driver's voice from the passenger's side, so that the driver's voice and passenger voice are well represented separately in the vehicle environment, without recognizing interfering sounds, and making it easy to accurately identify and process the driver's voice and passenger voice separately.

[0069] Example 3

[0070] Figure 4 This is a structural schematic diagram of an in-vehicle voice processing device provided in Embodiment 3 of the present invention. (Combined with...) Figure 4 The device includes: a voice acquisition module 410, a noise residual estimation module 420, and a noise and residual suppression module 430.

[0071] The voice acquisition module 410 is used to acquire voice in the driver and passenger directions, and extract the first voice in the driver direction and the first voice in the passenger direction from the voice in the driver and passenger directions.

[0072] The noise residual estimation module 420 is used to estimate the noise residual of the first voice in the driver's direction and the first voice in the passenger's direction respectively, and to determine the first estimated noise residual and the second estimated noise residual.

[0073] The noise and residual suppression module 430 is used to perform noise and residual suppression on the first speech in the driver's direction and the first speech in the passenger's direction based on the first estimated noise residual and the second estimated noise residual, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

[0074] Optionally, the voice acquisition module 410 includes:

[0075] The voice acquisition unit is used to extract the first voice from the driver's direction and the first voice from the passenger's direction from the voice in the driver's and passenger's directions using a beamforming algorithm.

[0076] Optionally, the first voice in the driver's direction includes the voice in the driver's direction, the voice residue in the passenger's direction, and a first noise; the first voice in the passenger's direction includes the voice in the passenger's direction, the voice residue in the driver's direction, and a second noise.

[0077] Optionally, the noise residual estimation module 420 includes:

[0078] The noise estimation unit is used to estimate the noise of the first speech in the driver's direction and the first speech in the passenger's direction, respectively, and to determine the first noise and the second noise.

[0079] The residual estimation unit is used to perform residual estimation on the first voice in the driver's direction and the first voice in the passenger's direction, respectively, and to determine the first residual and the second residual.

[0080] The first estimated noise residual determination unit is used to determine the first estimated noise residual based on the first noise and the first residual.

[0081] The second estimated noise residue determination unit is used to determine the second estimated noise residue based on the second noise and the second residue.

[0082] Optionally, the first residue is the difference between the first voice and the second noise in the direction of the passenger; the second residue is the difference between the first voice and the second noise in the direction of the driver.

[0083] Optionally, the first estimated noise residual determination unit is specifically used for:

[0084] The maximum values ​​of each frequency band of the first noise and the first residual are determined as the first estimated noise residual.

[0085] Optionally, the second estimated noise residual determination unit is specifically used for:

[0086] The maximum values ​​of each frequency band of the second noise and the second residual are determined as the first estimated noise residual.

[0087] Optionally, the noise and residual suppression module 430 includes:

[0088] The second voice acquisition unit in the driver's direction is used to take the difference between the first voice in the driver's direction and the first estimated noise residue as the second voice in the driver's direction.

[0089] The second voice acquisition unit in the passenger direction is used to take the difference between the first voice in the passenger direction and the second estimated noise residue as the second voice in the passenger direction.

[0090] The in-vehicle voice processing device provided in the embodiments of the present invention can execute the in-vehicle voice processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0091] Example 4

[0092] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention, as shown below. Figure 5 As shown, the device includes:

[0093] One or more processors 510, Figure 5 Take the 510 processor as an example;

[0094] Memory 520;

[0095] The device may also include an input device 530 and an output device 540.

[0096] The processor 510, memory 520, input device 530, and output device 540 in the device can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0097] Memory 520, as a non-transitory computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to an in-vehicle voice processing method in this embodiment of the invention (e.g., attached...). Figure 4 The diagram shows a voice acquisition module 410, a noise residual estimation module 420, and a noise and residual suppression module 430. The processor 510 executes various functional applications and data processing of the computer device by running software programs, instructions, and modules stored in the memory 520, thereby implementing an in-vehicle voice processing method according to the above method embodiment.

[0098] Acquire the voice from the driver and passenger directions, and extract the first voice from the driver's direction and the first voice from the passenger's direction from the voice from the driver and passenger directions;

[0099] Noise residual estimation is performed on the first voice in the driver's direction and the first voice in the passenger's direction respectively to determine the first estimated noise residual and the second estimated noise residual.

[0100] Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

[0101] The memory 520 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 520 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 520 may optionally include memory remotely located relative to the processor 510, and these remote memories can be connected to the terminal device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0102] Input device 530 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the computer device. Output device 540 may include display devices such as a display screen.

[0103] Example 5

[0104] Embodiment 5 of the present invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an in-vehicle voice processing method as provided in this embodiment of the present invention:

[0105] Acquire the voice from the driver and passenger directions, and extract the first voice from the driver's direction and the first voice from the passenger's direction from the voice from the driver and passenger directions;

[0106] Noise residual estimation is performed on the first voice in the driver's direction and the first voice in the passenger's direction respectively to determine the first estimated noise residual and the second estimated noise residual.

[0107] Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

[0108] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0109] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0110] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0111] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0112] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A vehicle-mounted voice processing method, characterized in that, include: Acquire the voice from the driver and passenger directions, and extract the first voice from the driver's direction and the first voice from the passenger's direction from the voice from the driver and passenger directions; Noise residual estimation is performed on the first voice in the driver's direction and the first voice in the passenger's direction respectively to determine the first estimated noise residual and the second estimated noise residual. The noise residual estimation includes residual estimates determined by considering the influence coefficient of the passenger's voice on the driver's voice or the influence coefficient of the driver's voice on the passenger's voice. Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction.

2. The method according to claim 1, characterized in that, Extracting the first voice from the driver's direction and the first voice from the passenger's direction from the voice in the driver's and passenger's directions includes: The beamforming algorithm is used to extract the first voice from the driver's direction and the first voice from the passenger's direction from the driver's and passenger's directions.

3. The method according to claim 1, characterized in that, The first voice in the driver's direction includes the voice in the driver's direction, the voice residue in the passenger's direction, and a first noise; the first voice in the passenger's direction includes the voice in the passenger's direction, the voice residue in the driver's direction, and a second noise.

4. The method according to claim 3, characterized in that, Noise residue estimation is performed on the first voice from the driver's side and the first voice from the passenger's side, respectively, to determine the first estimated noise residue and the second estimated noise residue, including: Noise estimation is performed on the first voice message from the driver's side and the first voice message from the passenger's side, respectively, to determine the first noise and the second noise. Residual estimations are performed on the first voice from the driver's direction and the first voice from the passenger's direction, respectively, to determine the first residual and the second residual. The first estimated noise residue is determined based on the first noise and the first residue; The second estimated noise residue is determined based on the second noise and the second residue.

5. The method according to claim 4, characterized in that, The first residual is the difference between the first voice and the second noise in the direction of the passenger; the second residual is the difference between the first voice and the second noise in the direction of the driver.

6. The method according to claim 4, characterized in that, Determining the first estimated noise residue based on the first noise and the first residue includes: The maximum value of each frequency band of the first noise and the first residual is determined as the first estimated noise residual; Determining the second estimated noise residue based on the second noise and the second residue includes: The maximum value of each frequency band of the second noise and the second residual is determined as the first estimated noise residual.

7. The method according to any one of claims 1-6, characterized in that, Based on the first estimated noise residue and the second estimated noise residue, noise and residue suppression are performed on the first speech in the driver's direction and the first speech in the passenger's direction, respectively, to obtain the second speech in the driver's direction and the second speech in the passenger's direction, including: The difference between the first voice in the driver's direction and the first estimated noise residual is taken as the second voice in the driver's direction; The difference between the first voice in the co-driver's direction and the second estimated noise residue is taken as the second voice in the co-driver's direction.

8. A vehicle-mounted voice processing device, characterized in that, include: The voice acquisition module is used to acquire voice from the driver and passenger directions, and extract the first voice from the driver direction and the first voice from the passenger direction from the driver and passenger directions. The noise residual estimation module is used to perform noise residual estimation on the first voice in the driver's direction and the first voice in the passenger's direction respectively, and determine the first estimated noise residual and the second estimated noise residual. The noise residual estimation includes residual estimates determined by considering the influence coefficient of the passenger's voice on the driver's voice or the influence coefficient of the driver's voice on the passenger's voice. The noise and residual suppression module is used to perform noise and residual suppression on the first voice in the driver's direction and the first voice in the passenger's direction based on the first estimated noise residual and the second estimated noise residual, respectively, to obtain the second voice in the driver's direction and the second voice in the passenger's direction.

9. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Vehicle-mounted multi-region-of-articulation interaction system and method

    CN109754803A

  • Vehicle-mounted voice recognition method and system

    CN110459234A