Speech processing system, speech processing method, medium, product and device
The sound source orientation is determined through the voice receiving module and the beam adjustment module, and the converging beam of the speaker array plays sound to the sound source, solving the interference problem of speaker sound on other personnel, realizing directional propagation and signal quality improvement.
Patent Information
- Application Number
- CN202410634706.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-05-21
AI Technical Summary
In scenes such as calls, meetings, navigation, video playback and square activities, the sound played by the speakers has a great interference to other people.
The sound source orientation is determined through the voice receiving module, and the beam adjustment module is used to control the converging beam of the speaker array to play the sound signal to the sound source orientation. The voice signal is picked up by a microphone array, and the signal quality is improved through the sound source separation module and echo cancellation processing modules.
It realizes directional propagation of sound, reduces interference to other personnel, and improves the clarity and quality of voice signals.
Smart Images

Figure CN118509765B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of multimedia technologies, and particularly to a voice processing system, a voice processing method, a medium, a product, and a device. Background Art
[0002] In related technologies, in scenarios such as calls, meetings, navigation, video playback, and square activities where speakers are used to play sounds, some people expect to receive relatively loud playback sounds, but these sounds may cause significant interference to the rest of the people. Summary of the Invention
[0003] To overcome the problems existing in related technologies, the present disclosure provides a voice processing system, a voice processing method, a medium, a product, and a device.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a voice processing system, where the voice processing system includes:
[0005] A voice receiving module, including a direction determination module and a microphone array, where the microphone array is configured to acquire voice signals in a target space, and the direction determination module is configured to determine the sound source direction of the sound source corresponding to the voice signals according to the voice signals;
[0006] A sound playback module, including a beam adjustment module and a speaker array for playing a to-be-played sound signal, where the beam adjustment module is configured to determine a converging beam in which the speaker array points to the sound source direction according to the sound source direction, so as to play the to-be-played sound signal to the sound source direction through the converging beam, and the to-be-played sound signal is sent by an object that performs a voice interaction with the sound source.
[0007] Optionally, the beam adjustment module is further configured to determine a converging beam pointing to the sound source direction based on equal sidelobe beamforming according to the sound source direction, so that the main lobe of the power radiation pattern of the speaker array points to the sound source direction, and the sidelobe signals corresponding to other sound source directions except the sound source direction are recessed.
[0008] Optionally, the direction determination module is further configured to determine a pickup beam in which the microphone array points to the sound source direction according to the sound source direction of the voice signals, so as to pick up the voice signals corresponding to the sound source direction through the pickup beam.
[0009] Optionally, when there are multiple sound sources in the target space, the voice processing system further includes:
[0010] A sound source separation module, configured to determine a pickup beam pointing to the sound source direction of each sound source among a plurality of sound sources, and determine a target sound source currently generating a sound signal from the plurality of sound sources;
[0011] The voice receiving module is further configured to determine a target pickup beam pointing to the sound source direction corresponding to the target sound source, so as to pick up the voice signal of the target sound source through the target pickup beam.
[0012] Optionally, for the pickup beam of each sound source, the main lobe signal of the pickup beam covers the sound source direction where the sound source is located, and the sidelobe signals corresponding to the sound source directions of the other sound sources except the sound source in the pickup beam are suppressed.
[0013] Optionally, the direction determination module is further configured to determine at least one activated sound area position among the plurality of sound areas through the activation states of the plurality of sound areas in the target space, and perform azimuth estimation based on the voice signal, so as to determine the sound source direction of the sound source corresponding to the voice signal within the at least one activated sound area position.
[0014] Optionally, the voice receiving module further includes at least one of the following modules:
[0015] An echo cancellation module, configured to perform echo cancellation processing on the voice signal picked up by the microphone array;
[0016] A non-linear post-processing module, configured to perform non-linear post-processing on the voice signal processed by the echo cancellation module;
[0017] A voiceprint enhancement module, configured to perform voiceprint enhancement processing and voice noise reduction processing on the voice signal according to the voiceprint information of the sound source;
[0018] A noise suppression module, configured to perform noise suppression processing on the voice signal;
[0019] An automatic gain module, configured to perform automatic gain processing on the voice signal so that the gain of the voice signal is within a first preset range;
[0020] An equalization module, configured to perform signal equalization processing on the voice signal.
[0021] Optionally, the sound playback module further includes at least one of the following modules:
[0022] A noise suppression module, configured to perform noise suppression processing on the sound signal to be played;
[0023] An automatic gain module, configured to perform automatic gain processing on the to-be-played sound signal so that the gain of the to-be-played sound signal is within a second preset range;
[0024] An equalization module, configured to perform signal equalization processing on the to-be-played sound signal.
[0025] Optionally, the target space is inside a vehicle cockpit.
[0026] According to a second aspect of the embodiments of the present disclosure, there is provided a voice processing method, including:
[0027] Obtaining a voice signal in a target space, and determining a sound source azimuth of a sound source corresponding to the voice signal in the target space according to the voice signal;
[0028] Determining a converging beam pointing to the sound source azimuth according to the sound source azimuth, so as to play a to-be-played sound signal to the sound source azimuth through the converging beam, where the to-be-played sound signal is sent by an object for voice interaction with the sound source.
[0029] According to a third aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the voice processing method provided in the second aspect of the present disclosure is implemented.
[0030] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the voice processing method provided in the second aspect of the present disclosure is implemented.
[0031] According to a fifth aspect of the embodiments of the present disclosure, there is provided a multimedia device, including the voice processing system provided in the first aspect of the present disclosure.
[0032] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0033] In the voice processing system of the present disclosure, by setting a voice receiving module to determine the sound source azimuth of the sound source generating the voice signal, and by a beam adjustment module, according to the sound source azimuth, determining a converging beam of the speaker array pointing to the sound source azimuth, so as to play a to-be-played sound signal to the sound source azimuth through the converging beam. In this way, according to the sound source azimuth of the sound source with the need to receive sound data, the converging beam of the speaker array can be controlled to point to the sound source azimuth, realizing directional sound propagation to some personnel, thereby reducing the interference degree to other personnel.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0035] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0036] Figure 1 is a block diagram of a voice processing system shown according to an exemplary embodiment.
[0037] Figure 2 is a schematic diagram of a convergent beam shown according to an exemplary embodiment.
[0038] Figure 3 is a schematic diagram of a pickup beam shown according to an exemplary embodiment.
[0039] Figure 4 is a schematic diagram of another pickup beam shown according to an exemplary embodiment.
[0040] Figure 5 is a schematic diagram of dividing a target space into multiple sound zones shown according to an exemplary embodiment.
[0041] Figure 6 is a block diagram of another voice processing system shown according to an exemplary embodiment.
[0042] Figure 7 is a flowchart of a voice processing method shown according to an exemplary embodiment.
[0043] Figure 8 is a flowchart of a voice signal processing method shown according to an exemplary embodiment.
[0044] Figure 9 is a flowchart of a method for processing a sound signal to be played shown according to an exemplary embodiment.
[0045] Figure 10 is a block diagram of a vehicle shown according to an exemplary embodiment. Detailed Description of Specific Embodiments
[0046] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0047] In some scenarios where sound is played through a speaker during a call, a meeting, navigation, video playback, a square activity, etc., for example, a user may wish to make a voice call or hold a video conference during a vehicle journey to improve work efficiency and travel efficiency. For example, a user may wish to engage in entertainment activities such as watching video programs during the journey. For example, a user may wish that during navigation, only the driver's seat receives a relatively large navigation announcement sound. In some other scenarios, some users wish to hold group activities in a square through a speaker, and so on. Some people expect to receive a relatively large playback sound, but these sounds may cause significant interference to the rest of the people in the same space.
[0048] See Figure 1 , Figure 1 is a block diagram of a voice processing system shown according to an exemplary embodiment. As Figure 1 shown, the voice processing system includes:
[0049] A voice receiving module, including a direction determination module and a microphone array. The microphone array is configured to acquire voice signals in a target space, and the direction determination module is configured to determine the sound source direction of the sound source corresponding to the voice signal according to the voice signal;
[0050] A sound playback module, including a beam adjustment module and a speaker array for playing a sound signal to be played. The beam adjustment module is configured to determine a converging beam in which the speaker array points to the sound source direction according to the sound source direction, so as to play the sound signal to be played to the sound source direction through the converging beam. The sound signal to be played is sent by an object that conducts voice interaction with the sound source.
[0051] Exemplarily, the microphone array includes a plurality of microphones. A microphone is an energy conversion device that can pick up voice signals generated by a sound source and convert the voice signals into electrical signals. The target space is any space, such as within the cockpit of a vehicle, or the area where an audio device is located, etc. Among them, the target space can be an enclosed space, a semi-enclosed space or an unobstructed space. An enclosed space is, for example, within the cockpit of a vehicle, a house, etc. A semi-enclosed space is, for example, a balcony, a square enclosed by a fence, etc. An unobstructed space is, for example, an open square, etc. Here, the type and size range of the target space are not limited.
[0052] Exemplarily, the sound signal to be played is sent by an object that conducts voice interaction with the sound source. For example, in scenarios such as a voice call scenario or a video conference scenario, the sound signal to be played is emitted by the other party in the voice call or the other party in the video conference with the device user corresponding to this voice processing system, and is collected, processed, and transmitted through the device used by the other party in the voice call or the other party in the video conference. In this case, it can be considered that the object that conducts voice interaction with the person or device of the local device is the person or device of the other party in the voice call or the other party in the video conference, that is, a person-to-person interaction or device-to-device interaction scenario. Another example is that in scenarios such as navigation, video playback, and square activities, the sound signal to be played is the navigation software, video playback program, audio playback program, etc. in the device corresponding to this voice processing system. In this case, it can be considered that the interaction object with the user of the local device is the local device or the application program in the local device, that is, a human-machine interaction scenario.
[0053] Exemplarily, the azimuth determination module can determine the sound source azimuth of the sound source corresponding to the voice signal according to the voice signal. Specifically, the azimuth estimation algorithm can adopt traditional algorithms such as Conventional Beamforming (CBF), Minimum Variance Distortionless Response (MVDR), Multiple Signal Classification (MUSIC), Compressed Sensing (CS), neural networks, etc. The beam adjustment module can determine the converging beam of the speaker array pointing to the sound source azimuth according to the sound source azimuth. In this way, the sound signal to be played can be played towards the sound source azimuth through the converging beam, so that the sound signal to be played points to the sound source, enabling the sound source to receive the sound data to be played directionally.
[0054] As an exemplary scenario, in the traditional in-vehicle hands-free call scenario, the driver connects the mobile phone to the in-vehicle Bluetooth to achieve hands-free calls, without the call zoning function, that is, without the ability to suppress interference from the co-driver and other passengers. Through the voice processing system of the present disclosure, multiple microphones are arranged in the cockpit to form a microphone array, and multiple speakers are arranged at the same time to form a speaker array. That is, it is possible to separately extract the voices of each passenger to determine the target passenger of the current call, and then use the speaker array to perform beam focusing to direct the voice resources to the target passenger, which can reduce the impact on other passengers.
[0055] In the voice processing system of the present disclosure, a voice receiving module is set to determine the sound source direction of the sound source that generates the voice signal, and a beam adjustment module determines a convergent beam in which the speaker array points to the sound source direction according to the sound source direction, so as to play the sound signal to be played in the sound source direction through the convergent beam. In this way, according to the sound source direction of the sound source that has the need to receive sound data, the convergent beam of the speaker array can be controlled to point to the sound source direction, realizing the directional propagation of sound to some personnel, thereby reducing the interference degree to other personnel.
[0056] As an optional embodiment, the beam adjustment module is further configured to determine a convergent beam pointing to the sound source direction based on equal sidelobe beamforming according to the sound source direction, so that the main lobe of the power radiation pattern of the speaker array points to the sound source direction, and the sidelobe signals corresponding to other sound source directions except the sound source direction are recessed.
[0057] Exemplarily, when the main lobe of the power radiation pattern of the speaker array points to the sound source direction for receiving the sound signal to be played, a recess can be presented in other sound source directions, which can greatly reduce the energy diffusion of the sound signal to be played in other sound source directions and further reduce the influence on other personnel. As Figure 2 shown in the convergent beam of the speaker array, the angular range occupied by the main lobe of the beam is from -20 degrees to 20 degrees to face the sound source direction. There is a recess in the range of 40 degrees to 60 degrees, and the sound signal energy that can be received by the personnel within the angular range of the recess is small, that is, the playback volume within this range is small.
[0058] Exemplarily, the power radiation pattern of the speaker can be adjusted by using conventional beam convergence technology. Here, the equal sidelobe beamforming method is adopted. Among them, it is assumed that there is a single-channel audio in each sound zone For the single-channel audio In the scenario of beam convergence and M speakers are set in the speaker array, the weighting and phase adjustment are performed on each speaker respectively to obtain , represents the weighting value of M speakers at a certain frequency point, generally a complex number, represents a complex number column vector with a total of M elements. The Fourier transform is performed on the audio to obtain , and then the calculation is performed , where . Each row of data of X is sent to different speakers in the speaker array respectively, so that different speakers can realize the mutual cancellation or in-phase enhancement of the diverse-played audio in space, thereby realizing the effect of beam convergence. Further, a convex optimization method can be adopted for the weighting value design, and the designed calculation formulas are as shown in formulas (1) and (2).
[0059] Formula (1): ;
[0060] Equation (2): s.t. , , , , .
[0061] Among them, 、 respectively represent the direction vectors of the speaker with respect to angles and ; is generally a relatively small number, such as 0.01; represents the constraint on the direction , for example, the signal intensity in the sidelobe angle range can be constrained to -30 dB; represents the set of sidelobe angles; W H represents the conjugate transpose of W; Considering that there is a certain angular interval between the two ears of a person and the estimation of the head position of a person may not be accurate, represents that the angular coverage allowing the main lobe to be distortion-free covers a certain range; represents the set of angles in the key suppression area. For example, when the driver needs to answer a voice broadcast and the passenger does not want to hear the voice broadcast, the area where the passenger is located is the key suppression area. For example, the signal intensity in the key suppression area can be suppressed to -60 dB.
[0062] As an optional embodiment, the azimuth determination module is further configured to determine a pickup beam of the microphone array pointing to the sound source azimuth according to the sound source azimuth of the voice signal, so as to pick up the voice signal corresponding to the sound source azimuth through the pickup beam.
[0063] Exemplarily, one or more microphone arrays can collect the voice signals of all sound sources in the target space, or can control the pickup beam of the microphone array to point to the sound source according to the sound source azimuth determined by the azimuth determination module, so as to specify the collection of the voice signal of the sound source, which can reduce the interference of the sound at other positions on the voice signal of the sound source.
[0064] As an optional embodiment, when there are multiple sound sources in the target space, the voice processing system further includes:
[0065] A sound source separation module, configured to determine a pickup beam pointing to each sound source azimuth according to the sound source azimuth of each sound source among the multiple sound sources, and determine a target sound source currently generating a sound signal from the multiple sound sources;
[0066] A voice receiving module, further configured to determine a target pickup beam pointing to the sound source azimuth corresponding to the target sound source, so as to pick up the voice signal of the target sound source through the target pickup beam.
[0067] Exemplarily, the main lobe in the target pickup beam points to the sound source azimuth corresponding to the target sound source that currently generates the sound signal. For a video conferencing scenario in a vehicle cockpit, if there is only one participant in the vehicle, the main beam is pointed to that participant; if there are multiple participants, multiple pickup beams can be designed to respectively extract the voices of different speaking participants, so as to directionally extract the voice signals of the speaking participants and filter out the remaining noise.
[0068] For example, in a scenario where both the driver and the passenger in the vehicle cockpit participate in the meeting, respectively determine the sound source azimuth of the speaker in the driver's position and the sound source azimuth of the speaker in the passenger's position, and then, based on the sound source azimuth in the driver's position and the sound source azimuth in the passenger's position, form a pickup beam pointing to the corresponding sound source azimuth based on adaptive beamforming. Among them, the main lobe of the pickup beam in the driver's position points to the sound source azimuth in the driver's position, and the side lobes of the pickup beam in the driver's position are noise and the voice signal in the passenger's position, so that the pickup beam pointing to the driver's position only extracts the voice signal in the driver's position. The main lobe of the pickup beam in the passenger's position points to the sound source azimuth in the passenger's position, and the side lobes of the pickup beam in the passenger's position are noise and the voice signal in the driver's position, so that the pickup beam pointing to the passenger's position only extracts the voice signal in the passenger's position. And these two extraction processes can be processed simultaneously through multiple threads to improve the pickup efficiency. In this way, when the driver is speaking, the voice signal in the driver's position can be extracted through the pickup beam in the driver's position, and when the passenger is speaking, it can be switched to extract the voice signal in the passenger's position through the pickup beam in the passenger's position.
[0069] As an optional embodiment, for the pickup beam of each sound source, the main lobe signal of the pickup beam covers the sound source azimuth where the sound source is located, and the side lobe signals corresponding to the sound source azimuths of the other sound sources except the sound source in the pickup beam are suppressed.
[0070] Exemplarily, according to the sound source azimuth, a pickup beam can be formed based on low side lobe beamforming, where the main lobe of the pickup beam points to the sound source azimuth, and the side lobes of the pickup beam are suppressed as noise, as Figure 3 shown, the main lobe of the pickup beam points to an azimuth between approximately -15 degrees and 15 degrees, and the side lobe signal intensity is suppressed between -30 dB and -60 dB. Or, in the case where there are multiple sound sources, pickup beams pointing to the corresponding sound source azimuths can be formed based on adaptive beamforming according to the sound source azimuth of each voice signal, where the side lobe signals corresponding to the sound source azimuths other than the sound source azimuth corresponding to the target sound source are suppressed as interference, as Figure 4As shown, in the beam pattern based on adaptive beamforming, the main lobe of the pickup beam points to an azimuth between approximately -10 degrees and 10 degrees. Most of the sidelobe signal intensities are suppressed between -30 dB and -60 dB, and there is another sound source at the 30-degree direction, and the sidelobe signal intensity corresponding to the 30-degree direction is suppressed to be less than -100 dB.
[0071] As an alternative embodiment, the azimuth determination module is further configured to determine at least one activated sound zone position among multiple sound zones in the target space through the activation states of the multiple sound zones in the target space, and perform azimuth estimation based on the voice signal to determine the sound source azimuth of the sound source corresponding to the voice signal within at least one sound zone position.
[0072] Exemplarily, sound zones can be divided in the target space, and the sound source azimuth of the voice signal in the activated sound zone position can be identified. For example, according to the distribution of the vehicle cockpit positions, the vehicle cockpit space can be evenly divided into multiple sound zones, as Figure 5 shown, the vehicle cockpit is divided into 4 sound zones. When a passenger rides in the vehicle, the passenger's sitting posture may not be at the center position of the sound zone position. For example, the passenger may tilt the body to make a call. Therefore, the sound source azimuth can be estimated within the angular coverage range of the activated sound zone position to determine the specific angle of the speaker relative to the center of the sound zone position. For example, the angular azimuth of the angular coverage range of the sound zone position is approximately 30 degrees to 60 degrees.
[0073] Exemplarily, when performing azimuth estimation on the voice signal, the sound source azimuth can be estimated within the angular coverage range of the activated sound zone position to determine the specific angle of the speaker at the center of the sound zone position, that is, the sound source azimuth of the sound source corresponding to the voice signal is obtained.
[0074] Exemplarily, in the related art, to determine at least one activated sound zone position among multiple sound zones and distinguish the target sound zone according to the wake-up word, in one way, a multi-sound zone voice interaction system can be used, but there are difficulties in mis-waking up the sound zone. In another way, a camera can be used to estimate the position of the speaker according to the lip movement, but when the lips are blocked, the speaker cannot be determined, and the camera solution has a high cost and high computational complexity. In yet another way, a pressure sensor can be installed under the seat, but if a heavy object is placed, it is impossible to determine whether the sound zone is an item or a passenger.
[0075] Exemplarily, in the vehicle cockpit scenario, for a vehicle equipped with Bluetooth for each seat, when a Bluetooth call is detected, according to the Bluetooth identification number of the target vehicle Bluetooth for which the Bluetooth call has been established, the target sound zone corresponding to the Bluetooth identification number can be determined, and further the sound source azimuth of the speaker in the activated sound zone position can be determined. An indicator light can also be set. When it is detected that a passenger in the sound zone position is speaking, the indicator light corresponding to the sound zone position will light up, indicating that it is detected that the passenger in this sound zone is speaking.
[0076] Exemplarily, as Figure 5 shown in the sound zone partition form, in the target space, only one microphone array and one speaker array can be set, or one microphone array can be set, and one speaker array can be set in each sound zone. It can also be set in other forms. When controlling the focusing beam, when only one speaker array is set in the target space, the pointing direction of the focusing beam of the speaker array can be controlled by controlling the audio played by each speaker of the speaker array. When one speaker array is set in each sound zone in the target space, the pointing direction of the focusing beam of the entire speaker array in the target space can be controlled by only controlling the audio played by the speaker array in the activated target sound zone, or by controlling the audio played by the speaker arrays in all sound zones.
[0077] Exemplarily, in order to accurately determine the position of the activated sound zone in the vehicle and avoid the situation of misawakening of the sound zone position, in one implementation, for the Bluetooth system of the vehicle in the related art, when a Bluetooth call is detected, the central control panel of the vehicle can be controlled to display the sound zone mode of the position of the sound zone that can be activated, where the sound zone mode includes the full vehicle sound zone mode and the local sound zone mode. Based on the full vehicle sound zone mode selected by the user on the central control panel, all seat positions of the vehicle are determined as the positions of the activated sound zones, and based on the seat positions selected by the user in the local sound zone mode, the corresponding positions of the activated sound zones are determined.
[0078] As an alternative embodiment, the voice receiving module further includes at least one of the following modules:
[0079] An echo cancellation module configured to perform echo cancellation processing on the voice signal picked up by the microphone array;
[0080] A non-linear post-processing module configured to perform non-linear post-processing on the voice signal processed by the echo cancellation module;
[0081] A voiceprint enhancement module configured to perform voiceprint enhancement processing and voice noise reduction processing on the voice signal according to the voiceprint information of the sound source;
[0082] A noise suppression module configured to perform noise suppression processing on the voice signal;
[0083] An automatic gain module configured to perform automatic gain processing on the voice signal so that the gain of the voice signal is within a first preset range;
[0084] An equalization module configured to perform signal equalization processing on the voice signal.
[0085] Exemplarily, due to the relatively strong reflection ability of parts such as the car window and door, the voice signal will produce echoes. Therefore, in order to improve the call quality of in-vehicle Bluetooth calls, before determining the sound source direction of the voice signal, the present disclosure performs echo cancellation on the voice signals received by each microphone array to avoid the multipath effect. For example, when performing acoustic echo cancellation (AEC) on the voice signals obtained by the microphone array, the echoes reflected through the car window, door, etc. can be removed through an echo cancellation algorithm. Among them, the echo cancellation algorithm can include, for example, the least mean square (LMS), normalized LMS (NLMS), recursive least square (RLS), partition block frequency domain adaptive filter (PBFDAF), frequency domain NLMS (FDNLMS), and other algorithms.
[0086] Exemplarily, due to the distortion of the speaker and the non-ideality of propagation, the echo cancellation module generally cannot perfectly solve the non-linear problem. Generally, a non-linear post-processing module is added to extract the useful information of the signal and calculate the scaling factor of the corresponding frequency points to further suppress the error signal in the output error signal.
[0087] Exemplarily, the user can choose to perform voiceprint registration during a call or a video conference. The voiceprint enhancement module uses voice enhancement technology to perform voice enhancement and noise reduction according to the user's voiceprint information.
[0088] Exemplarily, since beamforming cannot completely suppress the noise outside the beam direction, and noise will be mixed into the main lobe of the beam, the voice signal in the extracted active sound area will have residual noise. The noise suppression module can use post-processing filtering and other technologies to continue filtering the extracted voice signal to improve the signal-to-noise ratio and thus improve the call clarity.
[0089] Exemplarily, the automatic gain control module can be used for automatic gain control (AGC). The purpose is to control the gain of the voice signal within a preset reasonable range to avoid sudden changes in sound intensity. Among them, the first preset range can be set according to actual needs.
[0090] Exemplarily, the equalization module can be used to equalize (Equaliser, EQ) or perform dynamic range control (Dynamic Range Control, DRC) on the voice signal. It is a signal amplitude adjustment method that compensates for the defects of the speaker and the sound field by adjusting electrical signals of various different frequencies, compensates and modifies various sound sources and other special effects, and can make the sound sound softer or louder.
[0091] It can be understood that in practical applications, one or more of the echo cancellation module, non-linear post-processing module, voiceprint enhancement module, noise suppression module, automatic gain module, and equalization module can be used to process the voice signal to enhance the signal quality of the voice signal.
[0092] As an alternative embodiment, the sound playback module further includes at least one of the following modules:
[0093] A noise suppression module, configured to perform noise suppression processing on the sound signal to be played;
[0094] An automatic gain module, configured to perform automatic gain processing on the sound signal to be played so that the gain of the sound signal to be played is within a second preset range;
[0095] An equalization module, configured to perform signal equalization processing on the sound signal to be played.
[0096] Exemplarily, in video conferencing and voice call scenarios, since there may be residual noise and circuit background noise in the voices of the meeting counterpart and the call counterpart, noise suppression can be performed on the voice before it is played by the speaker. Specifically, traditional noise suppression algorithms such as Wiener filtering, spectral subtraction, and signal subspace filtering can be used for noise suppression, or neural network algorithms can also be used for noise suppression.
[0097] Exemplarily, the automatic gain control module can be used for automatic gain control (Auto Gain Control, AGC). The purpose is to control the gain of the sound signal to be played within a preset reasonable range to avoid sudden changes in sound intensity. Among them, the second preset range can be set according to actual needs.
[0098] Exemplarily, the equalization module can be used to equalize (Equaliser, EQ) or perform dynamic range control (Dynamic Range Control, DRC) on the voice signal to be played. It is a signal amplitude adjustment method that compensates for the defects of the speaker and the sound field by adjusting electrical signals of various different frequencies, compensates and modifies various sound sources and other special effects, and can make the sound sound softer or louder.
[0099] It can be understood that in practical applications, one or more of an echo cancellation module, a non-linear post-processing module, a voiceprint enhancement module, a noise suppression module, an automatic gain module, and an equalization module can be used to process the voice signal to enhance the signal quality of the sound signal to be played.
[0100] As an alternative embodiment, the target space is a vehicle cockpit.
[0101] As a specific example, the processing flow of the voice processing system of the present disclosure is exemplified in a conference scenario, and reference can be made to Figure 6 , where the microphone array collects the voice signals of the present participants. After the echo cancellation module performs echo cancellation processing on the voice signals, the azimuth determination module estimates the sound source azimuth of the sound source corresponding to the voice signals. The sound source separation module determines the current pickup beam according to the sound source azimuth. For example, if the rear passengers start the conference at the same time, two beams are designed. When the main lobe of the beam points to the left passenger, the signals of the right passenger are suppressed, so as to extract the voice signals of the left passenger; when the main lobe of the beam points to the right passenger during driving, the voice signals of the left passenger are suppressed, so as to extract the voice signals of the right passenger. After the sound source separation module processes, the noise suppression module filters the voice signals, and the non-linear post-processing module processes the voice signals to extract the useful signals in the voice signals. The automatic gain module performs automatic gain control on the voice signals, and the equalization module performs equalization processing on the voice signals to adjust the amplitude of the voice signals. Finally, the voice signals processed by the equalization module are transmitted to the other party of the conference through the communication system. The other party of the conference here can be a person or a device.
[0102] In the above scenario, during the process of transmitting the voice signals of the other party of the conference to the present participants, the voice signals of the other party of the conference are obtained through the participating device of the other party of the conference as the sound signals to be played, and are transmitted to the device end of the present participants through the communication system. The voice processing system of this device end performs noise suppression processing on the sound signals to be played through the noise suppression module, and the automatic gain module performs automatic gain control on the sound signals to be played. The equalization module performs equalization processing on the sound signals to be played to adjust the amplitude of the sound signals to be played. The beam adjustment module controls the focusing beam of the speaker to point to the azimuth where the present participants are located according to the sound source azimuth determined by the azimuth determination module, that is, the azimuth of the present participants.
[0103] Referring to Figure 7 Figure 7 is a flowchart of a voice processing method shown according to an embodiment of the present disclosure. The voice processing method can be applied to a multimedia device and includes the following steps.
[0104] S701, obtain the voice signals in the target space, and determine the sound source azimuth of the sound source corresponding to the voice signals in the target space according to the voice signals;
[0105] S702. Determine a converging beam pointing to the sound source direction according to the sound source direction, and play a sound signal to be played in the sound source direction through the converging beam. The sound signal to be played is sent by an object for voice interaction with the sound source.
[0106] The present disclosure determines the sound source direction of the sound source corresponding to the voice signal in the target space through the voice signal in the target space, and determines a converging beam pointing to the sound source direction according to the sound source direction, so as to play the sound signal to be played in the sound source direction through the converging beam. In this way, according to the sound source direction of the sound source with the need to receive sound data, the converging beam of the speaker array can be controlled to point to the sound source direction, realizing the directional propagation of sound to some personnel, thereby reducing the interference degree to other personnel.
[0107] Exemplarily, according to the sound source direction, the pickup beam of the microphone array in the target space can be controlled to point to the sound source direction to realize the directional pickup of the voice signal of the sound source, and the converging beam of the speaker array in the target space can be controlled to point to the sound source direction to realize the directional playback of sound to the sound source direction.
[0108] Among them, as Figure 8 shown, in step S801, the voice signal in the target space can be picked up by using the microphone array; in step S802, according to the voice signal, the sound source direction of the sound source corresponding to the voice signal in the target space can be determined; in step S803, echo cancellation processing can be performed on the voice signal; in steps S804a and S804b, source separation processing and / or voiceprint enhancement processing can be performed on the voice data after echo cancellation processing. Through source separation processing, the voice data corresponding to multiple sound sources in the target space can be respectively extracted, and through voiceprint enhancement processing, the voice data of the sound source in the voice signal can be extracted according to the voiceprint information corresponding to the sound source; in step S805, noise suppression processing can be performed on the voice signal after source separation processing and / or voiceprint enhancement processing to filter out the noise in the voice signal; in step S806, non-linear post-processing can be performed on the voice data after noise suppression processing to extract the useful signal in the voice signal; in step S807, automatic gain control can be performed on the voice signal after non-linear post-processing to make the gain of the voice signal within a preset reasonable range, avoiding the voice signal from being too large or too small; in step S808, equalization processing can be performed on the voice data after automatic gain control to adjust the amplitude of the voice signal.
[0109] Among them, as Figure 9As shown, after obtaining the sound signal to be played, in step S901, noise suppression processing can be performed on the sound signal to be played to filter out the noise in the sound signal to be played; in step S902, automatic gain control can be performed on the sound signal to be played after noise suppression processing, so that the gain of the sound signal to be played is within a preset reasonable range, avoiding the sound signal from being too loud or too soft; in step S903, equalization processing can be performed on the sound signal to be played after automatic gain control to adjust the amplitude of the sound signal to be played; in step S904, it is determined whether the current way of playing the sound signal to be played is the speaker playing mode. If so, step S905 is executed. Otherwise, the processing flow ends, and the sound signal to be played is played through other playing modes; in step S905, the sound source orientation corresponding to the object to receive the sound signal to be played is determined, where the sound source orientation determined in step S802 can be directly used; in step S906, according to the sound source orientation, by beam focusing of the speaker, the beam of the speaker array is directed to the sound source orientation, so as to achieve the purpose of directional playing of sound to the sound source orientation.
[0110] The present disclosure also provides a non-transitory computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the voice processing method of the present disclosure is implemented.
[0111] The present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, the voice processing method of the present disclosure is implemented.
[0112] The present disclosure also provides a multimedia device, including the voice processing system of the present disclosure.
[0113] Exemplarily, the multimedia device is installed on a vehicle. For example, the multimedia device is the central control device of the vehicle. Or, the multimedia device is a mobile phone, a tablet computer, a sound device, etc. Installing the voice processing system of the present disclosure on the multimedia device can achieve the purpose of directional playing of the sound signal to be played by the speaker array based on the sound source orientation corresponding to the acquired voice signal, so that the sound signal intensity in other areas can be weakened, reducing the impact on other people.
[0114] Figure 10 It is a block diagram of a vehicle 1000 shown according to an exemplary embodiment. For example, the vehicle 1000 can be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 1000 can be an autonomous vehicle or a semi-autonomous vehicle.
[0115] Refer to Figure 10, vehicle 1000 may include various subsystems. For example, an infotainment system 1010, a perception system 1020, a decision control system 1030, a drive system 1040, and a computing platform 1050. Among them, vehicle 1000 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of vehicle 1000 may be interconnected in a wired or wireless manner.
[0116] In some embodiments, the infotainment system 1010 may include a communication system, an entertainment system, a navigation system, and the like.
[0117] The perception system 1020 may include several sensors for sensing information about the environment around vehicle 1000. For example, the perception system 1020 may include a global positioning system (the global positioning system may be a GPS system, a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device.
[0118] The decision control system 1030 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0119] The drive system 1040 may include components that provide motive power for vehicle 1000. In one embodiment, the drive system 1040 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine can convert the energy provided by the energy source into mechanical energy.
[0120] Some or all functions of vehicle 1000 are controlled by the computing platform 1050. The computing platform 1050 may include at least one processor 1051 and a memory 1052. The processor 1051 may execute instructions 1053 stored in the memory 1052.
[0121] The processor 1051 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0122] The memory 1052 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0123] In addition to the instructions 1053, the memory 1052 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 1052 can be used by the computing platform 1050.
[0124] In an embodiment of the present disclosure, the processor 1051 can execute the instructions 1053 to complete all or part of the steps of the above-described voice processing method.
[0125] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0126] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A voice processing system, characterized in that, The voice processing system includes: A voice receiving module, including a direction determination module and a microphone array. The microphone array is configured to acquire voice signals in a target space, and the direction determination module is configured to determine the sound source direction of the sound source corresponding to the voice signals according to the voice signals; A sound playback module, including a beam adjustment module and a speaker array for playing sound signals to be played. The beam adjustment module is configured to determine a convergent beam in which the speaker array points to the sound source direction according to the sound source direction, so as to play the sound signals to be played to the sound source direction through the convergent beam. The sound signals to be played are sent by an object that conducts voice interaction with the sound source; The beam adjustment module is further configured to determine a convergent beam pointing to the sound source direction based on equal sidelobe beamforming according to the sound source direction, so that the main lobe of the power radiation pattern of the speaker array points to the sound source direction, and the sidelobe signals corresponding to other sound source directions except the sound source direction are recessed, so as to reduce the energy of the sound signals to be played at the other sound source directions; In the equal sidelobe beamforming, the signal intensity of the sidelobe signals corresponding to the sidelobe angle set and the key suppression area angle set is constrained. The sidelobe signals corresponding to the key suppression area angle set are the sidelobe signals corresponding to the other sound source directions, and the sidelobe signals corresponding to the sidelobe angle set are the sidelobe signals other than the sidelobe signals corresponding to the other sound source directions.
2. The voice processing system according to claim 1, wherein The direction determination module is further configured to determine a pickup beam in which the microphone array points to the sound source direction according to the sound source direction of the voice signals, so as to pick up the voice signals corresponding to the sound source direction through the pickup beam.
3. The voice processing system according to claim 2, wherein When there are multiple sound sources in the target space, the voice processing system further includes: A sound source separation module, configured to determine a pickup beam pointing to each sound source direction according to the sound source direction of each sound source among the multiple sound sources, and determine a target sound source that currently generates sound signals from the multiple sound sources; The voice receiving module is further configured to determine a target pickup beam pointing to the sound source direction corresponding to the target sound source, so as to pick up the voice signals of the target sound source through the target pickup beam.
4. The voice processing system according to claim 2, characterized in that For the pickup beam of each sound source, the main lobe signal of the pickup beam covers the sound source direction where the sound source is located, and the sidelobe signals corresponding to the sound source directions of the remaining sound sources except the sound source in the pickup beam are suppressed.
5. The voice processing system according to claim 1, wherein The direction determination module is further configured to determine at least one activated sound zone position in the multiple sound zones through the activation states of the multiple sound zones in the target space, and perform direction estimation according to the voice signals, so as to determine the sound source direction of the sound source corresponding to the voice signals within the at least one activated sound zone position.
6. The voice processing system according to any one of claims 1-5, characterized in that, The voice receiving module further includes at least one of the following modules: An echo cancellation module, configured to perform echo cancellation processing on the voice signals picked up by the microphone array; A non-linear post-processing module, configured to perform non-linear post-processing on the speech signal processed by the echo cancellation module; A voiceprint enhancement module, configured to perform voiceprint enhancement processing and speech noise reduction processing on the speech signal according to the voiceprint information of the sound source; A noise suppression module, configured to perform noise suppression processing on the speech signal; An automatic gain module, configured to perform automatic gain processing on the speech signal so that the gain of the speech signal is within a first preset range; An equalization module, configured to perform signal equalization processing on the speech signal.
7. The voice processing system according to any one of claims 1-5, characterized in that, The sound playback module further includes at least one of the following modules: A noise suppression module, configured to perform noise suppression processing on the sound signal to be played; An automatic gain module, configured to perform automatic gain processing on the sound signal to be played so that the gain of the sound signal to be played is within a second preset range; An equalization module, configured to perform signal equalization processing on the sound signal to be played.
8. The voice processing system according to any one of claims 1-5, characterized in that, The target space is inside a vehicle cockpit.
9. A voice processing method, characterized in that, Including: Obtain a speech signal in the target space, and determine the sound source azimuth of the sound source corresponding to the speech signal in the target space according to the speech signal; Determine a converging beam pointing to the sound source azimuth according to the sound source azimuth, and play a sound signal to be played to the sound source azimuth through the converging beam, where the sound signal to be played is sent by an object that performs a voice interaction with the sound source; Determining a converging beam pointing to the sound source azimuth according to the sound source azimuth includes: According to the sound source azimuth, determine a converging beam pointing to the sound source azimuth based on equal sidelobe beamforming, so that the main lobe of the power radiation pattern of the speaker array points to the sound source azimuth, and the sidelobe signals corresponding to other sound source azimuths except the sound source azimuth are sunken, so as to reduce the energy diffusion of the sound signal to be played in other sound source azimuths; In the equal sidelobe beamforming, the signal intensity of the sidelobe signals corresponding to the sidelobe angle set and the key suppression area angle set is constrained, the sidelobe signals corresponding to the key suppression area angle set are the sidelobe signals corresponding to other sound source azimuths, and the sidelobe signals corresponding to the sidelobe angle set are the sidelobe signals other than the sidelobe signals corresponding to other sound source azimuths.
10. A non-transitory computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instruction is executed by a processor, it implements the speech processing method described in claim 9.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the speech processing method described in claim 9.
12. A multimedia device, characterized in that, Including the speech processing system according to any one of claims 1-8.
Citation Information
Patent Citations
Sidelobe suppression method suitable for orbital angular momentum three-dimensional imaging sonar
CN112083430A
Microphone array on an aircraft to determine position and
CN114598962A
Array loudspeaker system for directional sound reinforcement and implementation method
CN117221800A
Voice signal processing method and device, vehicle and storage medium
CN117789741A