A voice processing method, device, apparatus and storage medium

By acquiring the orientation and signal-to-noise ratio of the speech signal, determining the target orientation and signal-to-noise ratio thresholds, and extracting the target speech features, the problems of speech separation and suppression of human voice interference in voice conferencing are solved, thus improving the quality of voice communication.

CN115641863BActive Publication Date: 2026-04-14LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2022-09-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively separate target speech and suppress human voice interference in voice conferencing, and blind source separation algorithms have sound source ordering problems, resulting in poor voice communication quality.

Method used

By acquiring the speech orientation and signal-to-noise ratio of the speech signal, the target orientation and signal-to-noise ratio thresholds are determined, target speech features are extracted, and speech separation and enhancement processing are performed under different signal-to-noise ratio conditions.

Benefits of technology

It improves voice communication quality under extremely low signal-to-noise ratio conditions, enhances the separation of speaker speech content, avoids speech damage and distortion, and has environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641863B_ABST
    Figure CN115641863B_ABST
Patent Text Reader

Abstract

The application discloses a speech processing method, device and equipment and a storage medium. The method comprises the following steps: acquiring a speech direction and a speech signal-to-noise ratio of a speech signal; determining that the speech direction is a target direction and the speech signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold; and extracting a target speech feature based on the speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech technology, and includes, but is not limited to, a speech processing method, apparatus, device, and storage medium. Background Technology

[0002] Working from home is an important way of working, and online voice conferencing is also an important way of communication when working from home. Therefore, obtaining good voice communication quality is a very important technical requirement.

[0003] Existing technical solutions detect the speaker's speaking direction using dual-microphone signals and control the microphone's on / off state during the meeting. There is no further speech enhancement or separation operation. Subsequent speech enhancement algorithms mainly remove noise and do not significantly suppress human voice interference. Blind source separation algorithms generally suffer from the problem of sound source ordering, making it impossible to determine which sound source signal is the target sound source. Summary of the Invention

[0004] In view of the above, embodiments of this application provide a voice processing method, apparatus, device, and storage medium.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a voice processing method, the method comprising:

[0007] Acquire the speech orientation and speech signal-to-noise ratio of the speech signal;

[0008] The speech direction is determined to be the target direction, and the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold;

[0009] Target speech features are extracted based on the speech signal.

[0010] Secondly, embodiments of this application provide a voice processing apparatus, the apparatus comprising:

[0011] The first acquisition module is used to acquire the speech orientation and speech signal-to-noise ratio of the speech signal;

[0012] The first determining module is used to determine that the speech orientation is a target orientation and that the speech signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold.

[0013] The first extraction module is used to extract target speech features based on the speech signal.

[0014] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the above-described method.

[0015] Fourthly, embodiments of this application provide a storage medium storing executable instructions for inducing a processor to execute the above-described method.

[0016] In this embodiment, the speech orientation and signal-to-noise ratio (SNR) of the speech signal are first obtained; then, the speech orientation is determined to be the target orientation, and the SNR is greater than or equal to a SNR threshold; finally, target speech features are extracted based on the speech signal. Thus, when the obtained speech signal has the target orientation and a high SNR, target speech features are extracted based on the speech signal. Corpus data that better reflects actual communication situations can be used for speaker speech feature extraction, thereby achieving better speaker speech content separation even under low SNR conditions, effectively enhancing the voice communication experience under extremely low SNR / SNR conditions. Attached Figure Description

[0017] Figure 1A A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0018] Figure 1B A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0019] Figure 1C A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0020] Figure 2A A schematic diagram illustrating the implementation process of a speech feature extraction method provided in this application embodiment;

[0021] Figure 2B A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0022] Figure 3A A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0023] Figure 3B A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0024] Figure 4A A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0025] Figure 4B A schematic diagram illustrating the implementation flow of a speech processing method provided in an embodiment of this application;

[0026] Figure 5A This is a schematic diagram of the composition structure of a voice processing device provided in an embodiment of this application;

[0027] Figure 5B This is a schematic diagram of the composition structure of a voice processing device provided in an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of the embodiments will be further described in detail below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0030] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0033] This application provides a speech processing method, such as... Figure 1A As shown, the method includes:

[0034] Step S110: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0035] During implementation, a voice detection component can be used to detect the direction of the voice signal. For example, sound signals can be collected using the dual microphones of a laptop computer to determine the direction of the voice and whether the speaker is speaking towards the laptop.

[0036] During implementation, multi-channel voice signals can be acquired using a microphone array on a laptop computer to obtain the voice signal-to-noise ratio.

[0037] Step S120: Determine that the speech direction is the target direction, and that the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold;

[0038] Here, the target orientation can be determined based on the actual application scenario, referring to the orientation of the voice signal to be collected. For example, when a user is participating in an online meeting using a laptop, the direction in which the user is speaking towards the laptop can be determined as the target orientation. Determining the voice orientation as the target orientation means determining that the collected voice is the voice facing the screen.

[0039] During implementation, while determining the speech direction as the target direction, it's also possible to determine that the speech signal-to-noise ratio (SNR) is greater than or equal to a SNR threshold. The SNR threshold can be determined based on the specific application scenario. The implementation can proceed sequentially: first, determine the speech direction as the target direction, then determine that the speech SNR is greater than or equal to the SNR threshold; alternatively, first, determine that the speech SNR is greater than or equal to the SNR threshold, then determine that the speech direction is the target direction; or, the speech direction can be determined as the target direction while simultaneously having a speech SNR greater than or equal to the SNR threshold. For example, in this case, it can be determined that the user is speaking towards the laptop screen, and the acquired speech signal is of good quality.

[0040] Step S130: Extract target speech features based on the speech signal.

[0041] Here, target speech features can be extracted based on the speech signal, and these target speech features can be stored as user speech features for application in scenarios with poor signal quality (low speech signal-to-noise ratio). They can also be used as reference speech features for speech separation to obtain a better quality speech signal.

[0042] In this embodiment, the speech orientation and signal-to-noise ratio (SNR) of the speech signal are first obtained; then, the speech orientation is determined to be the target orientation, and the SNR is greater than or equal to a SNR threshold; finally, target speech features are extracted based on the speech signal. Thus, when the obtained speech signal has the target orientation and a high SNR, target speech features are extracted based on the speech signal. Corpus data that better reflects actual communication situations can be used for speaker speech feature extraction, thereby achieving better speaker speech content separation even under low SNR conditions, effectively enhancing the voice communication experience under extremely low SNR / SNR conditions.

[0043] This application provides a speech processing method, such as... Figure 1B As shown, the method includes:

[0044] Step S110: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0045] Step S120: Determine the direction of the voice as the target direction;

[0046] Step S140: Determine that the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold;

[0047] During implementation, steps S120 and S140 are executed sequentially to determine the speech orientation as the target orientation, while ensuring the speech signal-to-noise ratio (SNR) is less than the SNR threshold. For example, in this case, it can be determined that the user is speaking towards the laptop screen, but the acquired speech signal is of poor quality.

[0048] Step S150: Extract the target speech signal from the speech signal that matches the target speech feature;

[0049] During implementation, the target speech features obtained after completing step S130 in the above embodiments can be used as a reference to extract the target speech signal that matches the target speech features from the speech signal. That is, in the case of low signal-to-noise ratio, the user's target speech signal can be extracted based on the target speech features.

[0050] Step S160: Enhance the target speech signal.

[0051] In this embodiment, the speech orientation and signal-to-noise ratio (SNR) of the speech signal are first obtained; then, the speech orientation is determined to be the target orientation; the SNR is determined to be less than the SNR threshold; finally, the target speech signal matching the target speech features is extracted from the speech signal; and the target speech signal is enhanced. In this way, speech separation can be performed even when the speech orientation is determined to be the target orientation and the SNR is less than the SNR threshold, meaning that a separation algorithm is necessary in low SNR situations. This allows users to obtain higher quality target speech signals, avoids speech damage and distortion caused by using the same speech separation algorithm at high SNRs, and has adaptive capabilities to the environment.

[0052] This application provides a speech processing method, such as... Figure 1C As shown, the method includes:

[0053] Step S110: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0054] Step S170: Determine that the voice direction is another direction besides the target direction;

[0055] Here, the target orientation can be determined based on the actual application scenario, referring to the orientation of the voice signal to be collected. For example, when a user is participating in an online meeting using a laptop, the direction in which the user is speaking towards the laptop can be determined as the target orientation. Determining the voice orientation as an orientation other than the target orientation, i.e., determining that the collected voice is not facing the screen, could be voice coming from the side or back of the laptop.

[0056] Step S150: Extract the target speech signal from the speech signal that matches the target speech feature;

[0057] Here, after performing step S170, step S150 in the above embodiments can be performed.

[0058] For example, when acquiring lateral speech from a laptop, it can be determined that the speech is corpus subject to interference. Therefore, the interference can be filtered out using a speech separation algorithm, which is beneficial for extracting the target speech signal more accurately.

[0059] Step S160: Enhance the target speech signal.

[0060] In this embodiment, the speech orientation and signal-to-noise ratio of the speech signal are first obtained; then, the speech orientation is determined to be an orientation other than the target orientation; finally, the target speech signal matching the target speech features is extracted from the speech signal; and the target speech signal is enhanced. In this way, speech separation can be performed even when the speech orientation is determined to be an orientation other than the target orientation, allowing users to obtain a higher quality target speech signal.

[0061] This application provides a voice processing method applied to a laptop computer, such as... Figure 2A As shown, the method includes:

[0062] Step S210: Detect the speaker's speaking direction using a speech detection component;

[0063] During implementation, sound signals can be collected using the dual microphones of a laptop computer, and the direction of the voice can be determined using the voice detection component installed in the laptop computer.

[0064] In some embodiments, multiple microphones located at different positions on the laptop can be used to collect sound signals to determine the speaker's speaking orientation. For example, if the sound signal collected by the microphone located on side B of the laptop is greater than the sound signal collected by the microphone located on side A of the laptop, it can be determined that the speaker is facing the screen.

[0065] Step S220: Determine whether the speaker is facing the screen while speaking;

[0066] For example, when a user (speaker) is using a laptop to participate in an online meeting, the content of the user speaking towards the screen can be coherent and clear, so it can be determined beforehand whether the speaker is facing the screen.

[0067] If it is determined that the speaker is facing the screen, proceed to step S230; if it is determined that the speaker is not facing the screen, proceed to step S260.

[0068] Step S230: Obtain the speech signal-to-noise ratio;

[0069] During implementation, multi-channel voice signals can be acquired through the microphone array on a laptop computer to estimate the signal-to-noise ratio of the current call state, thereby determining the ratio of voice signal to noise signal.

[0070] Step S240: Determine whether the speech signal-to-noise ratio is greater than the signal-to-noise ratio threshold;

[0071] Here, the signal-to-noise ratio threshold can be determined based on the actual application scenario.

[0072] If the speech signal-to-noise ratio is determined to be greater than or equal to the signal-to-noise ratio threshold, step S250 is executed; if the speech reconstruction ratio is determined to be less than the signal-to-noise ratio threshold, step S260 is executed.

[0073] Step S250: Extract the speaker's speech features using the current speech signal;

[0074] During implementation, since the speech signal-to-noise ratio (SNR) is greater than or equal to the SNR threshold, meaning the current speech environment has relatively low noise and the speaker's speech quality is high, the acquired speech signal can be saved, and the pre-registered speaker feature vector can be fine-tuned without performing speaker speech separation. Here, fine-tuning means updating the pre-registered speaker feature vector based on the acquired speaker feature vector, to be applied to scenarios where the speech SNR is less than the SNR threshold and speaker speech separation is required.

[0075] By continuously recording the target speaker's speech data and updating the pre-registered speaker feature vectors, speech data that better reflects the actual communication situation can be used for speaker feature extraction. This results in better speaker speech content separation during speech separation, enhancing the voice conferencing communication experience under extremely low signal-to-noise ratio / signal-to-interference ratio conditions.

[0076] Step S260: Using the updated speech features, perform speech separation of the target speaker on the original signal to obtain the separated speech signal.

[0077] During implementation, since the speech signal-to-noise ratio (SNR) is lower than the SNR threshold (i.e., the current speech environment is quite noisy), a deep learning model can be used to separate the speech of a specific speaker from the mixed speech signal using pre-stored speaker feature vectors or updated speaker feature vectors, in order to extract the enhanced signal of the target speaker. This achieves better speaker speech content separation and enhances the voice conferencing communication experience under extremely low SNR / SNR conditions.

[0078] In this embodiment, the speaker's orientation is first detected. If it is determined that the speaker is not facing the screen, speech separation of the target speaker is performed, and the speaker's enhanced signal is extracted. If the speaker's orientation is detected to be facing the screen, the speech signal-to-noise ratio (SNR) is then obtained. If the SNR is determined to be greater than or equal to a threshold, the speaker's speech features are extracted using the current speech signal. If the SNR is determined to be less than the threshold, the updated speech features are used to perform speech separation of the target speaker, and the speaker's enhanced signal is extracted. In this way, in a high-quality speech environment, the target speaker's speech data can be continuously recorded and the pre-registered speaker feature vector can be updated to use speech data that better reflects the actual communication situation for speaker feature extraction. Speech separation of a specific speaker is performed only when necessary, avoiding speech damage and distortion caused by using the same speech separation algorithm. Furthermore, this method is adaptive to the environment.

[0079] In existing technologies, it is necessary to register the speech of the target speaker in advance and extract the corresponding speech features. Due to the influence of the target speaker's voice state and environmental factors, the corpus data of a single registration process is often insufficient, which limits the accuracy of speaker speech feature extraction and affects the separation performance.

[0080] In some embodiments, such as Figure 2B As shown, step S250 "using the current signal to extract speaker features" can be achieved through the following steps:

[0081] Step S251: Save the current voice signal;

[0082] Step S252: Extract or update the speaker's speech features using the speaker feature extraction component.

[0083] In some embodiments, if the speaker's voice features are not pre-stored, the speaker's voice features can be extracted using a speaker feature extraction component, and then used in scenarios where voice separation is required.

[0084] In some embodiments, if the speaker's voice features are pre-stored, the speaker's voice features can be updated using a speaker feature extraction component.

[0085] In this embodiment of the application, by continuously recording the target speaker's speech data and updating the pre-registered speaker feature vector, speech data that is more consistent with the actual communication situation can be used for speaker feature extraction, thereby achieving a better speaker speech content separation effect and enhancing the voice conference communication experience under extremely low signal-to-noise ratio / signal-to-interference ratio conditions.

[0086] This application provides a speech processing method, such as... Figure 3A As shown, the method includes:

[0087] Step S310: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0088] Step S320: Determine that the speech direction is the target direction, and that the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold;

[0089] Step S330: Extract target speech features based on the speech signal;

[0090] Step S340: Determine that the voice direction is other than the target direction;

[0091] During implementation, step S330 is executed sequentially, meaning that if the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the speech orientation can be determined to be an orientation other than the target orientation. For example, in this case, it can be determined that the acquired speech signal is a high-quality speech signal, but the user is not speaking towards the laptop screen.

[0092] Step S350: Extract the target speech signal from the speech signal that matches the target speech features;

[0093] During implementation, the target speech features obtained after completing step S330 in the above embodiments can be used as a reference to extract the target speech signal that matches the target speech features from the speech signal. That is, when the speech direction is other than the target direction, the user's target speech signal is extracted based on the target speech features.

[0094] Step S360: Enhance the target speech signal.

[0095] In this embodiment, the speech orientation and signal-to-noise ratio (SNR) of the speech signal are first obtained; the speech orientation is determined to be the target orientation, and the SNR is greater than or equal to the SNR threshold; then, the speech orientation is determined to be an orientation other than the target orientation; finally, the target speech signal matching the target speech features is extracted from the speech signal; and the target speech signal is enhanced. In this way, even when the SNR is determined to be greater than or equal to the SNR threshold, but the speech orientation is an orientation other than the target orientation (i.e., the speech requiring separation algorithm processing is not oriented towards the target orientation), speech separation can be performed using the target speech features obtained when the speech orientation is the target orientation, allowing users to obtain higher quality target speech signals.

[0096] This application provides a speech processing method, such as... Figure 3B As shown, the method includes:

[0097] Step S310: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0098] Step S370: Determine that the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold;

[0099] Here, multi-channel voice signals are acquired through the microphone array on the laptop, and the signal-to-noise ratio (SNR) of the current call state is estimated to determine if the voice SNR is less than the SNR threshold.

[0100] Step S350: Extract the target speech signal from the speech signal that matches the target speech features;

[0101] Step S360: Enhance the target speech signal.

[0102] In this embodiment, the speech orientation and signal-to-noise ratio (SNR) of the speech signal are first obtained; then, it is determined that the SNR is less than a SNR threshold; finally, a target speech signal matching the target speech features is extracted from the speech signal; and the target speech signal is enhanced. In this way, speech separation can be performed even when the SNR is determined to be less than the SNR threshold, allowing users to obtain a higher quality target speech signal.

[0103] In some embodiments, step S130 "extracting target speech features based on the speech signal" can be implemented through the following steps:

[0104] Step 131: Obtain initial speech features;

[0105] During implementation, users can pre-set voice features. For example, before conducting an online meeting, users' voices can be collected as preset voice features, i.e., initial voice features.

[0106] Step 132: Update the initial speech features based on the speech signal to obtain the target speech features.

[0107] Here, due to differences in different environments or the content of a user's speech, the target speech features of the same user may vary. During implementation, by continuously recording the target speaker's speech data and updating the pre-registered speaker feature vector (initial speech features), speech data that better reflects actual communication situations can be used for speaker feature extraction.

[0108] In this embodiment, preset target speech features are first obtained; then, the target speech features are updated based on the speech signal. This allows for updating the target speech features using the speech signal in appropriate scenarios to obtain target speech features that better reflect the actual communication context.

[0109] This application provides a voice processing method applied to a laptop computer, such as... Figure 4A As shown, the method includes:

[0110] Step S401: Obtain the speech signal-to-noise ratio;

[0111] During implementation, multi-channel voice signals can be collected through the microphone array on a laptop computer to estimate the signal-to-noise ratio of the current call status.

[0112] Step S402: Determine whether the speech signal-to-noise ratio is greater than the signal-to-noise ratio threshold;

[0113] During implementation, the speech signal-to-noise ratio (SNR) can be distinguished as high or low based on a preset SNR threshold. If it is determined to be a high SNR state (speech SNR greater than or equal to the SNR threshold), then step S403 is executed; if it is determined to be a low SNR state (speech SNR less than the SNR threshold), then step S406 is executed.

[0114] Step S403: Detect the speaker's speaking direction using a speech detection component;

[0115] Here, the method described in step S210 above can be used to detect the speaker's speaking direction.

[0116] Step S404: Determine whether the speaker is facing the screen while speaking;

[0117] If it is determined that the speaker is facing the screen, proceed to step S405; if it is determined that the speaker is not facing the screen, proceed to step S406.

[0118] Step S405: Extract the speaker's speech features using the current speech signal;

[0119] During execution, the following can be used: Figure 2B The steps shown are for extracting the speaker's speech features.

[0120] Step S406: Using the updated speech features, perform speech separation of the target speaker on the original speech signal to obtain the separated speech signal.

[0121] Here, the method of step S260 above can be used to perform speech separation of the target speaker on the original speech signal using the updated speech features, and obtain the separated speech signal.

[0122] In this embodiment, the speech signal-to-noise ratio (SNR) is first obtained; if the SNR is less than a threshold, speech separation is performed; if the SNR is greater than or equal to the threshold, the speaker's orientation is detected; if the speaker is facing the screen, their speech features are extracted; and if the speaker is facing the screen, speech separation is performed again. This allows for speech separation of specific speakers only when necessary; in high-quality speech environments, the target speaker's speech data can be continuously recorded and the pre-registered speaker feature vectors updated, using more realistic speech data for speaker feature extraction; this effectively avoids speech impairment and distortion caused by using the same speech separation algorithm, and the method is adaptable to different environments.

[0123] This application provides a speech processing method, such as... Figure 4B As shown, it includes the following steps:

[0124] Step S411: Obtain the speech orientation and speech signal-to-noise ratio of the speech signal;

[0125] Step S412: Process the speech signal based on the speech orientation and speech signal-to-noise ratio;

[0126] Specifically, when the speech direction is determined to be the target direction and the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, target speech features are extracted based on the speech signal.

[0127] If the speech orientation is determined to be non-target orientation, or the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold, a target speech signal matching the target speech features is extracted from the speech signal to enhance the target speech signal.

[0128] In this embodiment, the speech orientation and speech signal-to-noise ratio (SNR) of the speech signal are first obtained; then, the speech signal is processed based on the speech orientation and SNR. This allows for the extraction of target speech features from the speech signal when the speech orientation is determined to be a target orientation and the SNR is greater than or equal to a SNR threshold; conversely, when the speech orientation is determined to be non-target orientation, or the SNR is less than the SNR threshold, a target speech signal matching the target speech features is extracted from the speech signal to enhance the target speech signal.

[0129] Based on the foregoing embodiments, this application provides two voice processing devices. The device includes various modules, each module including sub-modules, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0130] Figure 5A This is a schematic diagram of the composition structure of the voice processing device provided in the embodiments of this application, as shown below. Figure 5A As shown, the device 500 includes:

[0131] The first acquisition module 501 is used to acquire the speech orientation and speech signal-to-noise ratio of the speech signal;

[0132] The first determining module 502 is used to determine that the speech direction is a target direction and that the speech signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold.

[0133] The first extraction module 503 is used to extract target speech features based on the speech signal.

[0134] In some embodiments, the apparatus further includes a second determining module, a second extraction module, and an enhancement module, wherein the second determining module is configured to determine that the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold; the second extraction module is configured to extract a target speech signal from the speech signal that matches the target speech feature; and the enhancement module is configured to enhance the target speech signal.

[0135] In some embodiments, the apparatus further includes a third determining module for determining that the voice orientation is an orientation other than the target orientation.

[0136] In some embodiments, the first extraction module 503 includes an acquisition submodule and an update submodule, wherein the acquisition submodule is used to acquire preset target speech features; and the update submodule is used to update the target speech features based on the speech signal.

[0137] Figure 5B This is a schematic diagram of the composition structure of the voice processing device provided in the embodiments of this application, as shown below. Figure 5B As shown, the device 510 includes:

[0138] The second acquisition module 511 is used to acquire the speech orientation and speech signal-to-noise ratio of the speech signal;

[0139] Processing module 512 is used to process the speech signal based on the speech orientation and speech signal-to-noise ratio;

[0140] Specifically, when the speech direction is determined to be the target direction and the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, target speech features are extracted based on the speech signal.

[0141] If the speech orientation is determined to be non-target orientation, or the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold, a target speech signal matching the target speech features is extracted from the speech signal to enhance the target speech signal.

[0142] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0143] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause electronic devices (such as mobile phones, tablets, laptops, desktop computers, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0144] Correspondingly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the speech processing method provided in the above embodiments.

[0145] Correspondingly, embodiments of this application provide an electronic device, Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the hardware entity of the device 600 includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the voice processing method provided in the above embodiments.

[0146] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) in the processor 602 and various modules in the electronic device 600. It can be implemented by flash memory or random access memory (RAM).

[0147] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0148] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0150] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0151] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0152] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0153] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0154] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0155] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0156] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0157] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0158] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A speech processing method, characterized in that, The method includes: Acquire the speech orientation and speech signal-to-noise ratio of the speech signal; The voice direction is determined as the target direction, wherein the target direction is the speaking direction facing the screen of the electronic device; When the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the initial speech features are updated based on the speech signal to obtain target speech features, so as to enhance the target speech signal when the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold. The target speech signal is a signal in the speech signal that matches the target speech features.

2. The method of claim 1, wherein after acquiring the speech orientation and speech signal-to-noise ratio of the speech signal, the method further comprises: The direction of the voice is determined to be an orientation other than the target orientation; Extract the target speech signal from the speech signal that matches the target speech feature; Enhance the target speech signal.

3. The method of claim 1, wherein after determining that the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the method further comprises: The direction of the voice is determined to be an orientation other than the target orientation; Extract the target speech signal from the speech signal that matches the target speech feature; Enhance the target speech signal.

4. The method of claim 1, wherein after acquiring the speech orientation and speech signal-to-noise ratio of the speech signal, the method further comprises: The speech signal-to-noise ratio is determined to be less than the signal-to-noise ratio threshold; Extract the target speech signal from the speech signal that matches the target speech feature; Enhance the target speech signal.

5. A speech processing method, the method comprising: Acquire the speech orientation and speech signal-to-noise ratio of the speech signal; The speech signal is processed based on the speech orientation and speech signal-to-noise ratio; Wherein, when the speech orientation is determined to be the target orientation, and the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, target speech features are extracted based on the speech signal to enhance the target speech signal when the speech orientation is determined to be not the target orientation or the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold; wherein, the target orientation is the speaking direction facing the screen of the electronic device, and the target speech signal is the signal in the speech signal that matches the target speech features.

6. A voice processing apparatus, the apparatus comprising: The first acquisition module is used to acquire the speech orientation and speech signal-to-noise ratio of the speech signal; The first determining module is used to determine that the voice direction is the target direction; The first extraction module is configured to update the initial speech features based on the speech signal to obtain target speech features when the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, so as to enhance the target speech signal when the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold; wherein the target orientation is the speaking direction facing the screen of the electronic device, and the target speech signal is the signal in the speech signal that matches the target speech features.

7. A voice processing apparatus, the apparatus comprising: The second acquisition module is used to acquire the speech orientation and speech signal-to-noise ratio of the speech signal; The processing module processes the speech signal based on the speech orientation and speech signal-to-noise ratio; Specifically, when the speech orientation is determined to be the target orientation, and the speech signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the initial speech features are updated based on the speech signal to obtain target speech features, so as to enhance the target speech signal when the speech orientation is determined to be non-target orientation or when the speech signal-to-noise ratio is less than the signal-to-noise ratio threshold; the target orientation is the speaking direction facing the screen of the electronic device, and the target speech signal is the signal in the speech signal that matches the target speech features.

8. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice wake-up intelligent control method and device, electronic equipment and storage medium

    CN113066488A

  • Speech enhancement method and device, equipment, storage medium and program

    CN113223552A