Speech enhancement method and conference system
By using AI models, feature extraction networks, and attention networks, the problem of traditional beamforming technology being unable to track the direction of target speech in a timely manner has been solved, achieving fast and accurate speech enhancement and improving the user experience of the conferencing system.
Patent Information
- Application Number
- CN202511756358.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-24
AI Technical Summary
In traditional conference systems, beamforming technology cannot track the target speech direction in a timely manner when switching speakers, resulting in suppression of the new speaker's voice and affecting the voice quality and user experience.
An AI model is used in conjunction with an omnidirectional full-frequency beam feature extraction network, an attention network, and a mask feature acquisition network. The trained AI model can quickly and accurately determine the direction of the target speech, switch the beam direction in a timely manner, and further enhance the target speech data through mask features.
It enables fast and accurate tracking of the target speech direction, reduces suppression of new speakers' speech, obtains clear speech data, and significantly improves the user experience.
Smart Images

Figure CN121565191A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, specifically to a speech enhancement method and a conferencing system. Background Technology
[0002] In traditional conferencing systems, speech enhancement technology primarily relies on beamforming to enhance speech from a specific direction. Beamforming extracts the speech signal by focusing on a specific direction, thereby suppressing noise and interference from other directions. However, both fixed-beam and adaptive-beam techniques involve calculating the target speech direction or target statistics.
[0003] Traditional statistical methods for obtaining these statistics typically require a considerable amount of time to complete. When a speaker changes, the update of the target speech direction or target statistics takes time, which can prevent the beam from switching to the new target speech direction in a timely manner. This can suppress some of the new speaker's voice, affecting the audio quality and user experience. Therefore, how to track the target speech direction in a timely manner, reduce or avoid suppression of the new speaker's voice, and thus obtain stable and clear speaker audio to improve the experience of meeting participants has become an urgent problem to be solved.
[0004] Accordingly, there is a need in the field for a new speech enhancement solution to address the aforementioned problems. Summary of the Invention
[0005] In order to overcome the above-mentioned deficiencies, this application is made to solve, or at least partially solve, the technical problem of how to track the target speech direction in a timely and accurate manner to improve the speech enhancement effect of beamforming.
[0006] In a first aspect, a speech enhancement method is provided, the method comprising: Acquire first voice array data in the first direction and second voice array data in the second direction; According to the beamforming settings in the first direction, the first voice array data is beamformed to obtain full-frequency beam data in the first direction; According to the beamforming settings in the second direction, the second voice array data is beamformed to obtain full-frequency beam data in the second direction; Based on the full-frequency beam data in the first direction and the full-frequency beam data in the second direction, the target angle in the first direction and the target angle in the second direction are obtained through a trained AI model; Based on the target angle in the first direction and the target angle in the second direction, beamforming is performed on the area array composed of the first voice array data and the second voice array data to obtain enhanced target voice data.
[0007] In the above-mentioned speech enhancement method, the AI model includes an omnidirectional full-frequency beam feature extraction network, a first attention network, and a second attention network. "Based on the first-direction full-frequency beam data and the second-direction full-frequency beam data, the trained AI model is used to obtain the first-direction target angle and the second-direction target angle," which includes: Based on the first direction full-frequency beam data and the second direction full-frequency beam data, the omnidirectional full-frequency beam features are obtained through the omnidirectional full-frequency beam feature extraction network; Based on the omnidirectional full-frequency beam characteristics and the first-direction full-frequency beam data, the first-direction target angle is obtained through the first attention network; Based on the omnidirectional full-frequency beam characteristics and the second-direction full-frequency beam data, the second-direction target angle is obtained through the second attention network; The first directional full-frequency beam data includes a first number of first low-frequency beam data and a first number of first high-frequency beam data, and the second directional full-frequency beam data includes a second number of second low-frequency beam data and a second number of second high-frequency beam data.
[0008] In the above-mentioned speech enhancement method, the omnidirectional full-frequency beam feature extraction network includes a first feature extraction network, a second feature extraction network, and a third feature extraction network, and the method for acquiring the omnidirectional full-frequency beam features includes: Based on the first low-frequency beam data and the second low-frequency beam data, omnidirectional low-frequency beam features are obtained through the first feature extraction network. Based on the first high-frequency beam data and the second high-frequency beam data, omnidirectional high-frequency beam features are obtained through the second feature extraction network; Based on the omnidirectional low-frequency beam features and the omnidirectional high-frequency beam features, the omnidirectional full-frequency beam features are obtained through the third feature extraction network.
[0009] In the above-mentioned speech enhancement method, the first feature extraction network, the second feature extraction network, and the third feature extraction network are all constructed based on convolutional networks.
[0010] In the above-mentioned speech enhancement method, the first attention network includes a first query feature acquisition network, a first key feature acquisition network, and a first weight acquisition network, and the method for obtaining the first directional target angle includes: Based on the full-frequency beam data in the first direction, the first query feature is obtained through the first query feature acquisition network; Based on the full-frequency beam data in the first direction, the first key feature is obtained through the first key feature acquisition network; Based on the first Query feature and the first Key feature, the first weighted value corresponding to each beam angle in the first direction is obtained through the first weight acquisition network. The beam angle corresponding to the first weighted value with the largest value is selected as the target angle in the first direction.
[0011] In the above-mentioned speech enhancement method, the second attention network includes a second query feature acquisition network, a second key feature acquisition network, and a second weight acquisition network. The method for obtaining the second directional target angle includes: Based on the second direction full-frequency beam data, the second query feature is obtained through the second query feature acquisition network; Based on the second direction full-frequency beam data, the second key feature is obtained through the second key feature acquisition network; Based on the second Query feature and the second Key feature, the second weighting network is used to obtain the second weighting value corresponding to each beam angle in the second direction; The beam angle corresponding to the second weighted value with the largest value is selected as the target angle in the second direction.
[0012] In the above-mentioned speech enhancement method, both the first query feature acquisition network and the second query feature acquisition network are constructed based on fully connected neural networks; Both the first Key feature acquisition network and the second Key feature acquisition network are constructed based on convolutional networks; Both the first weight acquisition network and the second weight acquisition network are constructed based on the Softmax function.
[0013] In the above-mentioned speech enhancement method, the AI model further includes a MASK feature acquisition network, and the method further includes: Based on the omnidirectional full-frequency beam characteristics, the target enhanced MASK characteristics are obtained through the MASK feature acquisition network. Based on the target enhanced MASK features and the enhanced target speech data, further enhanced target speech data is obtained.
[0014] In the above-mentioned speech enhancement method, the MASK feature acquisition network is constructed based on a gated recurrent neural network and a fully connected neural network.
[0015] In a second aspect, a conference system is provided, the conference system comprising: Multiple first voice acquisition devices, wherein the multiple first voice acquisition devices form a uniform or non-uniform linear array and are arranged linearly along a first direction; Multiple second voice acquisition devices, wherein the multiple second voice acquisition devices form a uniform or non-uniform linear array and are arranged linearly along a second direction; A conference terminal configured to implement the speech enhancement method as described in any of the preceding claims.
[0016] The above-mentioned technical solutions of this application have at least one or more of the following beneficial effects: Utilizing the efficient data processing and learning capabilities of AI models, the direction of target speech can be quickly and accurately determined, and the beam direction can be switched in a timely manner to avoid or reduce suppression of the speech of new speakers in the conference, thereby obtaining clear speech and improving the user experience. Furthermore, by further enhancing the target speech data through target augmentation MASK features, even clearer speech can be obtained, greatly improving the user experience. Attached Figure Description
[0017] The disclosure of this application will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0018] Figure 1 This is a schematic diagram of a conference system according to an embodiment of this application; Figure 2 (a) is a schematic flowchart of a speech enhancement method according to an embodiment of this application. Figure 2 (b) is a schematic diagram of the structure of an AI model according to an embodiment of this application; Figure 3 This is a schematic flowchart of the main steps of a speech enhancement method according to an embodiment of this application; Figure 4 (a) is a schematic flowchart of a speech enhancement method according to another embodiment of this application. Figure 4 (b) is a schematic diagram of the structure of an AI model according to another embodiment of this application; Figure 5 This is a schematic flowchart of the main steps of a speech enhancement method according to another embodiment of this application. Detailed Implementation
[0019] Some embodiments of this application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.
[0020] In the description of this application, "module" and "processor" can include hardware, software, or a combination of both. A module can include hardware circuitry, various suitable sensors, communication ports, memory, and may also include software components, such as program code, or a combination of software and hardware. A processor can be a central processing unit, microprocessor, image processor, digital signal processor, or any other suitable processor. The processor has data and / or signal processing capabilities. The processor can be implemented in software, in hardware, or a combination of both. Computer-readable storage media includes any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The terms "at least one A or B" or "at least one of A and B" have a similar meaning to "A and / or B" and can include only A, only B, or A and B. The singular terms "a" or "this" can also include plural forms.
[0021] First, please refer to the appendix. Figure 1 , Figure 1 This is a schematic diagram of a conference system according to an embodiment of the present application. In the embodiment of the present application, the conference system includes a first voice acquisition device array 101 composed of a plurality of first voice acquisition devices, a second voice acquisition device array 102 composed of a plurality of second voice acquisition devices, and a conference terminal 103.
[0022] In the conference room of the conference system, a large conference screen 104 is mounted on the wall. Multiple first voice acquisition devices installed at the top of the conference screen 104 form a uniform linear array and are arranged linearly along a first direction (horizontal direction). Multiple second voice acquisition devices installed on one side of the conference screen 104 form a uniform linear array and are arranged linearly along a second direction (vertical direction).
[0023] The spacing between the voice acquisition devices can be adjusted according to acoustic design requirements to ensure effective acquisition of voice signals from all directions within the conference room. For example, the spacing between horizontally installed first voice acquisition devices can be set to 3 cm, while the spacing between vertically installed second voice acquisition devices can be appropriately widened to 5 cm.
[0024] like Figure 1In the conference system shown, the first voice acquisition device array 101 consists of 8 microphones (first voice acquisition devices) used to acquire first voice array data. The beamforming setting in the first direction (horizontal direction) includes: in the low frequency band (100 Hz - 3500 Hz), using all 8 first voice acquisition devices to perform beamforming in the horizontal direction to form 9 beams in 9 directions (first number), namely [0, 40, 60, 75, 90, 105, 120, 140, 180] degrees; in the high frequency band (3500 Hz - 7500 Hz), using the middle 6 first voice acquisition devices to form the above 9 beams in 9 directions (first number), that is, the first number of beams in the first direction is 9.
[0025] The second voice acquisition device array 102 consists of four microphones (second voice acquisition devices) used to acquire second voice array data. The beamforming setting in the second direction (vertical direction) includes: in both low and high frequency bands, all four second voice acquisition devices are used to perform beamforming in the vertical direction, forming beams in seven directions (second quantity), namely [0, 50, 70, 90, 110, 130, 180] degrees, and the second quantity of the second direction beams is seven.
[0026] It should be noted that, in other embodiments, the number of the first voice acquisition device, the number of the second voice acquisition device, the first number of the first directional beams, the second number of the second directional beams, and the number and location selection of microphones used for beamforming can be designed by those skilled in the art according to actual circumstances. Without departing from the principles of this application, all such modified or replaced technical solutions will fall within the protection scope of this application.
[0027] Please refer to the appendix for further details. Figure 2 ,in, Figure 2 (a) is a schematic flowchart of a speech enhancement method according to an embodiment of this application. Figure 2 (b) is a schematic diagram of the structure of an AI model according to an embodiment of this application.
[0028] like Figure 2As shown in (a), the overall approach of this embodiment is to perform beamforming (first beamforming) on the first and second speech array data respectively, to obtain full-frequency beam data in the first direction corresponding to each preset beam angle in the first direction and full-frequency beam data in the second direction corresponding to each preset beam angle in the second direction. After feature extraction, the full-frequency beam data in the first and second directions is input into a trained AI model to detect the optimal beam angle (target angle in the first and second directions) in real time. Then, based on the optimal beam angle obtained by the AI model, beamforming (second beamforming) is performed on the area array data composed of the first and second speech array data, thereby obtaining enhanced target speech data with better quality.
[0029] As an example, the beamforming method can employ a delay-sum beamforming algorithm to weight the voice data acquired by the first voice acquisition device array and / or the second voice acquisition device array, and extract the voice data within a set beam direction range.
[0030] Please refer to the appendix for further details. Figure 3 and combined Figure 1 and Figure 2 This application describes the speech enhancement method. Figure 3 This is a schematic flowchart illustrating the main steps of a speech enhancement method according to an embodiment of this application. Applied to... Figure 1 The conference terminal 101 shown in this application includes the following voice enhancement method: Step S301: Obtain the first voice array data in the first direction and the second voice array data in the second direction; Step S302: According to the beamforming settings in the first direction, beamform the first voice array data to obtain full-frequency beam data in the first direction; Step S303: According to the beamforming settings in the second direction, beamform the second voice array data to obtain full-frequency beam data in the second direction; Step S304: Based on the full-frequency beam data in the first direction and the full-frequency beam data in the second direction, obtain the target angle in the first direction and the target angle in the second direction through the trained AI model; Step S305: Based on the first target angle and the second target angle, beamforming is performed on the area array composed of the first voice array data and the second voice array data to obtain enhanced target voice data.
[0031] In step 301, the conference terminal 101 can acquire the first voice array data of the eight microphones in the first voice acquisition device array 101 and the second voice array data of the four microphones in the second voice acquisition device array 102 via wired or wireless connection.
[0032] In step 302, beamforming is performed on the first voice array data according to the pre-beam angle and the first quantity in the first direction to obtain the first direction full-frequency beam data corresponding to each beam direction in the first direction.
[0033] As an example, the first quantity is 9. At this time, the first low-frequency beam data in the first direction full-frequency beam data includes: bf_l0, bf_l1, ..., bf_l8, and the first high-frequency beam data includes: bf_h0, bf_h1, ..., bf_h8.
[0034] In step 303, beamforming is performed on the second voice array data according to the pre-beam angle and the second quantity in the second direction to obtain the second direction full-frequency beam data corresponding to each beam direction in the second direction.
[0035] As an example, the second quantity is 7. At this time, the second low-frequency beam data in the second direction full-frequency beam data includes: bf_l9, bf_l10, ..., bf_l15, and the second high-frequency beam data includes: bf_h9, bf_h10, ..., bf_h15.
[0036] Before inputting the trained AI model, it is usually necessary to preprocess the first-direction full-frequency beam data and the second-direction full-frequency beam data to facilitate AI model processing. Specifically, the spectral characteristics of the first-direction full-frequency beam data and the second-direction full-frequency beam data are obtained respectively. As an example, the relevant spectral characteristics can be obtained using short-time Fourier transform.
[0037] The spectral characteristics of the first low-frequency beam data include: BF_l0, BF_l1, ..., BF_l8, and the spectral characteristics of the first high-frequency beam data include: BF_h0, BF_h1, ..., BF_h8.
[0038] The spectral characteristics of the second low-frequency beam data include: BF_l9, BF_l10, ..., BF_l15, and the spectral characteristics of the second high-frequency beam data include: BF_h9, BF_h10, ..., BF_h15.
[0039] Furthermore, the logarithmic magnitude of each spectral feature is obtained as the model input spectral feature of the AI model. For example, FEA = log(abs(input_spec) + (1e-7)).
[0040] The model input spectral features of the first low-frequency beam data include: FEA_BF_l0, FEA_BF_l1, ..., FEA_BF_l8, and the model input spectral features of the first high-frequency beam data include: FEA_BF_h0, FEA_BF_h1, ..., FEA_BF_h8.
[0041] The model input spectral features of the second low-frequency beam data include: FEA_BF_l9, FEA_BF_l10, ..., FEA_BF_l15, and the model input spectral features of the second high-frequency beam data include: FEA_BF_h9, FEA_BF_h10, ..., FEA_BF_h15.
[0042] At this time, the omnidirectional low-frequency model input spectrum features include the model input spectrum features of the first low-frequency beam data and the model input spectrum features of the second low-frequency beam data, specifically including: FEA_BF_l0, FEA_BF_l1, ..., FEA_BF_l15.
[0043] The omnidirectional high-frequency model input spectral features include the model input spectral features of the first high-frequency beam data and the model input spectral features of the second high-frequency beam data, specifically including: FEA_BF_h0, FEA_BF_h1, ..., FEA_BF_h15.
[0044] Preferably, the feature data input to the AI model may also include voice data from one or more first voice acquisition devices and / or second voice acquisition devices, which are used as comparison voice data to further improve the accuracy of the AI model output.
[0045] Comparing speech data also requires performing the aforementioned operations such as spectral feature extraction and logarithmic amplitude. As an example, we select the signal data MIC0 from microphone number 0. The low-frequency spectral feature used as the model input for the AI model is Fea_MIC0_l, and the high-frequency spectral feature is Fea_MIC0_h.
[0046] At this point, the low-frequency spectral features of the model input include FEA_BF_l0, FEA_BF_l1, ..., FEA_BF_l15, Fea_MIC0_l; the high-frequency spectral features of the model input include FEA_BF_h0, FEA_BF_h1, ..., FEA_BF_h15, Fea_MIC0_h.
[0047] It should be noted that whether to add comparative speech data to the input data of the AI model, the amount of comparative speech data, and the position of the microphone, etc., can be selected by those skilled in the art based on the actual situation. Furthermore, the spectrum calculation method, amplitude calculation method, etc., can also be selected by those skilled in the art based on the actual situation. Without departing from the principles of this application, all such modified or replaced technical solutions will fall within the protection scope of this application.
[0048] In step S304, omnidirectional full-frequency beam features are obtained through the omnidirectional full-frequency beam feature extraction network 20, wherein the omnidirectional full-frequency beam feature extraction network 20 includes a first feature extraction network, a second feature extraction network and a third feature extraction network.
[0049] Specifically, the low-frequency spectral features (FEA_BF_l0, FEA_BF_l1, ..., FEA_BF_l15, Fea_MIC0_l) of the model are input into the first feature extraction network to obtain omnidirectional low-frequency beam features.
[0050] The high-frequency spectral features (FEA_BF_h0, FEA_BF_h1, ..., FEA_BF_h15, Fea_MIC0_h) of the model are input into the second feature extraction network to obtain omnidirectional high-frequency beam features.
[0051] Both the first and second feature extraction networks are constructed based on convolutional neural networks. As an example, the first and second feature extraction networks have the same structure, both including two levels of convolutional neural networks. Each convolutional layer contains [12, 16] channels, with a kernel size of 4×5. Downsampling is performed along the frequency dimension, with a stride of (1, 2) per level, and the ReLU activation function is used. Through the first and second feature extraction networks, amplitude difference information between different beams in the low / high frequency range can be obtained, thereby acquiring spatial features.
[0052] The omnidirectional low-frequency beam features and omnidirectional high-frequency beam features are combined (e.g., channel dimension splicing), and the combined feature data is input into the third feature extraction network to obtain the omnidirectional full-frequency beam features.
[0053] The third feature extraction network is also built upon a convolutional neural network. As an example, the third feature extraction network consists of two convolutional neural networks, each containing [24, 32] channels, with a kernel size of 4×5. Downsampling is performed along the frequency dimension, with a stride of (1, 2) per level, and the ReLU activation function is used. This third feature extraction network better integrates high- and low-frequency features.
[0054] In this application, a simplified attention mechanism is used to locate the direction of the target speech, that is, the first direction target angle and the second direction target angle are obtained through the first attention network 30 and the second attention network 40, respectively.
[0055] Specifically, the model corresponding to the full-frequency beam data in the first direction is input into the first full-frequency spectrum features (FEA_BF_l0, FEA_BF_l1, ..., FEA_BF_l8, FEA_BF_h0, FEA_BF_h1, ..., FEA_BF_h8) and then into the first Key feature acquisition network to obtain the first Key feature; the omnidirectional full-frequency beam features are input into the first Query feature acquisition network to obtain the first Query feature; the inner product of the first Key feature and the first Query feature is processed by the softmax function to obtain the first weighted value corresponding to each beam angle in the first direction. The magnitude of the first weighted value represents the probability that the target speech comes from that beam, and the beam angle with the largest first weighted value is the target angle in the first direction.
[0056] The model corresponding to the second-direction full-frequency beam data is input into the second full-frequency spectrum features (FEA_BF_l9, FEA_BF_l10, ..., FEA_BF_l15, FEA_BF_h9, FEA_BF_h10, ..., FEA_BF_h15) and then into the second Key feature acquisition network to obtain the second Key features. The omnidirectional full-frequency beam features are input into the second Query feature acquisition network to obtain the second Query features. The inner product of the second Key features and the second Query features is processed by the softmax function to obtain the second weighted value corresponding to each beam angle in the second direction. The magnitude of the second weighted value represents the probability that the target speech comes from that beam. The beam angle with the largest second weighted value is the target angle in the second direction.
[0057] As an example, both the first key feature acquisition network and the second key feature acquisition network are built based on convolutional neural networks, while both the second query feature acquisition network and the second query feature acquisition network are built based on fully connected neural networks.
[0058] In step S305, the first direction target angle and the second direction target angle are passed to the second beamforming to perform beamforming on the area array composed of the first speech array data and the second speech array data to obtain enhanced target speech data. At this time, the enhanced target speech data can be either time domain data or spectral feature data.
[0059] In this application, by utilizing the efficient data processing and learning capabilities of AI models, it is possible to quickly and accurately track changes in the direction of the target speech, switch beam directions in a timely manner, avoid or reduce the suppression of the speech of new speakers in the conference, thereby obtaining clear speech and improving the user experience.
[0060] In another embodiment, such as Figure 4 As shown, where, Figure 4 (a) is a schematic flowchart of a speech enhancement method according to another embodiment of this application. Figure 4 (b) is a schematic diagram of the structure of an AI model according to another embodiment of this application.
[0061] like Figure 4 As shown in (a), the overall concept of another embodiment of this application is that, in Figure 2 Based on the enhanced target speech data obtained in the embodiment shown in (a), the AI model adds a MASK feature acquisition network; through the target enhancement MASK features output by the AI model, the enhanced target speech data is further enhanced, thereby improving the quality of the target speech.
[0062] Please refer to the appendix for further details. Figure 5 and combined Figure 4 This application describes another embodiment of a speech enhancement method. Figure 5 This is a schematic flowchart illustrating the main steps of a speech enhancement method according to another embodiment of this application. Applied to... Figure 1 The conference terminal 101 shown, in another embodiment, includes a voice enhancement method comprising: Step S501: Obtain the first voice array data in the first direction and the second voice array data in the second direction; Step S502: According to the beamforming settings in the first direction, beamform the first voice array data to obtain full-frequency beam data in the first direction; Step S503: According to the beamforming settings in the second direction, beamform the second voice array data to obtain full-frequency beam data in the second direction; Step S504: Based on the full-frequency beam data in the first direction and the full-frequency beam data in the second direction, the target angle in the first direction, the target angle in the second direction, and the target enhancement MASK features are obtained through the trained AI model; Step S505: Based on the first direction target angle and the second direction target angle, beamforming is performed on the area array composed of the first speech array data and the second speech array data to obtain enhanced target speech data; Step S506: Based on the target enhancement MASK features and the enhanced target speech data, obtain further enhanced target speech data.
[0063] The method of step S501 is the same as the scheme of step S301 above, the method of step S502 is the same as the scheme of step S302 above, the method of step S503 is the same as the scheme of step S303 above, and the method of obtaining the first target angle and the second target angle in step S503 is the same as the method of obtaining the first target angle and the second target angle in step S303 above, and will not be repeated here.
[0064] like Figure 4As shown in (b), the AI model also includes a MASK feature acquisition network. In step S504, the omnidirectional full-frequency beam features are input into the MASK feature acquisition network to obtain target-enhanced MASK features. As an example, the MASK feature acquisition network is constructed based on a gated recurrent neural network and a fully connected neural network, wherein the GRU has 256 hidden units and 2 layers.
[0065] The method of step S505 is the same as that of step S305 mentioned above, and will not be repeated here.
[0066] In step S506, the target enhanced MASK feature and the enhanced target speech data (at this time, the enhanced target speech data should be spectral feature data, and it needs to be multiplied with the aforementioned spectral feature calculation method using the same parameters) are multiplied to obtain further enhanced target speech data.
[0067] The enhanced target speech data is then subjected to a short-time inverse Fourier transform to obtain the time-domain target speech signal. It is then further processed and optimized as needed, such as filtering and gain adjustment, before the speech signal is output to the speakers of the conference system for participants to listen to.
[0068] In this application, target speech data is further enhanced by target-enhanced MASK features, resulting in clearer and more accurate speech, which greatly improves the user experience.
[0069] It should be noted that although the steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of this application, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent to the technical solutions described in this application and therefore will also fall within the protection scope of this application.
[0070] Those skilled in the art will understand that all or part of the processes in the method of the above-described embodiment can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0071] The relevant user personal information that may be involved in the various embodiments of this application is processed in strict accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, based on the reasonable purpose of the business scenario, and includes personal information that users actively provide or that is generated as a result of using the product / service, as well as personal information obtained with user authorization.
[0072] The personal information processed in this application will vary depending on the specific product / service scenario and will be subject to the specific scenario in which the user uses the product / service. This may involve the user's account information, device information, or other related information. This application will treat the user's personal information and its processing with the utmost diligence.
[0073] This application attaches great importance to the security of users' personal information and has taken reasonable and feasible security protection measures that comply with industry standards to protect users' information and prevent unauthorized access, disclosure, use, modification, damage or loss of personal information.
[0074] The technical solution of this application has been described above with reference to one embodiment shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.
Claims
1. A speech enhancement method, characterized in that, The method includes: Acquire first voice array data in the first direction and second voice array data in the second direction; According to the beamforming settings in the first direction, the first voice array data is beamformed to obtain full-frequency beam data in the first direction; According to the beamforming settings in the second direction, the second voice array data is beamformed to obtain full-frequency beam data in the second direction; Based on the full-frequency beam data in the first direction and the full-frequency beam data in the second direction, the target angle in the first direction and the target angle in the second direction are obtained through a trained AI model; Based on the target angle in the first direction and the target angle in the second direction, beamforming is performed on the area array composed of the first voice array data and the second voice array data to obtain enhanced target voice data.
2. The speech enhancement method according to claim 1, characterized in that, The AI model includes an omnidirectional full-frequency beam feature extraction network, a first attention network, and a second attention network. "Based on the first-direction full-frequency beam data and the second-direction full-frequency beam data, the trained AI model is used to obtain the first-direction target angle and the second-direction target angle," which includes: Based on the first direction full-frequency beam data and the second direction full-frequency beam data, the omnidirectional full-frequency beam features are obtained through the omnidirectional full-frequency beam feature extraction network; Based on the omnidirectional full-frequency beam characteristics and the first-direction full-frequency beam data, the first-direction target angle is obtained through the first attention network; Based on the omnidirectional full-frequency beam characteristics and the second-direction full-frequency beam data, the second-direction target angle is obtained through the second attention network; The first directional full-frequency beam data includes a first number of first low-frequency beam data and a first number of first high-frequency beam data, and the second directional full-frequency beam data includes a second number of second low-frequency beam data and a second number of second high-frequency beam data.
3. The speech enhancement method according to claim 2, characterized in that, The omnidirectional full-frequency beam feature extraction network includes a first feature extraction network, a second feature extraction network, and a third feature extraction network. The method for acquiring the omnidirectional full-frequency beam features includes: Based on the first low-frequency beam data and the second low-frequency beam data, omnidirectional low-frequency beam features are obtained through the first feature extraction network. Based on the first high-frequency beam data and the second high-frequency beam data, omnidirectional high-frequency beam features are obtained through the second feature extraction network; Based on the omnidirectional low-frequency beam features and the omnidirectional high-frequency beam features, the omnidirectional full-frequency beam features are obtained through the third feature extraction network.
4. The speech enhancement method according to claim 3, characterized in that, The first feature extraction network, the second feature extraction network, and the third feature extraction network are all constructed based on convolutional networks.
5. The speech enhancement method according to claim 2, characterized in that, The first attention network includes a first query feature acquisition network, a first key feature acquisition network, and a first weight acquisition network. The method for obtaining the target angle in the first direction includes: Based on the full-frequency beam data in the first direction, the first query feature is obtained through the first query feature acquisition network; Based on the full-frequency beam data in the first direction, the first key feature is obtained through the first key feature acquisition network; Based on the first Query feature and the first Key feature, the first weighted value corresponding to each beam angle in the first direction is obtained through the first weight acquisition network. The beam angle corresponding to the first weighted value with the largest value is selected as the target angle in the first direction.
6. The speech enhancement method according to claim 5, characterized in that, The second attention network includes a second query feature acquisition network, a second key feature acquisition network, and a second weight acquisition network. The method for obtaining the second direction target angle includes: Based on the second direction full-frequency beam data, the second query feature is obtained through the second query feature acquisition network; Based on the second direction full-frequency beam data, the second key feature is obtained through the second key feature acquisition network; Based on the second Query feature and the second Key feature, the second weighting network is used to obtain the second weighting value corresponding to each beam angle in the second direction; The beam angle corresponding to the second weighted value with the largest value is selected as the target angle in the second direction.
7. The speech enhancement method according to claim 6, characterized in that, Both the second query feature acquisition network and the second query feature acquisition network are constructed based on fully connected neural networks; Both the second key feature acquisition network and the second key feature acquisition network are constructed based on convolutional networks; Both the second weight acquisition network and the second weight acquisition network are constructed based on the Softmax function.
8. The speech enhancement method according to any one of claims 2 to 7, characterized in that, The AI model also includes a MASK feature acquisition network, and the method further includes: Based on the omnidirectional full-frequency beam characteristics, the target enhanced MASK characteristics are obtained through the MASK feature acquisition network. Based on the target enhanced MASK features and the enhanced target speech data, further enhanced target speech data is obtained.
9. The speech enhancement method according to claim 8, characterized in that, The MASK feature acquisition network is constructed based on gated recurrent neural networks and fully connected neural networks.
10. A conference system, characterized in that, The conference system includes: Multiple first voice acquisition devices, wherein the multiple first voice acquisition devices form a uniform linear array or a non-uniform linear array, and are arranged linearly along a first direction; Multiple second voice acquisition devices, wherein the multiple second voice acquisition devices form a uniform linear array or a non-uniform array, and are arranged linearly along a second direction; A conference terminal configured to implement the speech enhancement method according to any one of claims 1 to 9.