Speech processing method, apparatus, device, storage medium, and program product
By acquiring spatial and visual information through an omnidirectional microphone array and a camera, and combining the direction of arrival (DOA) and lip-reading model to generate an audiovisual mask, the problem of inaccurate speech separation in multi-speaker conferences is solved, achieving efficient speech recognition and conference decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-03-05
- Publication Date
- 2026-07-14
AI Technical Summary
In multi-speaker conferences, existing technologies suffer from inaccurate speech separation due to factors such as room reverberation, background noise, and similar acoustic characteristics of speakers. This affects the speech recognition module's ability to accurately identify the content of each speaker, leading to information omissions and errors in conference decision-making.
By using an omnidirectional microphone array and a camera to acquire the speaker's spatial and visual information, a spatial time-frequency mask is calculated using the direction of arrival angle, and a visual time-frequency mask is generated by combining it with a lip-reading model. These are then fused to form an audiovisual mask to separate independent speech signals.
It improves the accuracy and robustness of speech separation, ensuring that the speech recognition module can accurately identify spoken content, thereby improving the efficiency of meeting debriefing and the effectiveness of decision execution.
Smart Images

Figure CN122392557A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video, and more particularly to a voice processing method, apparatus, device, storage medium, and program product. Background Technology
[0002] In meeting scenarios, the application of speech separation technology is crucial for improving communication efficiency and ensuring accurate recording of meeting content. Effectively separating the speech signals of overlapping speakers provides clearer input for subsequent speech recognition tasks, thereby improving accuracy. However, in the complex acoustic environment of multi-speaker meetings, factors such as room reverberation, background noise, and similar acoustic characteristics of speakers limit the effectiveness of speech separation, leading to inaccurate separation. This results in speech recognition modules failing to accurately identify the content of each speaker, further causing information omissions during meeting debriefings, impacting the execution of meeting decisions, and even leading to business losses. Summary of the Invention
[0003] This application provides a speech processing method, apparatus, device, storage medium, and program product to solve the technical problem of inaccurate speech separation in multi-speaker conferences.
[0004] In a first aspect, this application provides a speech processing method, comprising: receiving mixed speech signals from multiple speakers in a conference using an omnidirectional microphone array; and capturing facial visual information of at least one speaker in the conference while speaking using a camera; wherein each speaker is located in a different spatial position in the conference room;
[0005] Based on the direction of arrival angle of each speaker relative to the microphone array, the spatial time-frequency mask of each speaker is obtained; for each speaker in at least one speaker, speech synthesis is performed on their facial visual information using a lip-reading model to obtain the visual time-frequency mask of each speaker.
[0006] The spatial time-frequency mask and visual time-frequency mask of any speaker are fused to form the speaker's audiovisual mask;
[0007] Apply the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
[0008] Secondly, this application provides a voice processing apparatus, comprising:
[0009] The acquisition module is used to receive mixed speech signals from multiple speakers in a conference using an omnidirectional microphone array; and to acquire facial visual information of at least one speaker in the conference while speaking using a camera; wherein each speaker is located in a different spatial position in the conference room;
[0010] The processing module is used to obtain the spatial time-frequency mask of each speaker based on the direction of arrival angle of each speaker relative to the microphone array; for each speaker among at least one speaker, speech synthesis is performed on their facial visual information using a lip-reading model to obtain the visual time-frequency mask of each speaker.
[0011] The fusion module is used to fuse the spatial time-frequency mask and visual time-frequency mask of any speaker to form the speaker's audiovisual mask;
[0012] The separation module is used to apply the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
[0013] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0014] The memory stores computer-executed instructions;
[0015] The processor executes computer execution instructions stored in the memory to implement the method as described in any of the first aspects.
[0016] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any of the first aspects.
[0017] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the first aspects.
[0018] The speech processing method, apparatus, device, storage medium, and program products provided in this application acquire spatial and visual information of different speakers using an omnidirectional microphone array and a camera. Based on the direction of arrival angle of each speaker relative to the microphone array, a spatial time-frequency mask for each speaker is obtained; the speaker's speech is synthesized using visual information to obtain a visual time-frequency mask for each speaker. The spatial time-frequency mask and visual time-frequency mask of any speaker are fused to form the speaker's audiovisual mask; applying the audiovisual mask of any speaker to a mixed speech signal yields the speaker's independent speech signal. The method of this application eliminates the problem of poor speech separation accuracy caused by complex meeting environments because visual information is unaffected by reverberation, background noise, and interfering speakers in the audio modality, thus improving the robustness of the speech separation method. The fusion of the spatial time-frequency mask and the visual time-frequency mask improves the confidence of the separation mask, thereby improving the accuracy of speech separation. Inputting the separated speech signal into a speech recognition module allows for accurate identification of the speaker's speech content, improving the ability of participants to obtain complete information during meeting debriefing, ensuring the effective execution of meeting decisions, and thus promoting business success. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] Figure 1 This is a schematic diagram illustrating an application scenario involved in an embodiment of this application;
[0021] Figure 2 A schematic flowchart of a speech processing method provided in an embodiment of this application;
[0022] Figure 3 A flowchart illustrating the second speech signal processing method provided in this application embodiment;
[0023] Figure 4 A flowchart illustrating the second speech signal processing method provided in this application embodiment;
[0024] Figure 5 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0026] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0029] It should be noted that the speech processing methods, apparatus, devices, storage media, and program products provided in this application can be used in the audio and video field, as well as in any other field. The application fields of the speech processing methods, apparatus, devices, storage media, and program products in this application are not limited.
[0030] Figure 1 This is a schematic diagram illustrating an application scenario involved in an embodiment of this application, such as... Figure 1 As shown, the specific application scenario of this application is multi-speaker conferences.
[0031] In meeting scenarios, the application of speech separation technology is of great significance for improving communication efficiency and ensuring accurate recording of meeting content. By effectively separating the speech signals of overlapping speakers, clearer input can be provided for subsequent speech recognition tasks, thereby improving the accuracy of speech recognition. This is an important guarantee for meeting participants to improve communication efficiency and review and understand meeting content.
[0032] However, in the complex acoustic environment of multi-speaker conferences, factors such as indoor reverberation, background noise, and similar acoustic characteristics of speakers can limit the effectiveness of speech separation, resulting in technical problems of inaccurate speech separation. This makes it impossible for the speech recognition module to accurately identify the content of each speaker, which in turn leads to information omissions by participants during meeting debriefing, affects the execution of meeting decisions, and may even cause business losses.
[0033] Therefore, improving the accuracy of speech separation in multi-speaker conferences has become an urgent technical problem to be solved.
[0034] The speech processing method, apparatus, device, storage medium, and program products provided in this application aim to solve the aforementioned technical problems of the prior art. Spatial and visual information of different speakers is acquired using an omnidirectional microphone array and a camera. A spatial time-frequency mask for each speaker is obtained based on the direction of arrival angle of each speaker relative to the microphone array; the speaker's speech is synthesized using the visual information to obtain a visual time-frequency mask for each speaker. The spatial time-frequency mask and visual time-frequency mask of any speaker are fused to form the speaker's audiovisual mask; applying the audiovisual mask of any speaker to a mixed speech signal yields the speaker's independent speech signal. The method of this application eliminates the problem of poor speech separation accuracy caused by complex meeting environments because the visual information is unaffected by reverberation, background noise, and interfering speakers in the audio modality, thus improving the robustness of the speech separation method. By fusing spatial time-frequency masks and visual time-frequency masks, the confidence level of the separation mask is improved, thereby enhancing the accuracy of speech separation. The separated speech signal is then input into the speech recognition module, which can accurately identify the speaker's content, improve the ability of participants to obtain complete information during meeting debriefing, ensure the effective execution of meeting decisions, and thus promote business success.
[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0036] The execution subject of this application embodiment can be a voice processing device or an electronic device equipped with a voice processing device. This application embodiment uses an electronic device as an example for illustration.
[0037] Figure 2 This is a flowchart illustrating a speech processing method provided in an embodiment of this application. Figure 2 As shown, the method may include, for example, the following steps:
[0038] S201. An electronic device uses an omnidirectional microphone array to receive mixed speech signals from multiple speakers in a meeting; and uses a camera to capture facial visual information of at least one speaker in the meeting while speaking; wherein each speaker is located in a different spatial position in the meeting room.
[0039] An omnidirectional microphone can be any microphone capable of receiving sound uniformly from all directions. An omnidirectional microphone array can be an array of multiple omnidirectional microphones. The omnidirectional microphones in the array can be arranged in a geometric layout, such as linear, circular, or rectangular. Omnidirectional microphone arrays can effectively capture the voices of all speakers in a conference environment, facilitating the acquisition of comprehensive audio data in complex acoustic environments and providing a foundation for subsequent speech separation and speech recognition.
[0040] The speaker's facial visual information during speech can be, for example, a video signal of the speaker speaking, which can be a silent time-frequency signal, and can include the movement and shape changes of the speaker's lips.
[0041] The camera can be any device capable of capturing dynamic video. There can be one or more cameras; for example, one camera can be used to capture the facial visual information of a single speaker; or one camera can be used to capture the facial visual information of multiple speakers; through computer vision and machine learning techniques, the facial visual information of multiple speakers can be separated, so that each facial visual information corresponds to a single speaker.
[0042] Since each speaker is in a different spatial location within the conference room, a microphone array can be used to determine the spatial location of each speaker.
[0043] S202. The electronic device obtains the spatial time-frequency mask of each speaker based on the direction of arrival angle of each speaker relative to the microphone array; for each speaker among at least one speaker, speech synthesis is performed on their facial visual information using a lip-reading model to obtain the visual time-frequency mask of each speaker.
[0044] The direction of arrival (DOA) is the directional angle of a sound source (speaker) relative to the receiving device (microphone array). It is a key parameter for sound source localization, determining the direction from which sound arrives at the microphone array. The ODA can be calculated using delay estimation, beamforming, or high-resolution algorithms.
[0045] For example, the mixed speech signal received by the microphone array can be converted to the time-frequency domain by using short-time Fourier transform, so as to obtain the speech signal received by each microphone in different times and frequencies.
[0046] A spatial time-frequency mask can be, for example, a matrix with the same dimensions as the time-frequency representation, where each element typically has a value between 0 and 1. Using microphone array signal processing techniques, the signal strength from different directions at each time-frequency point can be estimated. The spatial time-frequency mask utilizes this directional information to generate a matrix where each element represents the signal strength from a specific direction at a specific time and frequency. The spatial time-frequency mask can be used to highlight the target speaker's signal while suppressing interference from other directions. By applying a spatial mask to the time-frequency representation, signals from the target direction can be amplified, thereby separating the speech signal from the target speaker.
[0047] Lip-reading models can be any technique capable of recognizing speech information by analyzing lip movements. Using a lip-reading model to synthesize speech from a speaker's facial visual information yields synthesized speech, from which a visual time-frequency mask for each speaker can be obtained. This visual time-frequency mask can be, for example, a matrix of the same dimension as the time-frequency representation, where each element typically has a value between 0 and 1. Visual time-frequency masks highlight the mixed speech signal components associated with visual activity. Applying visual video masks to the time-frequency representation enhances signals relevant to the target speaker, thereby separating the speech signal from the target speaker.
[0048] S203. The electronic device fuses the spatial time-frequency mask and visual time-frequency mask of any speaker to form the speaker's audiovisual mask;
[0049] Through fusion, the generated audiovisual mask combines the complementarity of spatial and visual information, enabling more accurate identification and separation of speech signals from different speakers.
[0050] For example, the speaker's spatial time-frequency mask value can be used as the base of an exponentiation operation, and the spline-interpolated visual time-frequency mask value can be used as the exponent, performing a power-law transformation to form the speaker's audiovisual mask. This audiovisual mask can be used as a separation mask. Spline interpolation allows the use of piecewise polynomials to connect the known data points of the visual time-frequency mask with smooth curves, smoothing and enhancing the quality of the visual time-frequency mask, making it more suitable for fusion with the spatial time-frequency mask. By applying a power-law transformation to both the spatial time-frequency mask and the spline-interpolated visual time-frequency mask, the higher-confidence components of the mask can be enhanced while suppressing the lower-confidence components, thus more effectively highlighting the target signal.
[0051] If the visual time-frequency mask value is large, it will reduce the confidence of the original spatial mask, thus decreasing its mask value; if the visual time-frequency mask value is small, it will increase the confidence of the original spatial mask, thus increasing its mask value; if the visual time-frequency mask confidence is low (i.e., the mask value is around 0.5), the original value of the spatial mask will not be changed after power-law transformation. By using a late-stage fusion method with linear interpolation through power-law transformation to fuse the spatial time-frequency mask and the visual time-frequency mask, the confidence of the separation mask can be increased. Using a separation mask with high confidence to separate independent speech from different speakers in a mixed speech signal can improve the accuracy of speech signal separation.
[0052] S204. The electronic device applies the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
[0053] The mixed speech signal can be converted to the time-frequency domain, and the generated audiovisual mask can be applied to the time-frequency representation of the mixed speech signal to extract the time-frequency representation of the speaker's speech signal.
[0054] In one example, an audiovisual mask is multiplied by the mixed speech signal to obtain the speaker's individual speech signal. For instance, the audiovisual mask is a matrix of the same dimension as the time-frequency representation of the mixed signal, where each element is typically between 0 and 1, representing the weight of the target speaker's signal at that time-frequency point. Multiplying the audiovisual mask point-by-point by the time-frequency representation of the mixed speech signal means that for each element in the time-frequency representation matrix, the corresponding audiovisual mask value is used for weighting; the resulting new time-frequency representation matrix after the point-by-point multiplication shows that the target speaker's signal is enhanced, while other speakers and noise are suppressed. This method is easy to implement.
[0055] In one example, the time-frequency representation of the mixed speech signal and the audiovisual mask can be input into a deep learning model to learn how to separate the speech signal of a single speaker from the mixed speech signal. This method can adapt to new data and environmental changes, and can respond quickly when the signal changes.
[0056] In one example, an adaptive filter can be designed using an audio mask as a reference signal to dynamically adjust the filter system in order to obtain the speaker's independent speech signal.
[0057] The time-frequency representation of the speaker's speech signal after separation can be converted back to the time domain through inverse short-time Fourier transform, thus obtaining the speaker's independent speech signal.
[0058] Subsequently, the speaker's independent speech signal can be input into the speech recognition module, which outputs the spoken content. This allows different recognized speech contents to correspond to different speaker identities.
[0059] Evaluation metrics for speech separation performance may include at least one of the following: Source-to-Distortion Ratio (SDR), Source-to-Interference Ratio (SIR), Word Error Rate (WER), etc.
[0060] In summary, the speech signal processing method provided in this application acquires spatial and visual information of different speakers using an omnidirectional microphone array and a camera. Based on the direction of arrival angle of each speaker relative to the microphone array, a spatial time-frequency mask for each speaker is obtained. The speaker's speech is synthesized using visual information to obtain a visual time-frequency mask for each speaker. The spatial time-frequency mask and visual time-frequency mask of any speaker are fused to form the speaker's audiovisual mask. Applying the audiovisual mask of any speaker to the mixed speech signal yields the speaker's independent speech signal. This method eliminates the problem of poor speech separation accuracy caused by complex meeting environments because visual information is unaffected by reverberation, background noise, and interfering speakers in the audio modality, thus improving the robustness of the speech separation method. The fusion of the spatial and visual time-frequency masks increases the confidence of the separation mask, thereby improving the accuracy of speech separation. Inputting the separated speech signal into a speech recognition module allows for accurate identification of the speaker's content, improving the ability of participants to obtain complete information during meeting debriefings, ensuring the effective execution of meeting decisions, and ultimately promoting business success.
[0061] Figure 3 A flowchart illustrating the second speech signal processing method provided in this application embodiment is shown below. Figure 3 As shown, in Figure 2 Based on the embodiments, this paper explains how to obtain the spatial time-frequency mask of each speaker based on the direction of arrival angle of each speaker relative to the microphone array.
[0062] S301. The electronic device calculates the time difference between the arrival of the voice signals of different speakers in the conference room at different microphones, and calculates the direction of arrival angle of each speaker relative to the microphone array.
[0063] When a speaker emits sound, the sound arrives at the various microphones in the microphone array at different times. By measuring these differences in arrival time, the direction of the sound can be inferred. Using this time difference information, combined with the geometric layout of the microphone array, geometric triangulation or signal processing algorithms can be used to calculate the direction of arrival (DOA) of the sound, thereby determining the direction of each speaker.
[0064] S302. The electronic device calculates the spatial spectrum of each speaker in different directions based on the direction of arrival angle of each speaker relative to the microphone array.
[0065] A spatial spectrum represents the energy distribution of a speaker's voice in various directions within space, reflecting the spatial characteristics of different speakers. For example, beamforming techniques can be used to scan different directions in space and calculate the signal intensity in each direction. Alternatively, the covariance matrix of the sound signal can be analyzed to generate a high-resolution spatial spectrum, displaying the precise location and intensity of the speaker's voice in space.
[0066] S303. The electronic device compares the spatial spectrum amplitude of different speakers at each time and frequency point to generate a spatial time and frequency mask for each speaker.
[0067] For each time-frequency point, beamforming or high-resolution algorithms are used to calculate the spatial spectrum from different directions, providing an energy distribution in each direction. At each time-frequency point, the spatial spectral amplitudes in different speaker directions are compared. The direction with the largest spatial spectral amplitude is identified, and the speaker corresponding to that direction is considered the major contributor at that time-frequency point. Based on the comparison results of the spatial spectral amplitudes, a spatial time-frequency mask can be generated for each speaker. The mask value is typically between 0 and 1, representing the signal weight at that time-frequency point.
[0068] Furthermore, at any given time-frequency point, a mask value of 1 can be assigned to the speaker with the largest spatial spectral amplitude, and a mask value of 0 can be assigned to other speakers to generate a spatial time-frequency mask for each speaker.
[0069] If the value of this time-frequency point is 1, it indicates that the speaker with the largest spatial spectrum amplitude dominates this time-frequency point. Other speakers and background noise are suppressed, thereby improving the signal-to-noise ratio of the target speech signal, effectively reducing interference from other speakers and signal aliasing, and improving the quality of speech separation.
[0070] In summary, the speech signal processing method provided in this application calculates the time difference of arrival of speech signals from different speakers in a conference room using a microphone array, and calculates the direction of arrival angle based on this information. This allows for precise location of each speaker, and then the generation of a spatial time-frequency mask for each speaker can highlight the signals of speakers located in different spatial positions in the time-frequency domain while suppressing interference from other directions. This method can be used in conjunction with visual time-frequency masks to further improve speech separation and the overall performance of the system.
[0071] Figure 4 A flowchart illustrating the second speech signal processing method provided in this application embodiment is shown below. Figure 4 As shown, in Figure 2Based on the embodiments, this paper explains how to use a lip-reading model to synthesize speech based on facial visual information of each speaker in at least one speaker, and how to obtain the visual time-frequency mask of each speaker.
[0072] S401. The electronic device uses a neural network-based lip-reading model to generate synthesized lip-reading speech from the facial visual information of each speaker during speech.
[0073] The method inputs the speaker's facial visual information into a lip-reading model, which captures the movement and shape changes of the speaker's lips. The lip-reading model extracts the shape, movement trajectory, and other facial features of the lips. Through machine learning algorithms, such as deep learning, the lip-reading model maps these visual features onto the speech content, thereby recognizing the words or sentences spoken by the speaker. This process generates synthesized lip-reading speech from the facial visual information of each speaker. The method in this application eliminates the problem of poor speech separation accuracy caused by complex meeting environments because the visual information is unaffected by reverberation, background noise, and interfering speakers in the audio modality, thus improving the robustness of the speech separation method.
[0074] S402. The electronic device compares the synthesized lip-reading speech with the mixed speech signal to determine the proportion of the lip-reading speech in the mixed speech signal;
[0075] For example, features can be extracted from the synthesized lip-reading speech and mixed speech signals using methods such as short-time Fourier transform or Mel-frequency cepstral coefficients, converting them into time-frequency feature representations. The similarity between the synthesized lip-reading speech and the mixed speech signal can be calculated using methods such as cosine similarity, correlation coefficient, or dynamic time warping. By calculating the proportion of the similarity score in the entire mixed speech signal, the percentage of synthesized lip-reading speech in the mixed speech signal can be analyzed.
[0076] S403. The electronic device generates a visual time-frequency mask for each speaker based on the proportion of lip-reading speech in the mixed speech signal.
[0077] Based on the proportion of lip-reading in the mixed speech signal, higher mask values can be assigned to time-frequency points with higher proportions, thus reflecting the speaker's dominance at that time-frequency point. The generated visual video mask can be applied to the mixed speech signal to highlight the target speaker's speech signal and suppress other interference.
[0078] In summary, the speech signal processing method provided in this application converts the speaker's panel visual information into synthesized lip-reading speech by using a neural network-based lip-reading model. By comparing the synthesized lip-reading speech with the mixed speech signal, the speaker's visual time-frequency mask is determined. This method can highlight the target speaker's signal while suppressing other speakers and background noise, thereby improving the signal-to-noise ratio. The visual video mask can be used in conjunction with the spatial time-frequency mask to further improve the speech separation effect and the overall performance of the system.
[0079] Figure 5 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Figure 5 As shown, the device may include, for example, the following modules: acquisition module 501, processing module 502, fusion module 503, and separation module 504;
[0080] Acquisition module 501 is used to receive mixed speech signals from multiple speakers in a meeting using an omnidirectional microphone array; and to acquire facial visual information of at least one speaker in the meeting while speaking using a camera; wherein each speaker is located in a different spatial position in the meeting room;
[0081] Processing module 502 is used to obtain the spatial time-frequency mask of each speaker based on the direction of arrival angle of each speaker relative to the microphone array; and for each speaker in at least one speaker, to use a lip-reading model to perform speech synthesis on their facial visual information to obtain the visual time-frequency mask of each speaker.
[0082] The fusion module 503 is used to fuse the spatial time-frequency mask and the visual time-frequency mask of any speaker to form the speaker's audiovisual mask;
[0083] The separation module 504 is used to apply the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
[0084] One possible implementation is that the processing module 502 is specifically used to calculate the time difference of the voice signals of different speakers in the conference room arriving at different microphones, and calculate the direction of arrival angle of each speaker relative to the microphone array; based on the direction of arrival angle of each speaker relative to the microphone array, calculate the spatial spectrum of each speaker in different directions; at each time-frequency point, compare the spatial spectrum amplitudes of different speakers to generate a spatial time-frequency mask for each speaker.
[0085] One possible implementation is that the processing module 502 is specifically used to assign a mask value of 1 to the speaker with the largest spatial spectral amplitude at any time-frequency point, and assign a mask value of 0 to other speakers, so as to generate a spatial time-frequency mask for each speaker.
[0086] One possible implementation is that the processing module 502 is specifically used to generate synthesized lip-reading speech from the facial visual information of each speaker when speaking using a lip-reading model based on a neural network; compare the synthesized lip-reading speech with the mixed speech signal to determine the proportion of the lip-reading speech in the mixed speech signal; and generate a visual time-frequency mask for each speaker based on the proportion of the lip-reading speech in the mixed speech signal.
[0087] One possible implementation is that the fusion module 503 is specifically used to take the value of the speaker's spatial time-frequency mask as the base of the exponentiation operation and the value of the visual time-frequency mask after spline interpolation as the exponent of the exponentiation operation, and perform a power-law transformation to form the speaker's audiovisual mask.
[0088] One possible implementation is a separation module 504, which is specifically used to perform a dot product operation between the audiovisual mask and the mixed speech signal to obtain the speaker's independent speech signal.
[0089] The voice processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0090] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device may include at least one processor 601 and a memory 602.
[0091] The memory 602 is used to store programs. Specifically, the program may include program code, which includes computer operation instructions.
[0092] The memory 602 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0093] The processor 601 is used to execute computer execution instructions stored in the memory 602 to implement the actions in the foregoing method embodiments. The processor 601 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0094] Optionally, the electronic device may also include a communication interface 603 for communication and interaction with external devices. In specific implementations, if the communication interface 603, memory 602, and processor 601 are implemented independently, the communication interface 603, memory 602, and processor 601 can be interconnected via a bus to complete communication between them.
[0095] Optionally, in a specific implementation, if the communication interface 603, memory 602, and processor 601 are integrated on a single chip, then the communication interface 603, memory 602, and processor 601 can communicate through an internal interface.
[0096] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), and a random access memory (RAM). Specifically, the computer-readable storage medium stores program instructions, which are used to implement the actions of the above-described method implementation.
[0097] This application also provides a computer program product including executable instructions stored in a readable storage medium. At least one processor of an electronic device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the electronic device to perform the actions described in the method embodiments.
[0098] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0099] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A speech processing method, characterized in that, include: Use an omnidirectional microphone array to receive mixed speech signals from multiple speakers in a conference; In addition, a camera is used to capture facial visual information of at least one speaker in the meeting while they are speaking; wherein each speaker is in a different spatial location in the meeting room; Based on the direction of arrival angle of each speaker relative to the microphone array, the spatial time-frequency mask of each speaker is obtained; for each speaker in at least one speaker, speech synthesis is performed on their facial visual information using a lip-reading model to obtain the visual time-frequency mask of each speaker. The spatial time-frequency mask and visual time-frequency mask of any speaker are fused to form the speaker's audiovisual mask; Apply the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
2. The method according to claim 1, characterized in that, The process of obtaining the spatial time-frequency mask for each speaker based on their direction-of-arrival angle relative to the microphone array specifically includes: Calculate the time difference between the arrival times of the speech signals of different speakers in the conference room at different microphones, and calculate the direction of arrival angle of each speaker relative to the microphone array; Based on the direction of arrival angle of each speaker relative to the microphone array, the spatial spectrum of each speaker in different directions is calculated. At each time-frequency point, the spatial spectrum amplitudes of different speakers are compared to generate a spatial time-frequency mask for each speaker.
3. The method according to claim 2, characterized in that, The step of comparing the spatial spectral amplitudes of different speakers at each time-frequency point to generate a spatial time-frequency mask for each speaker specifically includes: At any given time-frequency point, a mask value of 1 is assigned to the speaker with the largest spatial spectral amplitude, and a mask value of 0 is assigned to the other speakers to generate a spatial time-frequency mask for each speaker.
4. The method according to claim 1, characterized in that, For each speaker among at least one speaker, speech synthesis is performed using a lip-reading model on their facial visual information to obtain a visual time-frequency mask for each speaker, specifically including: By using a neural network-based lip-reading model, synthesized lip-reading speech is generated from the facial visual information of each speaker during speech. The synthesized lip-reading speech is compared with the mixed speech signal to determine the proportion of lip-reading speech in the mixed speech signal; Based on the proportion of lip-reading speech in the mixed speech signal, a visual time-frequency mask is generated for each speaker.
5. The method according to claim 1, characterized in that, The process of fusing the spatial time-frequency mask and visual time-frequency mask of any speaker to form the speaker's audiovisual mask specifically includes: The speaker's spatial time-frequency mask is used as the base of the exponentiation operation, and the visual time-frequency mask after spline interpolation is used as the exponent of the exponentiation operation. A power-law transformation is then performed to form the speaker's audiovisual mask.
6. The method according to any one of claims 1 to 5, characterized in that, The step of applying an audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal specifically includes: The audiovisual mask is multiplied by the mixed speech signal to obtain the speaker's independent speech signal.
7. A voice processing device, characterized in that, include: The acquisition module is used to receive mixed speech signals from multiple speakers in a conference using an omnidirectional microphone array; In addition, a camera is used to capture facial visual information of at least one speaker in the meeting while they are speaking; wherein each speaker is in a different spatial location in the meeting room; The processing module is used to obtain the spatial time-frequency mask of each speaker based on the direction of arrival angle of each speaker relative to the microphone array; for each speaker among at least one speaker, speech synthesis is performed on their facial visual information using a lip-reading model to obtain the visual time-frequency mask of each speaker. The fusion module is used to fuse the spatial time-frequency mask and visual time-frequency mask of any speaker to form the speaker's audiovisual mask; The separation module is used to apply the audiovisual mask of any speaker to the mixed speech signal to obtain the speaker's independent speech signal.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 6.