Voice separation method, device, equipment, medium and program product

By combining data collected from microphones and voice pickup sensors, voice activity detection and source analysis are performed, solving the problems of accuracy and computational load in voice separation in multi-person conversation scenarios, and achieving efficient voice separation results.

CN121096360APending Publication Date: 2025-12-09LUXSHARE PRECISION TECH(NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511554649.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

In multi-person conversation scenarios, mixed human voices result in lower accuracy for automatic speech recognition. Traditional human voice separation models are large in size and difficult to apply to embedded terminal systems. Low-frequency signals collected by vibration sensors are difficult to use directly for speech extraction and have limited retraining effects.

Method used

By combining microphones and voice pickup sensors to collect voice data, and through voice activity detection and source analysis, the source of the voice can be determined and separated in mixed speech. The voice pickup sensor is used to help identify the separation object, thereby improving the separation accuracy and reducing the amount of computation.

Benefits of technology

It improves the accuracy of near-field wearer voice data extraction, reduces the computational load of voice separation processing, and enhances the accuracy and efficiency of voice separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096360A_ABST
    Figure CN121096360A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice separation method and device, recognition, a medium and a program product. Comprising the steps of obtaining to-be-processed voice data; the to-be-processed voice data comprises first voice data and second voice data, the first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor; voice activity detection is performed on the first voice data and the second voice data, and a first activity identifier and a second activity identifier are determined; performing voice source analysis on the to-be-processed voice data through the first activity identifier and the second activity identifier, and determining a voice source analysis result; and when the voice source analysis result indicates that the to-be-processed voice data is the mixed voice, performing voice separation on the to-be-processed voice data. The voice collection condition of the voice pickup sensor is introduced into voice separation to assist in defining the object needing voice separation, so that the voice separation accuracy is improved, and the calculation amount required by voice separation processing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a speech separation method, apparatus, device, medium, and program product. Background Technology

[0002] With the continuous development of smart terminals, artificial intelligence (AI) translation and other voice recognition-based control functions have become standard features of terminal products. However, in multi-person conversation scenarios, mixed human voices will significantly interfere with the ability of automatic speech recognition (ASR) to extract the desired speech, resulting in low accuracy of text based on extracted speech, making it difficult to translate correctly or identify the correct control requests. Furthermore, traditional voice separation models are too large to be suitable for embedded terminal systems.

[0003] While introducing vibration sensors for voice pickup can suppress ambient noise and human voices, primarily focusing on extracting the vibration signals of the smart device wearer's own speech, thus facilitating the separation of different voice types, the low-frequency signals picked up by the vibration sensors are difficult to directly use for ASR (Automatic Speech Recognition) to extract the desired speech. Retraining the ASR model using low-frequency signals is costly, and the effectiveness of retrained models in extracting the desired speech remains limited, failing to meet the speech recognition accuracy requirements of smart devices. Summary of the Invention

[0004] This invention provides a speech separation method, apparatus, device, medium, and program product, which incorporates speech acquisition data from a vibration sensor into speech separation to help identify the target to be separated, thereby improving the accuracy of speech separation and reducing the computational load required for speech separation processing.

[0005] In a first aspect, embodiments of the present invention provide a speech separation method, including:

[0006] Acquire the voice data to be processed; the voice data to be processed includes first voice data and second voice data, the first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor;

[0007] Perform voice activity detection on the first voice data and the second voice data respectively to determine the first activity identifier and the second activity identifier;

[0008] The speech source analysis is performed on the speech data to be processed using the first and second activity identifiers to determine the speech source analysis results.

[0009] When the speech source analysis results indicate that the speech data to be processed is mixed speech, speech separation is performed on the speech data to be processed.

[0010] Secondly, embodiments of the present invention also provide a speech separation device, comprising:

[0011] The data acquisition module is used to acquire the voice data to be processed; the voice data to be processed includes first voice data and second voice data, the first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor.

[0012] The activity identifier determination module is used to perform voice activity detection on the first voice data and the second voice data respectively, and determine the first activity identifier and the second activity identifier;

[0013] The source result determination module is used to perform speech source analysis on the speech data to be processed using the first activity identifier and the second activity identifier, and determine the speech source analysis result.

[0014] The speech separation module is used to separate the speech data to be processed when the speech source analysis results indicate that the speech data to be processed is mixed speech.

[0015] Thirdly, embodiments of the present invention also provide a speech separation device, comprising:

[0016] At least one processor; and a memory communicatively connected to the at least one processor;

[0017] The memory stores a computer program that can be executed by at least one processor, which enables the at least one processor to implement the speech separation method of any embodiment of the present invention.

[0018] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the speech separation method of any embodiment of the present invention.

[0019] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program, which, when executed by a processor, is used to perform the speech separation method of any embodiment of the present invention.

[0020] This invention provides a speech separation method, apparatus, recognition, medium, and program product. The method involves acquiring speech data to be processed, including first speech data and second speech data. The first speech data is collected by a microphone, and the second speech data is collected by a speech pickup sensor. Speech activity detection is performed on both the first and second speech data to determine a first activity identifier and a second activity identifier. Speech source analysis is then performed on the speech data to be processed based on the first and second activity identifiers to determine the speech source analysis result. When the speech source analysis result indicates that the speech data to be processed is mixed speech, speech separation is performed. By adopting the above technical solution, a speech pickup sensor is introduced into the acquisition of speech data, allowing simultaneous acquisition of speech data through both the speech pickup sensor and the microphone, thus improving the accuracy of extracting the wearer's own speech data in the near field. Furthermore, by analyzing the activity of the speech data collected by the speech pickup sensor and the microphone, the possible sources of speech in the speech data to be processed are analyzed, and speech separation processing is only performed on the speech data that is determined to be likely mixed speech after analysis. By incorporating the speech acquisition data from the speech pickup sensor into speech separation to help identify the target speech to be separated, the accuracy of speech separation is improved and the amount of computation required for speech separation processing is reduced.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a speech separation method provided in Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of a speech separation method provided in Embodiment 2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of a speech separation device provided in Embodiment 3 of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a speech separation device provided in Embodiment 4 of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a speech separation method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where mixed human voices may be collected by intelligent voice devices in multi-person conversation scenarios. The method can be executed by a speech separation device, which can be implemented by software and / or hardware, and can be configured within a speech separation device. Optionally, the speech separation device can be a laptop, desktop computer, smart tablet, augmented reality (AR) glasses, virtual reality (VR) glasses, or true wireless stereo (TWS) earphones, etc. This embodiment of the present invention does not impose any limitations on this.

[0031] like Figure 1 As shown in the figure, the speech separation method provided by this embodiment of the invention specifically includes the following steps:

[0032] S101. Obtain the voice data to be processed.

[0033] The voice data to be processed includes first voice data and second voice data. The first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor.

[0034] In this embodiment, the voice data to be processed can be specifically understood as the voice data obtained from a multi-person conversation scenario, which may involve multiple people speaking at the same time, resulting in the superposition and mixing of human voices. It needs to be separated to determine the voice data that needs to be used for speech recognition.

[0035] In this embodiment, the first voice data can be specifically understood as voice data collected by one or more microphones installed on the smart terminal based on air vibrations. It is understood that the first voice data is data collected by the microphones from all sound waves in the air, which has a wide sound pickup range but is more susceptible to environmental interference.

[0036] In this embodiment, the second voice data can be specifically understood as voice data collected by a voice pick-up sensor (VPU) installed on a smart terminal, based on the conduction of sound vibrations from the human skeleton or vocal cords. It can be understood that the second voice data is data collected by the VPU based on human conduction, focusing more on near-field voice pickup and being less affected by ambient noise transmitted through the air, but the collected data tends to be low-frequency data.

[0037] Specifically, when a smart terminal needs to perform voice recognition and voice-based control operations, a voice acquisition device, including a voice processing unit (VPU), installed on the smart terminal can collect voice data from the conversation scenario. Voice data collected through multiple acquisition methods can then be combined as the voice data to be processed. Optionally, the voice acquisition device installed on the smart terminal may include a VPU and at least one microphone. Different voice acquisition devices can collect corresponding types of voice data, which together constitute the voice data to be processed.

[0038] It is understandable that, since different voice acquisition devices are used for acquisition, the voice data to be processed can be easily processed separately according to different acquisition methods, or it can be processed as a whole. This embodiment of the invention does not limit this.

[0039] S102. Perform voice activity detection on the first voice data and the second voice data respectively, and determine the first activity identifier and the second activity identifier.

[0040] In this embodiment, Voice Activity Detection (VAD) can be specifically understood as a detection method for automatically distinguishing speech segments (i.e., human voices) and non-speech segments (i.e., silence, background noise, and ambient sounds) in an audio signal, and can be used to detect whether there is a person speaking in the audio.

[0041] In this embodiment, the first activity identifier and the second activity identifier are respectively used to identify whether human voices exist in the first voice data and the second voice data.

[0042] Specifically, voice activity detection is performed on the first voice data to determine whether there is a human voice in the first voice data, and a first activity identifier corresponding to the first voice data is obtained based on the determination result; voice activity detection is performed on the second voice data to determine whether there is a human voice in the second voice data, and a second activity identifier corresponding to the second voice data is obtained based on the determination result.

[0043] S103. Perform speech source analysis on the speech data to be processed using the first activity identifier and the second activity identifier, and determine the speech source analysis result.

[0044] In this embodiment, voice source analysis can be specifically understood as an analysis used to distinguish which party in a multi-person conversation scenario the voice contained in the voice data to be processed is speaking. For example, voice source analysis can be used to distinguish whether the voice data to be processed comes from the near field (i.e., a person wearing a smart terminal device) or the far field (i.e., a person other than a person wearing a smart terminal device) in a multi-person conversation scenario.

[0045] Specifically, since the first activity identifier can be used to indicate whether the voice data to be processed contains human voices collected from the entire scene, and the second activity identifier can be used to indicate whether the voice data to be processed contains human voices collected from people wearing smart terminal devices, combining the two can comprehensively analyze whether the voice data to be processed contains human voices, and whether the contained human voices come from the far field, near field, or both in a multi-person conversation scene. Thus, the above analysis results can be determined as the voice source analysis results of the voice data to be processed.

[0046] S104. When the speech source analysis result indicates that the speech data to be processed is mixed speech, perform speech separation on the speech data to be processed.

[0047] In this embodiment, mixed speech can be specifically understood as the speech data to be processed containing both far-field and near-field human voices.

[0048] Specifically, when the speech source analysis results indicate that the speech data to be processed is mixed speech, it can be assumed that the speech data to be processed may contain human voices from both the far field and the near field. Since the second speech data must be human voices from the near field, the first speech data in the speech data to be processed can be analyzed a second time to determine the source of the speech contained therein. If it is determined that it does indeed contain human voices from both the far field and the near field, the speech data to be processed can be treated as a whole for speech separation. The separated speech can be identified to determine whether it belongs to the far field or the near field in a multi-person conversation scenario, so as to perform subsequent processing on the separated speech corresponding to the far field or the near field.

[0049] The technical solution of this embodiment acquires voice data to be processed, including first voice data and second voice data. The first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor. Voice activity detection is performed on the first and second voice data to determine a first activity identifier and a second activity identifier. Voice source analysis is then performed on the voice data to be processed based on the first and second activity identifiers to determine the voice source analysis result. When the voice source analysis result indicates that the voice data to be processed is mixed voice, voice separation is performed. By adopting the above technical solution, a voice pickup sensor is introduced into the voice data acquisition process, allowing simultaneous acquisition of voice data by both the voice pickup sensor and the microphone, thus improving the accuracy of extracting the wearer's own voice data in the near field. Furthermore, by analyzing the activity of the voice data collected by the voice pickup sensor and the microphone, the possible voice sources in the voice data to be processed are analyzed, and voice separation processing is only performed on the voice data to be processed that is determined to be likely mixed voice after analysis. By introducing the voice acquisition situation of the voice pickup sensor into the voice separation process to help identify the objects to be separated, the accuracy of voice separation is improved, and the computational load required for voice separation processing is reduced.

[0050] Example 2

[0051] Figure 2This is a flowchart of a speech separation method provided in Embodiment 2 of the present invention. The technical solution of this embodiment further optimizes the above-mentioned optional technical solutions. Based on the information indicating the existence of speech activity contained in the first and second activity identifiers, it determines whether there is speech data that needs to be separated in the speech data to be processed. It judges whether the existing speech data originates solely from the near field or far field, or is a mixed speech containing both near and far field speech, and obtains the corresponding speech source analysis result. Then, it performs direction of arrival analysis on the first speech data in the speech data to be processed that is determined to be mixed speech. By determining whether the first speech data originates from the mouth direction of the wearer of the smart terminal device, it determines whether the first speech data is near field speech. If the first speech data is determined to be near field speech, it is determined that speech separation is not required for the speech data to be processed. This achieves a secondary judgment on whether speech separation is needed for the speech data to be processed, and only performs speech separation on the speech data to be processed that is ultimately determined to contain mixed near and far field speech. Simultaneously, it can extract and compare the voiceprint features of the two separated speech streams, making the identified speech source more accurate, improving the accuracy of speech separation, and reducing the amount of computation required for speech separation processing.

[0052] like Figure 2 As shown, the speech separation method provided in Embodiment 2 of the present invention specifically includes the following steps:

[0053] S201. Obtain the voice data to be processed.

[0054] The voice data to be processed includes first voice data and second voice data. The first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor.

[0055] S202. Based on the first speech data and the second speech data, determine the first short-time zero-crossing rate and the first pitch period corresponding to the first speech data, and the second short-time zero-crossing rate and the second pitch period corresponding to the second speech data.

[0056] In this embodiment, the Short-Time Zero Crossing Rate (ZCR) can be specifically understood as the number of times the audio signal crosses the zero level per unit time in speech data, and can be used to distinguish speech (especially unvoiced sounds such as "s" and "sh") from background noise in speech data. The Pitch Period can be specifically understood as a parameter used in speech data to describe the characteristics of voiced sounds (such as vowels "a", "i", "u", etc.). It can be understood that since voiced sounds are generated by the vibration of the vocal cords, and this vibration is periodic, the waveform of the voiced sound also has obvious periodic repetition. The time interval between two adjacent repetitive waveforms can be used as the pitch period. In this embodiment of the invention, by calculating the pitch period, it can be used to distinguish whether there are voiced sound features in the speech data, and thus determine whether there is speech activity in the speech data.

[0057] Specifically, based on the first speech sampling signal corresponding to each sampling time in the first speech data, and the sign of each first speech sampling signal, the first short-time zero-crossing rate and the first pitch period of the first speech data are determined; based on the second speech sampling signal corresponding to each sampling time in the second speech data, and the sign of each second speech sampling signal, the second short-time zero-crossing rate and the second pitch period of the second speech data are determined.

[0058] In some examples, the short-time zero-crossing rate is calculated as follows:

[0059]

[0060] in,

[0061] Where n represents the nth frame in the speech data; N is the maximum number of sampling frames in the speech data; This refers to the speech sample signal at the m-th sampling point within the n-th frame; The minimum time index of the sampling point within the nth frame; is the maximum time index of the sampling point within the nth frame; sgn is the sign function, which takes the value 1 when the input value is non-negative and 0 when the input value is negative; Let be the window in the nth frame.

[0062] In some examples, the fundamental period is determined as follows:

[0063] The signal similarity between each audio data sampling frame can be calculated first. Where k is the accumulated value, the signal similarity corresponding to the nth frame in the speech data is determined by the above formula, and then the maximum value among the signal similarities is determined as the pitch period. .

[0064] S203. Based on the preset short-time zero-crossing rate threshold, the preset pitch period range, the first short-time zero-crossing rate, the first pitch period, the second short-time zero-crossing rate, and the second pitch period, determine the first activity identifier of the first speech data and the second activity identifier of the second speech data, respectively.

[0065] In this embodiment, the preset short-time zero-crossing rate threshold can be understood as a short-time zero-crossing rate boundary value pre-set to distinguish whether the speech data contains human voice, based on the characteristics of speech data that actually contains human voice. In some examples, the preset short-time zero-crossing rate threshold can be 0.6, or it can be pre-set according to the actual situation. This embodiment of the invention does not limit this.

[0066] In this embodiment, the preset pitch period range can be understood as a pitch period interval pre-set based on the characteristics of actual human voice speech data to distinguish whether the speech data contains voiced sounds. In some examples, the preset pitch period range can be (8, 94), or it can be pre-set according to the actual situation. This embodiment of the invention does not limit this.

[0067] Specifically, the first short-time zero-crossing rate is compared with a preset short-time zero-crossing rate threshold, and the first pitch period is compared with a preset pitch period range. When the first short-time zero-crossing rate is greater than the preset short-time zero-crossing rate threshold and the first pitch period is within the preset pitch period range, the first activity identifier of the first speech data is determined as an identifier indicating the presence of speech activity; otherwise, the first activity identifier of the first speech data is determined as an identifier indicating the absence of speech activity. Similarly, the second short-time zero-crossing rate is compared with a preset short-time zero-crossing rate threshold, and the second pitch period is compared with a preset pitch period range. When the second short-time zero-crossing rate is greater than the preset short-time zero-crossing rate threshold and the second pitch period is within the preset pitch period range, the second activity identifier of the second speech data is determined as an identifier indicating the presence of speech activity; otherwise, the second activity identifier of the second speech data is determined as an identifier indicating the absence of speech activity.

[0068] S204. Perform speech source analysis on the speech data to be processed using the first activity identifier and the second activity identifier, and determine the speech source analysis results.

[0069] Optionally, the voice data to be processed is analyzed for voice source using the first activity identifier and the second activity identifier to determine the voice source analysis results, including the following cases:

[0070] When both the first activity identifier and the second activity identifier indicate that there is no voice activity, the voice source analysis result is determined to be no voice source;

[0071] When both the first activity identifier and the second activity identifier indicate the presence of speech activity, the speech source analysis result is determined to be mixed speech;

[0072] When the first activity identifier indicates that a voice activity exists, and the second activity identifier indicates that a voice activity does not exist, the voice source analysis result is determined to be far-field voice.

[0073] When the first activity identifier indicates that the voice activity does not exist, and the second activity identifier indicates that the voice activity exists, the voice source analysis result is determined to be near-field voice.

[0074] Specifically, when both the first and second activity indicators indicate that voice activity does not exist, it can be assumed that neither the microphone nor the VPU in the smart terminal device has collected sound data containing human voice. In this case, the voice source analysis result can be determined as no voice source. Furthermore, since there is actually no voice in the voice data to be processed, there is no need for voice separation processing. When both the first and second activity indicators indicate that voice activity exists, it can be assumed that both the microphone and the VPU in the smart terminal device have collected sound data containing human voice. Since the sound data collected by the microphone has a wider coverage area, it is difficult to distinguish between near-field and far-field speech. In this case, it can be temporarily assumed that the voice data to be processed is a mixed near-field and far-field speech containing both near-field and far-field speech. At this point, the speech source analysis result can be determined as mixed speech, and it can be considered that speech separation processing is required for the speech data to be processed. When the first activity indicator indicates that speech activity exists, and the second activity indicator indicates that speech activity does not exist, it can be considered that only the microphone in the smart terminal device has collected sound data containing human voice. Since the VPU has not collected sound data containing human voice, it can be considered that there is no near-field speech in the speech data to be processed. Therefore, it can be further determined that the sound data containing human voice collected by the microphone is far-field speech. At this point, the speech source analysis result can be determined as far-field speech. When the first activity indicator indicates that speech activity does not exist, and the second activity indicator indicates that speech activity exists, it can be considered that only the VPU in the smart terminal device has collected sound data containing human voice. Since the VPU only collects near-field speech, the speech source analysis result can be determined as near-field speech.

[0075] S205. When the speech source analysis result indicates that the speech data to be processed is mixed speech, determine the direction of arrival of the first speech data; when the direction of arrival is the preset mouth direction, execute S206; when the direction of arrival is not the preset mouth direction, execute S207.

[0076] In this embodiment, the direction of arrival (DOA) can be specifically understood as the information obtained from which the first speech data was collected after performing DOA estimation on the first speech data. For example, the DOA can be determined using beamforming algorithms, multiple signal classification (MUSIC) algorithms, and rotation-invariant subspace (ESPRIT) algorithms, etc., and this embodiment of the invention does not impose any limitations on these methods.

[0077] In this embodiment, the preset mouth direction can be understood as directional information pre-set based on the mouth position of the person wearing the smart terminal device, used to assist in determining the direction of speech.

[0078] Specifically, when the speech source analysis result indicates that the speech data to be processed is mixed speech, in order to further distinguish whether the first speech data is indeed composed of both far-field and near-field speech, the direction of arrival (DOA) of the first speech data can be determined by the DOA algorithm, and the composition of the first speech data can be determined by comparing the DOA with a preset mouth direction. When the DOA is the preset mouth direction, it can be assumed that both the first and second speech data come from the wearer of the smart terminal device, and thus both belong to near-field speech, at which point S206 is executed; when the DOA is not the preset mouth direction, it can be assumed that the speech data to be processed contains not only the second speech data belonging to near-field speech, but also the first speech data that necessarily contains speech data that does not belong to near-field speech, at which point S207 is executed.

[0079] For example, an exemplary method for determining the direction of arrival is given here, which may include the following steps:

[0080] 1) When calculating the direction of arrival (DOA), the first speech data x in the time domain acquired by the microphone can be transformed into the frequency domain using a time-frequency transform (TF-F). This time-frequency transform can be performed using a Fast Fourier Transform (FFT) to obtain the corresponding frequency domain signal x(k). The first speech data includes speech data acquired by two microphones. and For example, the resulting frequency domain signal can be represented as follows: and Where k can represent a frequency domain subscript, and in the Fast Fourier Transform, k can be taken as: .

[0081] 2) Based on the angle parameter, the plane area perpendicular to the microphone in the space around the microphone is divided into a preset number of regions n, for example, n=360 / 60=6, and the angle ranges corresponding to each region n are 0-60, 60-120, ..., 300-360.

[0082] The time delay for the first voice data transmission is: Where d is the distance between elements in the microphone array, c is the speed of speech signal transmission, i.e., the speed of sound; when there are two or more microphones, i is the microphone number. The angular frequency at point k can be expressed as... This can be used to indicate the number of radians that the first speech data rotates in space per second, where, It is the sampling frequency, and 256 is the maximum value of k mentioned above.

[0083] 3) Based on the determined time delay and angular frequency, the steering vector can be determined: .

[0084] 4) Multiply the frequency signal data in each region by the corresponding steering vector and sum them to calculate the total speech energy picked up by the microphone in each region. For example, the total energy picked up in the nth region (n=1,2,3,4,5,6) can be written as:

[0085]

[0086] in, express . conjugate.

[0087] 5) Take the direction of maximum energy as the direction of arrival: .

[0088] S206. Redetermine the speech source analysis results as near-field speech, and determine that the speech data to be processed does not require speech separation.

[0089] Specifically, since both the first and second voice data come from the wearer of the smart terminal device, it can be assumed that all voice data in the voice data to be processed is near-field voice. At this time, the voice source analysis result of the voice data to be processed can be redefined as near-field voice. Meanwhile, since there is no far-field voice that needs to be separated, it can be determined that the voice data to be processed does not need to be separated.

[0090] S207. The speech data to be processed is separated into first separated speech data and second separated speech data through a pre-built speech separation model.

[0091] In this embodiment, the speech separation model can be specifically understood as a neural network model used to classify and process the input speech data. For example, the speech separation model can be a UNET network, which may include multiple sets of symmetrical encoder-decoder structures, and the connection part is implemented through gated recurrent units (GRUs).

[0092] Specifically, the speech data to be processed is input as a whole into a pre-built speech separation model for processing, and the two speech data obtained after separation are used as the first separated speech data and the second separated speech data, respectively.

[0093] Optionally, after separating the speech data to be processed into first separated speech data and second separated speech data using a pre-built speech separation model, the method further includes:

[0094] Voiceprint features are extracted from the first and second separated speech data respectively to determine the first and second voiceprint features.

[0095] The first and second voiceprint features are matched with the far-field and near-field voiceprint features respectively for similarity. Based on the similarity matching results, the voice source analysis results of the first and second separated speech data are re-determined.

[0096] Specifically, since the voiceprint information generated by different people is significantly different, in order to more clearly determine whether the first and second separated speech data belong to the far-field speech or the near-field speech in a multi-person conversation scenario, voiceprint features can be extracted from the first and second separated speech data respectively. The extracted voiceprint features are then used as the first and second voiceprint features, respectively. These features can then be matched with the far-field voiceprint features corresponding to the far-field speech in the multi-person conversation scenario, and the near-field voiceprint features corresponding to the near-field speech. The matching result with the highest similarity is determined as the required similarity matching result. For example, if the first voiceprint feature has the highest similarity to the far-field voiceprint feature, and the second voiceprint feature has the second highest similarity, then the first separated speech data corresponding to the first voiceprint feature should be considered to belong to the far-field speech, and the second separated speech data corresponding to the second voiceprint feature should be determined to belong to the near-field speech. Based on the above conclusions, the speech source analysis results of the first and second separated speech data are then re-determined.

[0097] Optionally, after determining the speech source analysis results of the speech data to be processed, speech recognition can be performed on the speech data from the corresponding source according to the speech recognition requirements of the smart terminal device. For example, for AI translation requirements, speech data whose speech source analysis results are far-field speech can be used as input to a pre-built automatic speech recognition model to determine the speech recognition result.

[0098] In this embodiment, the Automatic Speech Recognition (ASR) model, which is pre-trained based on the actual situation, is capable of recognizing human voices in the input speech data and converting them into text.

[0099] The technical solution of this embodiment, based on the information indicating the existence of voice activity contained in the first and second activity identifiers, determines whether there is voice data that needs to be separated in the voice data to be processed. It judges whether the existing voice data originates solely from the near field or far field, or is a mixture of near and far field voices, obtaining the corresponding voice source analysis result. Then, it performs direction-of-arrival analysis on the first voice data in the voice data to be processed that is determined to be mixed voice. By determining whether the first voice data originates from the direction of the mouth of the wearer of the smart terminal device, it determines whether the first voice data is near field voice. If the first voice data is determined to be near field voice, it is determined that there is no need to separate the voice data to be processed. This achieves a secondary judgment on whether the voice data to be processed needs to be separated, and only performs voice separation on the voice data to be processed that is ultimately determined to contain a mixture of near and far field voices. Simultaneously, it can extract and compare the voiceprint features of the two separated voices, making the determined voice source of the separated voice more accurate, improving the accuracy of voice separation, and reducing the computational load required for voice separation processing.

[0100] Example 3

[0101] Figure 3 This is a schematic diagram of the structure of a speech separation device provided in Embodiment 3 of the present invention, as shown below. Figure 3 As shown, the speech separation device includes a data acquisition module 31, an activity identifier determination module 32, a source result determination module 33, and a speech separation module 34.

[0102] The system includes a data acquisition module 31 for acquiring voice data to be processed, which includes first voice data and second voice data. The first voice data is acquired by a microphone, and the second voice data is acquired by a voice pickup sensor. An activity identifier determination module 32 is used to perform voice activity detection on the first voice data and the second voice data respectively, and determine the first activity identifier and the second activity identifier. A source result determination module 33 is used to perform voice source analysis on the voice data to be processed using the first activity identifier and the second activity identifier, and determine the voice source analysis result. A voice separation module 34 is used to perform voice separation on the voice data to be processed when the voice source analysis result indicates that the voice data to be processed is mixed voice.

[0103] The technical solution of this invention introduces a voice pickup sensor into the acquisition of voice data, enabling simultaneous acquisition of voice data via both the voice pickup sensor and microphone, thus improving the accuracy of extracting the wearer's own voice data in the near field. Furthermore, by analyzing the activity of the voice data acquired by the voice pickup sensor and microphone, the potential sources of voice data to be processed are analyzed, and voice separation processing is performed only on voice data that is determined to be likely mixed speech. By incorporating the voice acquisition data of the voice pickup sensor into the voice separation process to help identify the objects requiring voice separation, the accuracy of voice separation is improved, and the computational load required for voice separation processing is reduced.

[0104] Optionally, the source result determination module 33 is specifically used for:

[0105] When both the first activity identifier and the second activity identifier indicate that there is no voice activity, the voice source analysis result is determined to be no voice source;

[0106] When both the first activity identifier and the second activity identifier indicate the presence of speech activity, the speech source analysis result is determined to be mixed speech;

[0107] When the first activity identifier indicates that a voice activity exists, and the second activity identifier indicates that a voice activity does not exist, the voice source analysis result is determined to be far-field voice.

[0108] When the first activity identifier indicates that the voice activity does not exist, and the second activity identifier indicates that the voice activity exists, the voice source analysis result is determined to be near-field voice.

[0109] Optional, the activity identifier determination module 32 is specifically used for:

[0110] Based on the first speech data and the second speech data, determine the first short-time zero-crossing rate and the first pitch period corresponding to the first speech data, and the second short-time zero-crossing rate and the second pitch period corresponding to the second speech data, respectively.

[0111] Based on the preset short-time zero-crossing rate threshold, the preset pitch period range, the first short-time zero-crossing rate, the first pitch period, the second short-time zero-crossing rate, and the second pitch period, the first activity identifier of the first speech data and the second activity identifier of the second speech data are determined respectively.

[0112] Optional, the voice separation module 34 is specifically used for:

[0113] Determine the direction of arrival of the first voice data;

[0114] If the direction of arrival of the first speech data is the preset mouth direction, then the speech source analysis result is redefined as near-field speech, and it is determined that the speech data to be processed does not need to be separated into speech.

[0115] If the direction of arrival of the first speech data is not the preset mouth direction, the speech data to be processed is separated into first separated speech data and second separated speech data by a pre-constructed speech separation model.

[0116] Optionally, after separating the speech data to be processed into first separated speech data and second separated speech data using a pre-built speech separation model, the method further includes:

[0117] Voiceprint features are extracted from the first and second separated speech data respectively to determine the first and second voiceprint features.

[0118] The first and second voiceprint features are matched with the far-field and near-field voiceprint features respectively for similarity. Based on the similarity matching results, the voice source analysis results of the first and second separated speech data are re-determined.

[0119] Optionally, the voice separation device also includes:

[0120] The speech recognition module is used to take speech data whose speech source analysis results are far-field speech as input to a pre-built automatic speech recognition model to determine the speech recognition result.

[0121] The speech separation device provided in this embodiment of the invention can execute the speech separation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0122] Example 4

[0123] Figure 4 This is a schematic diagram of a voice separation device according to Embodiment 4 of the present invention. The voice separation device 40 can represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The voice separation device 40 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0124] like Figure 4As shown, the voice separation device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 can also store various programs and data required for the operation of the voice separation device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0125] Multiple components in the voice separation device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, optical disk, etc.; and a communication unit 49, such as a network card, modem, wireless transceiver, etc. The communication unit 49 allows the voice separation device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0126] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as speech separation methods.

[0127] In some embodiments, the speech separation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on the speech separation device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the speech separation method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the speech separation method by any other suitable means (e.g., by means of firmware).

[0128] Optionally, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the speech separation method provided in any embodiment of the present invention.

[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a voice separation device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the voice separation device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0134] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0135] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0136] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A speech separation method, characterized in that, include: Acquire voice data to be processed; the voice data to be processed includes first voice data and second voice data, the first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor; Perform voice activity detection on the first voice data and the second voice data respectively to determine the first activity identifier and the second activity identifier; The speech source analysis is performed on the speech data to be processed using the first activity identifier and the second activity identifier to determine the speech source analysis result. When the speech source analysis result indicates that the speech data to be processed is mixed speech, speech separation is performed on the speech data to be processed.

2. The speech separation method according to claim 1, characterized in that, The step of performing voice source analysis on the voice data to be processed using the first activity identifier and the second activity identifier, and determining the voice source analysis result, includes: When both the first activity identifier and the second activity identifier indicate that there is no voice activity, the voice source analysis result is determined to be no voice source; When both the first activity identifier and the second activity identifier indicate the presence of voice activity, the voice source analysis result is determined to be mixed voice; When the first activity identifier indicates that a voice activity exists, and the second activity identifier indicates that a voice activity does not exist, the voice source analysis result is determined to be far-field voice. When the first activity identifier indicates that the voice activity does not exist, and the second activity identifier indicates that the voice activity exists, the voice source analysis result is determined to be near-field voice.

3. The speech separation method according to claim 1, characterized in that, The step of performing voice activity detection on the first voice data and the second voice data respectively, and determining the first activity identifier and the second activity identifier, includes: Based on the first speech data and the second speech data, determine the first short-time zero-crossing rate and the first pitch period corresponding to the first speech data, and the second short-time zero-crossing rate and the second pitch period corresponding to the second speech data, respectively. Based on the preset short-time zero-crossing rate threshold, the preset pitch period range, the first short-time zero-crossing rate, the first pitch period, the second short-time zero-crossing rate, and the second pitch period, the first activity identifier of the first speech data and the second activity identifier of the second speech data are determined respectively.

4. The speech separation method according to claim 1, characterized in that, The process of performing speech separation on the speech data to be processed includes: Determine the direction of arrival of the first voice data; If the direction of arrival of the first speech data is a preset mouth direction, then the speech source analysis result is redefined as near-field speech, and it is determined that the speech data to be processed does not need to be separated into speech. If the direction of arrival of the first speech data is not the preset mouth direction, the speech data to be processed is separated into first separated speech data and second separated speech data by a pre-constructed speech separation model.

5. The speech separation method according to claim 4, characterized in that, After separating the speech data to be processed into first separated speech data and second separated speech data using a pre-built speech separation model, the process further includes: Voiceprint features are extracted from the first and second separated speech data respectively to determine the first and second voiceprint features. The first and second voiceprint features are matched with far-field and near-field voiceprint features respectively for similarity. Based on the similarity matching results, the voice source analysis results of the first and second separated speech data are re-determined.

6. The speech separation method according to any one of claims 1-5, characterized in that, Also includes: The speech data, which is far-field speech according to the analysis result of the speech source, is used as the input of the pre-built automatic speech recognition model to determine the speech recognition result.

7. A speech separation device, characterized in that, include: The data acquisition module is used to acquire voice data to be processed; the voice data to be processed includes first voice data and second voice data, the first voice data is collected by a microphone, and the second voice data is collected by a voice pickup sensor. The activity identifier determination module is used to perform voice activity detection on the first voice data and the second voice data respectively, and determine the first activity identifier and the second activity identifier; The source result determination module is used to perform voice source analysis on the voice data to be processed using the first activity identifier and the second activity identifier, and determine the voice source analysis result. The speech separation module is used to perform speech separation on the speech data to be processed when the speech source analysis result indicates that the speech data to be processed is mixed speech.

8. A speech separation device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the speech separation method according to any one of claims 1-6.

9. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the speech separation method as described in any one of claims 1-6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the speech separation method as described in any one of claims 1-6.