Voiceprint separation method of smart glasses, electronic device and storage medium

By performing time-synchronized processing of audio and image data on smart glasses and fusing multiple features into the voiceprint separation neural network model, the problem of low voiceprint separation accuracy in multi-person speaking scenarios is solved, achieving a more efficient voiceprint separation effect and improving the translation and communication capabilities of smart glasses.

CN119274571BActive Publication Date: 2025-11-21湖北星纪魅族集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411369583.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-11-21
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

Existing voiceprint separation methods have low accuracy in scenarios where multiple people are speaking simultaneously, making it difficult to effectively distinguish between different sound sources.

Method used

By performing time-synchronized processing on the target mixed audio data collected by the microphone on the smart glasses and the target image data collected by the camera, and combining the voiceprint separation neural network model, voiceprint separation is performed by fusing audio and image features, including using audio features such as the direction of arrival of the sound source signal and the energy of the sound source signal, as well as visual features such as the anchor frame position and confidence level in the image.

Benefits of technology

It improves the accuracy of voiceprint separation, enabling effective separation of the voiceprint of the identified object in complex environments, thus enhancing the translation and communication capabilities of smart glasses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274571B_ABST
    Figure CN119274571B_ABST
Patent Text Reader

Abstract

The application relates to a voiceprint separation method of smart glasses, an electronic device and a storage medium. The method comprises the following steps: performing time synchronization processing on target mixed audio data collected by a microphone on the smart glasses and target image data of an identified object collected by a camera; obtaining a first feature and a second feature of the target mixed audio data, and a third feature of the target image data, wherein the first feature is a basic attribute feature of the target mixed audio; inputting the first feature into a voiceprint separation neural network model, and fusing the second feature and the third feature in the voiceprint separation neural network model to separate the sound of the identified object from the target mixed audio data. The application solves the technical problem that the voiceprint separation method in the related art generally has low accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and in particular to a method for separating voiceprints in smart glasses, electronic devices, and storage media. Background Technology

[0002] Smart glasses, also known as smart glasses or wearable computing glasses, are glasses that integrate microcomputers and advanced electronics to provide functions similar to smartphones or tablets, displaying information directly in the user's field of vision. Smart glasses are finding increasingly diverse applications, one of which is translation, recognizing the speaker's speech, translating it into the wearer's native language, and displaying it on the glasses' screen, enabling the wearer to easily communicate with foreigners.

[0003] However, voiceprint separation methods in related technologies generally suffer from low accuracy, especially in scenarios where multiple people are speaking simultaneously, where the differentiation is not ideal. Currently, no effective solution has been proposed to address this problem. Summary of the Invention

[0004] This application provides a voiceprint separation method, electronic device, and storage medium for smart glasses, in order to at least solve the technical problem of low accuracy that is commonly found in voiceprint separation methods in related technologies.

[0005] According to one aspect of the embodiments of this application, a voiceprint separation method for smart glasses is provided, comprising: performing time synchronization processing on target mixed audio data collected by a microphone on the smart glasses and target image data of an identified object collected by a camera; obtaining a first feature and a second feature of the target mixed audio data, and a third feature of the target image data, wherein the first feature is a basic attribute feature of the target mixed audio; inputting the first feature into a voiceprint separation neural network model, and fusing the second feature and the third feature in the voiceprint separation neural network model to separate the voice of the identified object from the target mixed audio data.

[0006] Optionally, the second feature includes at least one of the following: direction of arrival of the sound source signal, time of arrival of the sound source signal, and energy of the sound source signal.

[0007] Optionally, the third feature includes the anchor box position with the highest confidence in the target image data and the confidence of the anchor box.

[0008] Optionally, the voiceprint separation neural network model includes an encoding network and a decoding network. The encoding network and the decoding network each include multiple network blocks. The decoding network further includes a fusion module, which is used to fuse the second feature and the third feature.

[0009] Optionally, the fusion module is located in the last network block of the decoding network.

[0010] Optionally, the network block includes multiple network layers, wherein the first layer is a convolutional neural network (CNN), the second layer is a memory processing unit, and the third layer is a CNN. The memory processing unit includes a recurrent neural network (RNN), a long short-term memory neural network (LSTM), or a merging processing module with a memory unit. The memory unit is used to store historical features, and the merging processing module is used to merge the historical features with the current features and update the current features to the memory unit.

[0011] Optionally, the fusion module is located before the memory processing unit.

[0012] Optionally, time synchronization processing of the target mixed audio data collected by the microphone on the smart glasses and the target image data of the identified object collected by the camera includes: obtaining the time difference between the audio transmission and the light transmission from the identified object; the time difference is used to correct the acquisition time of the target mixed audio data forward.

[0013] Optionally, the time difference is obtained based on the quotient of the estimated distance corresponding to the current application scenario and the speed of sound propagation.

[0014] Optionally, the smart glasses include a microphone array consisting of multiple microphones, and the acquisition of the time difference between audio transmission and light transmission from the object to be identified includes: determining the time when the target mixed audio data arrives at each microphone; calculating the positional distance of the object to be identified based on the time when the target mixed audio data arrives at each microphone and the array shape of the microphone array; and calculating the time difference based on the quotient of the positional distance and the speed of sound propagation.

[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a processor, and a memory storing a program, the program including instructions that, when executed by the processor, cause the processor to perform the methods in any of the above embodiments.

[0016] According to another aspect of the embodiments of this application, a non-transitory machine-readable medium storing computer instructions for causing a computer to perform the methods in any of the above embodiments is also provided.

[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program, which, when executed by a computer's processor, causes the computer to perform the methods in any of the above embodiments.

[0018] The beneficial effects of the embodiments of this application are as follows:

[0019] In this embodiment, given the different acquisition delays of sound and video, with sound having a relatively large delay, time synchronization processing is performed on the target mixed audio data acquired by the microphone on the smart glasses and the target image data of the identified object acquired by the camera. Furthermore, when performing voiceprint separation using the voiceprint separation neural network model, in addition to using the features of the target mixed audio data itself (the first feature), the second feature of the target mixed audio data and the third feature of the target image data are also fused, improving the accuracy of voiceprint separation and thus solving the technical problem of low accuracy commonly found in related technologies.

[0020] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a voiceprint separation method for smart glasses according to an embodiment of this application;

[0023] Figure 2a This is a diagram showing the position distribution of microphones on smart glasses according to an embodiment of this application;

[0024] Figure 2b This is a diagram showing the position distribution of microphones on another type of smart glasses according to an embodiment of this application;

[0025] Figure 3 This is a schematic diagram of echo cancellation according to an embodiment of this application;

[0026] Figure 4 This is another echo cancellation schematic diagram according to an embodiment of this application;

[0027] Figure 5 This is a flowchart of a silent detection method according to an embodiment of this application;

[0028] Figure 6 This is a schematic diagram of the camera position according to an embodiment of this application;

[0029] Figure 7 This is a schematic diagram of the generalized cross-correlation delay estimation method according to an embodiment of this application;

[0030] Figure 8 This is a schematic diagram of a far-field sound source location estimation method using two microphones according to an embodiment of this application;

[0031] Figure 9 This is a schematic diagram of a near-field sound source localization estimation method using three microphones according to an embodiment of this application;

[0032] Figure 10 This is a flowchart of a first feature acquisition method according to an embodiment of this application;

[0033] Figure 11 This is a schematic diagram of a YOLO parameter list according to an embodiment of this application;

[0034] Figure 12 This is a schematic diagram of another YOLO parameter list according to an embodiment of this application;

[0035] Figure 13 This is a diagram of the neural network model architecture for voiceprint separation according to an embodiment of this application;

[0036] Figure 14 This is a schematic diagram of a network block structure according to an embodiment of this application;

[0037] Figure 15 This is a schematic diagram of another voiceprint separation neural network model structure according to an embodiment of this application;

[0038] Figure 16 This is a schematic diagram of another voiceprint separation neural network model structure according to an embodiment of this application;

[0039] Figure 17 This is a schematic diagram of another network block structure according to an embodiment of this application;

[0040] Figure 18 This is a schematic diagram of another voiceprint separation neural network model structure according to an embodiment of this application;

[0041] Figure 19 This is a schematic diagram of the voiceprint separation device for smart glasses according to an embodiment of this application;

[0042] Figure 20 This is a schematic diagram of the structure of the electronic device in this embodiment. Detailed Implementation

[0043] Embodiments of this embodiment will now be described in more detail with reference to the accompanying drawings. While some embodiments of this embodiment are shown in the drawings, it should be understood that this embodiment can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this embodiment. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of this embodiment.

[0044] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:

[0045] CPU, or Central Processing Unit, also known as a general-purpose processor, is the "brain" of intelligent devices.

[0046] DSP stands for Digital Signal Processing.

[0047] ASR stands for Automatic Speech Recognition.

[0048] AR stands for Augmented Reality.

[0049] VAD, or Voice Activity Detection, aims to detect the presence of a voice signal.

[0050] Speech, human voice.

[0051] Noise, non-human voice, or ambient noise.

[0052] YOLO, "You Only Look Once," is an algorithm that uses convolutional neural networks for object detection.

[0053] An anchor box for object detection describes the coordinates of four points of a rectangle in an image that contains the object of interest.

[0054] MFCCs, Mel Frequency Cepstral Coefficents, are cepstral parameters extracted in the Mel-scale frequency domain and are a widely used feature in automatic speech and speaker recognition.

[0055] FBANK, filter bank, is an audio feature.

[0056] RNN stands for Recurrent Neural Network.

[0057] LSTM, Long Short-Term Memory Neural Network

[0058] GRU, Gate Circular Unit, functions similarly to LSTM.

[0059] In related technologies, voiceprint separation neural network models only use the feature information of the audio itself for voiceprint separation, resulting in low accuracy. To solve this problem, this application provides a related solution, which is described in detail below. The application scenarios of this application include, but are not limited to: scenarios where smart glasses are used for translation, and scenarios where smart glasses are used to separate and store the speaker's speech in a meeting. That is, this application can be used in any application scenario where the target voice needs to be separated in an environment with high background noise.

[0060] Figure 1 This is a flowchart of a voiceprint separation method for smart glasses according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0061] Step S102: Time synchronization processing is performed on the target mixed audio data collected by the microphone on the smart glasses and the target image data of the identified object collected by the camera.

[0062] It should be noted that the aforementioned target mixed audio data can include audio data from multiple different sound sources, such as audio data of the smart glasses wearer, audio data of the non-wearer of the smart glasses, audio data of the person speaking, and environmental audio data. The target image data of the identified object can be image data of the speaker closest to the smart glasses wearer, or an image captured by the smart glasses' camera. It is understood that the smart glasses' camera field of view is usually directly in front of the wearer, and the wearer is often facing the object to be identified. Therefore, the object to be identified will be within the camera's field of view, and the target image captured by the camera can include the object to be identified. Thus, the target image can be used in the voiceprint separation processing for that object.

[0063] In some embodiments of this application, the microphones may include multiple microphones, for example, forming a microphone array. Where there are two microphones, embodiments of this application provide, as follows: Figures 2a-2bThe diagram shows the microphone placement on the smart glasses. The microphones on the smart glasses continuously record audio, transmitting the collected mixed audio data to the wake-up module (a hardware and software combination that enables the smart glasses to quickly activate from sleep or standby mode and begin responding to user commands) every certain time interval (one frame) for processing. The mixed audio data generally includes three channels: microphone 1, which collects audio from a microphone close to the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; microphone 2, which collects audio from a microphone further away from the wearer's mouth; and microphone 2, which is more suitable for recording environmental sounds or non-wearer audio. The sound played by the smart glasses' speaker is also recorded. During voice interaction, the signal formed after the sound played by the speaker is collected by the microphone is called an echo. Since the purpose of the microphone is to collect human speech, the echo needs to be eliminated, otherwise it will interfere with the microphone's collection of human voice and speech recognition. Therefore, echo cancellation can be performed on the above mixed audio data to determine the target mixed audio data. In one possible implementation, this application embodiment provides... Figure 3 , Figure 4 The two echo cancellation methods shown are as follows. Figure 3 , 4 As shown, echo cancellation estimates the delay of the echo signal based on the signal played from the speaker (Text-to-Speech (TTS) signal) and the signal captured from the recording. The primary delay estimation method is Kalman filtering. After obtaining the delay, the echo signal and the recorded signal are aligned based on this delay, and the echo signal is subtracted from the recorded data. For example, a filter can be designed whose output signal has the same waveform as the echo but opposite phase, and this is added to the recorded data to cancel out the echo signal. The main filter algorithm is Normalized Least Mean Square (NLMS). In practical applications, echo residue caused by nonlinearity must also be considered, requiring additional modules to eliminate echo residue. Neural network-based echo cancellation algorithms can also be used. Low-power echo cancellation refers to echo cancellation running on low-power devices with relatively low computational power. If it is based on a DSP, a fixed-point program is generally used; if it is based on a neural network, the model size is generally small.

[0064] In addition, the dual microphones require silence detection for the audio after echo cancellation on each microphone. During voice interaction, users are silent for more than half the time, resulting in more than half of the recorded audio signal being silent. Processing this silent audio with speech recognition wastes computing resources, and transmitting it to the cloud for recognition also wastes network bandwidth. Therefore, it is necessary to detect and remove the silent portions before passing them to subsequent processing modules. The silence detection process is similar to the wake-up process. In one possible implementation, this application embodiment provides... Figure 5 The flowchart shown illustrates the noise detection method. Figure 5 As shown, this includes: feature extraction from audio data; determining the acoustic model; when using a neural network as the acoustic model, the fully connected layer multiplies the input feature vector with the weight matrix and adds a bias to form the output vector; softmax / sigmoid is usually used in the last layer of classification tasks to convert the output vector of the fully connected layer into a probability distribution; the probability output by the softmax / sigmoid layer represents the confidence level of each category. In a Voice Activity Detection (VAD) system, a threshold is usually set to determine whether the current audio segment is speech or silence. Two-way talk means both parties are talking. The simplest method is to perform silence detection based on the signals from two microphones. If human voice is detected, it means that the person corresponding to that microphone is talking. Microphone 1 (MIC1) corresponds to the wearer, and microphone 2 (MIC2) corresponds to the non-wearer / conversation participant / external environment. The results of two-way talk can include: no one is talking; only the wearer is talking; only the non-wearer is talking; and both the wearer and non-wearer are talking. When both the wearer and non-wearer are detected speaking, the above voiceprint separation steps can be performed. If no one is detected speaking, the process can return directly to continue collecting the next frame of data.

[0065] In some embodiments of this application, if the wearer and a non-wearer are speaking simultaneously, beamforming is required to enhance the signal from the non-wearer, such as the signal acquired by MIC2 in the example above. In far-field voice interaction scenarios, especially with environmental interference, the acquired signal quality is poor, making subsequent recognition difficult. Therefore, a microphone array composed of multiple microphones is used to enhance the signal in the target direction and suppress the signal in the non-target direction, thereby improving the signal-to-noise ratio. The simplest method for microphone arrays to enhance speech is to first find the target direction (the microphone at 0 degrees in front of the wearer), calculate the delay between the other microphones and the microphone in front, and then sum the delays of all microphone signals to achieve the enhancement purpose. In practical applications, the sound signal source and any microphone are not necessarily at 0 degrees in front of the wearer, and the timing and phase of the signals received by all microphones are inconsistent. Therefore, beamforming is needed to enhance one or all microphone signals. The General Side-lobe Canceller (GSC) algorithm can be used with post-filtering. Smart glasses have two microphones. The first microphone is closer to the wearer's mouth. Because of this proximity, it captures a stronger audio signal from the wearer compared to other microphones. The second microphone is farther from the mouth, capturing a lower-energy audio signal, but it's better suited for recording ambient sounds or recordings of non-wearers. Beamforming utilizes this two-microphone array to amplify one of the signals. For example, in a translation scenario, it amplifies the voice of the non-wearer.

[0066] In some embodiments of this application, based on the result of the dual-talk judgment, the signal to be identified is transmitted to the mobile phone via Bluetooth connection between the smart glasses and the terminal (mobile phone, tablet computer, etc.) for further processing. As shown in Table 1.

[0067] Table 1

[0068] Dual-talk judgment result Transmitted content No one spoke No transmission Only the wearer is speaking. No transmission Only the non-wearer was speaking. Transmit the second signal Both the wearer and the non-wearer were speaking. Beamforming transmits mixed signals

[0069] In some embodiments of this application, when the camera collects target image data of the identified object, at least one camera can be deployed on the smart glasses, facing forward, either in the center or on either the left or right sides. In one possible implementation, a... Figure 6The diagram shown illustrates a camera placement method, specifically a center-mounted camera. The purpose of taking photos is to capture images of the other person's lips, using these time-lapse photos to determine if the person is speaking. Photo taking is generally only triggered in translation scenarios, capturing 2-8 frames per second. A lower capture frequency reduces computational load, but too low a frequency may lead to misjudgments. The size and resolution of photos taken in translation scenarios are fixed, for example, 480*600 pixels, consistent with the sample data used during training.

[0070] After acquiring the target mixed audio data and target image data, since the acquisition delays of sound and image are different, with sound delay being relatively large, this application embodiment proposes a step of time synchronization processing for the target mixed audio data and target image data. This step includes: acquiring the time difference between the audio transmission from the object being identified and the light transmission (light transmission is the process of light traveling from the shooting scene through the camera lens and finally reaching the image sensor), and using this time difference to correct the acquisition time of the target mixed audio data forward.

[0071] For example, if the calculated time difference is Δt, and the audio signal lags behind the image, the audio data can be timed forward by Δt during subsequent voiceprint separation to ensure synchronization between the audio and video data. For instance, target image data 1 and target mixed audio data 1 are acquired at time point t1. Since light travels very fast, target image data 1 can be considered the actual image at time point t1. However, sound travels much slower than light, so target mixed audio data 1 is not actually generated at time point t1. The audio data generated at time point t1 should be target mixed audio data 2 acquired at time point t1 + Δt. Therefore, this application advances target mixed audio data 2 to time point t1, i.e., performs time synchronization processing with target image data 1. This allows for joint processing of target mixed audio data 2 and target image data 1, improving the accuracy of voiceprint separation.

[0072] The determination of the aforementioned time difference can include: Method 1, estimation based on empirical values; Method 2, sound source localization based on multiple microphones. In some embodiments of this application, Method 1 can be obtained by dividing the estimated distance corresponding to the current application scenario by the speed of sound propagation. For example, in a translation scenario, the speaker is interested in the voice of the person speaking to the person wearing glasses. Typically, the person speaking to the person wearing glasses stands about 1 meter away from the person wearing glasses. The sound transmission delay at a distance of 1 meter is 1000 / 340 = 2.94 milliseconds, meaning that the audio arrives approximately 2.94 ms slower than the video. This 2.94 ms delay needs to be compensated for in the audio and video data. Assuming that the audio is captured in 10 ms frames, with a sampling rate of 16000 and a bit width of 16 bits, each frame contains 16000 sampled data points. The 2.94 ms delay corresponds to 4704 sampled data points, meaning that the current frame data needs to be shifted forward by 4704 sampled points to be synchronized with the currently captured photo.

[0073] Method 2 mentioned above includes: determining the arrival time of the target mixed audio data at each microphone; calculating the location distance of the identified object based on the arrival time of the target mixed audio data at each microphone and the array shape of the microphone array; and calculating the time difference based on the quotient of the location distance and the speed of sound propagation. For example, taking audio-video synchronization compensation for glasses with three or more microphones as an example, the audio-video time synchronization compensation described in Method 1 above is based on a rough estimate of empirical values. Since the conversation partners move their positions and orientations at any time, the estimation of the delay is not very accurate. This application proposes a reliable method for dynamically detecting the distance of conversation partners, namely, sound source localization based on multiple microphones. For example, the Time Difference of Arrival (TDOA) method. The sound source localization method based on TDOA generally consists of two steps: Step 1, calculating the time difference of the sound source signal arriving at the microphone array (time delay estimation); Step 2, establishing a sound source localization model through the geometry of the microphone array and solving it to obtain the location distance of the identified object (localization estimation).

[0074] Regarding step 1 above, there are many methods for estimating the delay of sound reaching different microphones. This embodiment is based on generalized cross-correlation delay estimation. In one possible implementation, this embodiment proposes... Figure 7 The diagram shown illustrates a delay estimation method based on generalized cross-correlation. Figure 7 As shown. Step 1 includes step 1.1: Calculating the cross-power spectrum of the sound signals reaching the two microphones. For example... Figure 7As shown, each signal is first windowed, then FFT-transformed to the frequency domain, one of which undergoes conjugate processing, and then multiplied to obtain the cross-power spectrum. Step 1.2: Cross-power spectrum weighting. Weighting is used to highlight and sharpen correlations and downplay uncorrelated ones. Weighting functions include cross-correlation function, smooth coherence transform, maximum likelihood weighting, and phase transform (PHAT) weighting. Among them, PHAT phase transform weighting is the most typical weighting function, with the formula 1.0 / np.abs(f1*np.conj(f2)), where f1 and f2 are the frequency domain transforms of the two signals, np.abs() is a function in the NumPy library used to calculate the absolute value (i.e., amplitude) of the input complex or real number array, and np.conj() is a function in the NumPy library used to calculate the conjugate value of the input complex number array. Step 1.3: Perform an Inverse Fast Fourier Transform (IFFT) on the weighted power spectrum to the time domain. Step 1.4: Calculate the time delay using peak detection. Specifically, find the point with the largest value in the time-domain signal, i.e., the peak value. In the context of time delay estimation, compare the similarity of two signals using this peak value to find the time difference (time delay) between them.

[0075] Step 2 above includes: The incident angle of the far-field sound source to the microphone array is equivalent to a plane wave; for different microphones, the only difference is the distance, and the incident angle is the same. The incident angle of the near-field sound source to the microphone array is equivalent to a spherical wave; for different microphones, the distance is different, and the incident angle is also different. The distance r between the sound source and the microphone is greater than... The time frame represents the far field, where d is the distance between the two microphones and λ is the sound wavelength. If the distance between the microphones on the glasses is greater than 3cm, the scene from the person speaking to the glasses microphones will definitely be a far field scenario. Figure 8 This is a schematic diagram of the far-field sound source location estimation method using two microphones provided in an embodiment of this application, as shown below. Figure 8 As shown, y1(k) and y2(k) are two microphones, and s(k) is the sound source signal. According to the trigonometric function relationship: delay ΔT = d cos(θ) / c, where c is the speed of sound 340 (m / s), d is the distance between the microphones, and θ is the incident angle. From step 1 above, we know the delay ΔT, so we can calculate the angle θ = arccos(ΔTc / d).

[0076] Positioning on a 2D plane requires 3 microphones; 2 microphones can only calculate angles. The following description assumes the glasses have 3 microphones. Figure 9 This is a schematic diagram of the near-field sound source localization estimation method using three microphones provided in an embodiment of this application, as shown below. Figure 9As shown, r1 is the distance between the sound source and microphone y1(k), r2 is the distance between the sound source and microphone y2(k), and r3 is the distance between the sound source and microphone y3(k). The time delay between microphone 1 and microphone 2, and the time delay between microphone 1 and microphone 3, can be calculated using the following equations:

[0077] ΔT12=(r2-r1) / c (Equation 1)

[0078] ΔT13=(r3-r1) / c (Equation 2)

[0079] The delays in Equations 1 and 2 above can be estimated based on the generalized cross-correlation, and the angle θ can also be calculated.

[0080] from Figure 9 The geometric relationship of the microphone shown can be obtained from the following equation:

[0081] (r2) 2 =(r1) 2 +(d) 2 +2*r1*d*cos(θ1)(Equation 3)

[0082] (r3) 2 =(r1) 2 +4*(d) 2 +4*r1*d*cos(θ1)(Equation 4)

[0083] In equations 1 to 4 above, there are three unknowns, r1, r2, and r3. These three unknowns can be solved using these four equations. Since there are three distances, in this embodiment, the average value of r1, r2, and r3 can be used as the location distance of the identified object. Furthermore, the time difference can be calculated based on the quotient of the location distance and the speed of sound propagation.

[0084] The distance between the speaker and the microphone has been calculated through the above steps. Since the speaker can move around at any time, the distance calculation should be updated regularly to make the audio and video delay compensation more accurate and the model will perform better when the voiceprint separation neural network model performs voiceprint separation.

[0085] Step S104: Obtain the first and second features of the target mixed audio data, and the third feature of the target image data, wherein the first feature is the basic attribute feature of the target mixed audio.

[0086] In some embodiments of this application, the second feature mentioned above includes at least one of the following: direction of arrival of the sound source signal, time of arrival of the sound source signal, and energy of the sound source signal. The third feature mentioned above includes the anchor frame position with the highest confidence in the target image data and the confidence level of the anchor frame.

[0087] For example, the first feature mentioned above could be FBANK or MFCC. The following describes the acquisition process of the first feature using FBANK as an example. Input: An audio segment, continuously acquired as input. The audio sampling rate is typically 16kHz, with a bit width of 16bit and a single channel. Output: Feature values. For example, if the audio segment is 160ms long, it is divided into 16 frames, each frame containing 40-dimensional data, resulting in 16 frames with features [16, 40]. Function: To handle real-time processing, the audio is segmented, typically 10ms per frame. Generally, a 10ms segment of audio is sufficient to extract meaningful features and results. In subsequent processing, to prevent spectral leakage, a method of partial overlap between consecutive frames is used. Each frame is 25ms long, with 15ms representing historical information—a 15ms overlap, meaning only 10ms are actually moved.

[0088] BANK (Filter-Bank) is obtained by summing the squares of the power spectrum of the Mel filter and then taking the logarithm. Compared to MFCC, it omits one step of discrete cosine transform calculation. The calculation steps of FBANK are as follows: Figure 10 As shown. Generally, the audio file is first pre-emphasized to enhance high-frequency signals. Then, it is framed and windowed. Frame segmentation divides the audio signal into 10ms frames, and windowing prevents spectral leakage. Features are calculated using 25ms of signal each time, meaning each frame is shifted by 10ms, effectively using 25ms of signal with 15ms of historical overlap. A Fourier transform is then performed to obtain the frequency domain signal, i.e., the spectrum, from the time-domain signal. The frequency domain values ​​at a certain time are accumulated to obtain the speech spectrum. This spectrum is then mapped to the Mel frequency scale using a Mel filter bank, and finally, the logarithm is taken to obtain the FBANK feature. In applications, a Mel filter bank is typically used with 40 or 80 filters, meaning each audio frame corresponds to 40 or 80 outputs.

[0089] When the second feature mentioned above includes the direction of arrival of the sound source signal, since the position and orientation of the microphone on the glasses are fixed, the direction of the speaker's and non-speaker's voices is already present in the audio after recording. If the direction of arrival of the sound source channel in the audio can be extracted, it can help determine whether the current speaker is a wearer or a non-wearer. In some embodiments of this application, sound source localization algorithms based on signal subspace methods, such as MUSIC (Multiple Signal Classification), beamforming methods, such as SRP-PHAT (Steered Response Power-Phase Transform), and phase transformation weighted controllable response power are proposed. Taking the MUSIC method as an example, the steps for direction estimation are as follows: Assume the glasses have N microphones and M sound sources, assuming M=2. Calculate the covariance matrix of the signal; perform eigenvalue decomposition on the covariance matrix; there are N sound source eigenvalues ​​and corresponding N eigenvectors, resulting in M ​​sound source subspaces and NM noise subspaces. Among the N sound source eigenvalues, the M largest eigenvalues ​​represent the sound source direction (i.e., the direction the sound signal arrives), and the NM smaller eigenvalues ​​represent the noise direction. Search for the angles of the two sound sources with the largest spatial spectrum. Sort the N eigenvalues ​​from largest to smallest, and for each angle, iterate through the first M eigenvalues, calculating the spatial spectrum. The angle with the largest spatial spectrum is the sound source direction. Iterate through each signal source (the two largest sound sources); iterate through each angle (e.g., θ (0~360), with a 1-degree interval); calculate the spatial spectrum. The direction corresponding to the angle with the largest spatial spectrum is the sound wave direction. The spatial spectrum formula is as follows:

[0090]

[0091] Where θ is the angle and α(θ) is the steering vector, which is the ideal response of the microphone array to a sound source in a certain direction. Let be the feature vector of the noise, (·) H This indicates the conjugate transpose.

[0092] Since the steering vector and noise are uncorrelated, their inner product is minimized. When the angle θ is the direction of the sound source, the steering vector and noise are orthogonal, or ideally, uncorrelated and orthogonal. The more orthogonal they are, the smaller the denominator becomes; ideally, the denominator is 0. The closer the denominator is to 0, the larger the P-value. The angle corresponding to the largest P-value is the direction of the sound source. The output of this step can be the spatial spectrum of the two sound sources obtained in the previous step, or it can be the spatial spectrum plus the corresponding angle value, used as subsequent multimodal features.

[0093] In some embodiments of this application, the sound source signal energy calculation in the second feature described above can be performed in either the time domain or the frequency domain. For example, frequency domain energy calculation uses a short-time FFT to calculate a frame of frequency domain signal, obtaining complex values ​​x at N (e.g., 256) frequency domain points, and then calculates the energy of the frequency domain signal. The energy calculation formula is: (x.real) 2 +x.imag 2 Using a moving average, the sliding window duration is 1 second, which is the average of the frequency domain energy values ​​of all frames within 1 second.

[0094] The third feature of the aforementioned target image data can be the target image portion of an image, obtaining the position and size of that target image portion. Commonly used target detection neural networks include YOLO and SSD. The following uses YOLO as an example to describe a target (e.g., lips) detection scheme.

[0095] YOLO is fast and can detect the location (anchor box) of an object and its category. It is a fully convolutional neural network, containing only convolutional, pooling, and fully connected layers, making it easy to accelerate to hardware. In one possible implementation, this application embodiment provides... Figure 11 The diagram shows a YOLO parameter list, and Figure 12 Another YOLO parameter list diagram is shown. In this embodiment, taking the first YOLO network as an example, its output includes three items: category confidence, anchor box confidence, and anchor boxes. There are 20 categories, which are 20 floating-point confidence values. This embodiment only cares about the lip category. There are 2 anchor box confidence values, which are 2 floating-point numbers. There are 2 anchor box position values, with 4 coordinate values ​​for each anchor box. Therefore, the output of the first YOLO network is 30 values. In this embodiment, since only the lip area is of interest, it is sufficient to detect the lip position in the image. Since there may be multiple people in front of the glasses wearer, only the lips of the closest person are of interest. That is, there can only be 1 or 0 detected targets. If there are multiple targets, the one with the largest size is selected, which is equivalent to the closest one. If the person is standing far away, that is, the lip size is relatively small, less than a certain threshold value (for example, less than 1 / 20 of the image), then it is not considered that lips have been detected, because this should not be a normal translation scenario. The output of this step is the image segmented from the anchor box with the highest confidence, along with the confidence score of the anchor box.

[0096] Step S106: Input the first feature into the voiceprint separation neural network model, and fuse the second feature and the third feature into the voiceprint separation neural network model to separate the voice of the identified object from the target mixed audio data.

[0097] In some embodiments of this application, the following are provided Figure 13The schematic diagram of the voiceprint separation neural network model shown is as follows: Figure 13 As shown, it includes an encoding network and a decoding network, both of which comprise multiple network blocks. Optionally, the decoding network further includes a fusion module for fusing the second and third features described above. In one possible implementation, embodiments of this application also provide... Figure 14 The diagram shows the structure of network blocks, including CNN, RNN / LSTM, and CNN. There are many ways to fuse components. For example, if features A and B are one-dimensional vectors, they can be fused together using the following method:

[0098] Dimensional expansion fusion: [A, B]

[0099] Plus fusion: A+B, provided that A and B have the same dimensions.

[0100] Multiplication fusion: A*B

[0101] This application's embodiments use dimensional expansion fusion. Since a dimensional expansion fusion method is used, it implies an increase in parameters. If fusion is performed earlier, the dimensions of subsequent networks will all need to increase. Therefore, it's best to place the fusion module towards the end of the entire neural network, such as placing it in the last network block of the decoding network. Figure 15 The diagram shows another neural network model structure for voiceprint separation, which can be achieved by simply adding parameters to the last sub-module.

[0102] Because multimodal features require temporal contextual information—for example, a single lip image cannot identify whether someone is speaking, but multiple images taken before and after it can create a video-like effect—it is necessary to retain information from a preceding period of time. This information can be stored using a recurrent neural network (RNN / LSTM). In other words, multimodal features need to be fused before the RNN / LSTM. Therefore, in some embodiments of this application, the architecture of the voiceprint separation neural network model can be that the aforementioned network block includes multiple network layers. The first layer of these multiple network layers is a convolutional neural network (CNN), the second layer is a recurrent neural network (RNN) or a long short-term memory (LSTM) neural network, and the third layer is a CNN. The fusion module is located before the RNN or LSTM, such as... Figure 16 The diagram shows another type of neural network model for voiceprint separation.

[0103] Alternatively, in one possible implementation, embodiments of this application provide as follows: Figure 17 The diagram of the network block structure shown is as follows: Figure 17As shown, the first layer of this multi-layer network is a Convolutional Neural Network (CNN), the second layer is a merging processing module with memory units, and the third layer is a CNN. The memory units are used to store historical features, and the merging processing module is used to merge historical features with current features and update the current features in the memory units. The fusion module is located before the merging processing module, as shown below. Figure 18 The diagram shows another type of voiceprint separation neural network model. The memory unit acts as a memory unit, expanding the time dimension. The feature data stored in the memory unit is similar to a 2D or multi-dimensional floating-point matrix.

[0104] After separating the target's voice from the mixed audio data, speech recognition can be performed on a mobile phone or on a cloud server. Generally, recognition on a cloud server yields better results, but also has higher latency due to network transmission delays, especially when there is no Wi-Fi or a weak network signal. In translation scenarios, translation is typically performed in the cloud. The translation result is then transmitted to the glasses. Alternatively, the phone receives the translated text and transmits it to the glasses via Bluetooth. Finally, the translation result is displayed on the glasses' screen.

[0105] Because there are short pauses in speech, after separating the speech of the identification target (non-wearer) in this embodiment, a "voice state" can be set. If there is one frame of silent data in the following period, the state is not considered to have changed from "voice state" to "silent state". Only after a certain period of time, such as 300ms of silence, will the non-wearer be considered to have entered the "silent state". As shown in Table 2.

[0106] Table 2

[0107] Dual-talk judgment result state No one spoke mute mode Only the wearer is speaking. mute mode Only the non-wearer was speaking. Human voice status Both the wearer and the non-wearer were speaking. Human voice status

[0108] In summary, this embodiment of the application, considering the different acquisition delays of sound and video (with sound having a relatively large delay), performs time synchronization processing on the target mixed audio data acquired by the microphone on the smart glasses and the target image data of the identified object acquired by the camera. Furthermore, when performing voiceprint separation using the voiceprint separation neural network model, in addition to using the features of the target mixed audio data itself (the first feature), it also integrates the second feature of the target mixed audio data and the third feature of the target image data, improving the accuracy of voiceprint separation and thus solving the technical problem of low accuracy in related voiceprint separation methods.

[0109] In addition, while related technologies use pre-fusion or mid-fusion, this application embodiment adopts post-fusion, specifically fusing at the last network block of the decoding network or before the memory processing unit of the decoding network, thereby reducing parameters and computational load.

[0110] Furthermore, since the audio features are temporal, lip movements are also time-dependent. To extract these temporal features, most of the designed networks use recurrent neural networks such as RNNs / LSTMs / GRUs. This leads to increased computation and parameters, making them unsuitable for low-power devices. This application's embodiment uses a merging processing module with memory units instead of recurrent neural networks, offering greater flexibility, fewer parameters, and lower computational cost.

[0111] Corresponding to the application scenarios and methods provided in the embodiments of this application, the embodiments of this application also provide a voiceprint separation device for smart glasses. For example... Figure 19 The diagram shown is a structural block diagram of a voiceprint separation device for smart glasses according to an embodiment of this application, comprising:

[0112] Synchronization module 1902 is used to perform time synchronization processing on the target mixed audio data collected by the microphone on the smart glasses and the target image data of the recognition object collected by the camera.

[0113] The acquisition module 1904 is used to acquire the first feature and the second feature of the target mixed audio data, and the third feature of the target image data, wherein the first feature is the basic attribute feature of the target mixed audio;

[0114] The separation module 1906 is used to input the first feature into the voiceprint separation neural network model, and fuse the second feature and the third feature in the voiceprint separation neural network model to separate the voice of the identified object from the target mixed audio data.

[0115] In some embodiments of this application, the second feature includes at least one of the following: the direction of arrival of the sound source signal, the arrival time of the sound source signal, and the energy of the sound source signal. The third feature includes the location of the anchor frame with the highest confidence in the target image data and the confidence level of the anchor frame.

[0116] The aforementioned voiceprint separation neural network model includes an encoding network and a decoding network. The encoding network and the decoding network each include multiple network blocks. The decoding network also includes a fusion module, which is used to fuse the second feature and the third feature.

[0117] In some embodiments of this application, the fusion module is located in the last network block of the decoding network.

[0118] The aforementioned network block includes multiple network layers. The first layer is a convolutional neural network (CNN), the second layer is a memory processing unit, and the third layer is a CNN. The memory processing unit includes a recurrent neural network (RNN), a long short-term memory neural network (LSTM), or a merging processing module with a memory unit. The memory unit is used to store historical features, and the merging processing module is used to merge the historical features with the current features and update the current features to the memory unit.

[0119] In some embodiments of this application, the fusion module is located before the memory processing unit.

[0120] The aforementioned synchronization module 1902 is further configured to acquire the time difference between audio transmission and light transmission from the object being identified; this time difference is used to correct the acquisition time of the target mixed audio data. Optionally, the time difference is obtained based on the quotient of the estimated distance corresponding to the current application scenario and the speed of sound propagation. The smart glasses include a microphone array composed of multiple microphones. Acquiring the time difference between audio transmission and light transmission from the object being identified includes: determining the time when the target mixed audio data arrives at each microphone; calculating the positional distance of the object being identified based on the arrival time of the target mixed audio data at each microphone and the array shape of the microphone array; and calculating the time difference based on the quotient of the positional distance and the speed of sound propagation.

[0121] The functions of each module in each device in the embodiments of this application can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0122] This application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the methods of this application embodiment.

[0123] This application also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the method of this application embodiment.

[0124] This application also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the methods of this application embodiment.

[0125] refer to Figure 20The present invention describes a structural block diagram of an electronic device that can serve as a server or client in embodiments of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the application described and / or claimed herein.

[0126] like Figure 20 As shown, the electronic device includes a computing unit 2001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 2002 or a computer program loaded into a random access memory (RAM) 2003 from a storage unit 2008. The RAM 2003 may also store various programs and data required for the operation of the electronic device. The computing unit 2001, ROM 2002, and RAM 2003 are interconnected via a bus 20020. An input / output (I / O) interface 2005 is also connected to the bus 20020.

[0127] Multiple components in the electronic device are connected to the I / O interface 2005, including: an input unit 2006, an output unit 2007, a storage unit 2008, and a communication unit 2009. The input unit 2006 can be any type of device capable of inputting information into the electronic device. The input unit 2006 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 2007 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 2008 may include, but is not limited to, a hard disk and an optical disk. The communication unit 2009 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network interface cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0128] The computing unit 2001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 2001 include, but are not limited to, CPUs, graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processors, controllers, microcontrollers, etc. The computing unit 2001 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of this application can be implemented as computer programs tangibly contained in a machine-readable medium, such as storage unit 2008. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 2002 and / or communication unit 2009. In some embodiments, the computing unit 2001 can be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).

[0129] Computer programs used to implement the methods of the embodiments of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0130] In the context of embodiments of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0131] It should be noted that the term "comprising" and its variations used in the embodiments of this application are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of this application are illustrative and not restrictive. Those skilled in the art should understand that, unless explicitly indicated otherwise in the context, they should be understood as "one or more".

[0132] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0133] The steps described in the method embodiments provided in this application can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of this application is not limited in this respect.

[0134] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily imply the same embodiment, nor does it imply independence from or alternative to other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.

[0135] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. Voiceprint separation methods for smart glasses, including: Time synchronization processing is performed on the target mixed audio data collected by the microphone on the smart glasses and the target image data of the object to be identified collected by the camera; Acquire a first feature and a second feature of the target mixed audio data, and a third feature of the target image data, wherein the first feature is a basic attribute feature of the target mixed audio; The first feature is input into the voiceprint separation neural network model, and the second feature and the third feature are fused in the voiceprint separation neural network model to separate the voice of the identified object from the target mixed audio data; The second feature includes at least one of the following: direction of arrival of the sound source signal, time of arrival of the sound source signal, and energy of the sound source signal; and / or the third feature includes the anchor frame position with the highest confidence in the target image data and the confidence of the anchor frame; The voiceprint separation neural network model includes an encoding network and a decoding network. The encoding network and the decoding network each include multiple network blocks. The decoding network also includes a fusion module, which is used to fuse the second feature and the third feature. The fusion module is located in the last network block of the decoding network. Each network block includes multiple network layers. The first layer of the multiple network layers is a convolutional neural network (CNN), the second layer is a memory processing unit, and the third layer is a CNN. The memory processing unit includes a recurrent neural network (RNN), a long short-term memory neural network (LSTM), or a merging processing module with a memory unit. The memory unit is used to store historical features. The merging processing module is used to merge the historical features with the current features and to update the current features to the memory unit.

2. The method according to claim 1, wherein, The fusion module is located before the memory processing unit.

3. The method according to claim 1, wherein, The time synchronization processing of the target mixed audio data collected by the microphone on the smart glasses and the target image data of the identified object collected by the camera includes: The time difference between the audio transmission and the light transmission from the identified object is obtained, and the time difference is used to correct the acquisition time of the target mixed audio data.

4. The method according to claim 3, wherein, The time difference is obtained by dividing the estimated distance corresponding to the current application scenario by the speed of sound propagation.

5. The method according to claim 3, wherein, The smart glasses include a microphone array consisting of multiple microphones, which acquires the time difference between audio transmission and light transmission from the object being identified, including: Determine the time when the target mixed audio data arrives at each microphone; The location distance of the identified object is calculated based on the arrival time of the target mixed audio data at each microphone and the array shape of the microphone array. The time difference is calculated based on the quotient of the location distance and the speed of sound propagation.

6. An electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1-5.

7. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Acoustic recognition model, method and system for Chinese and English mixed speech

    CN110930980A

  • Voice separation method and device, storage medium and electronic device

    CN113593587A