Data processing method, electronic device, storage medium, and computer program product
By setting up a microphone array on smart glasses and performing silence detection and sound source identification, the difficulty of speech recognition when the wearer and non-wearer speak at the same time is solved, improving the accuracy of the translation function and the device's battery life.
Patent Information
- Application Number
- CN202410930876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-07-11
AI Technical Summary
When the wearer and non-wearer speak at the same time, smart glasses have difficulty effectively distinguishing speech, resulting in decreased speech recognition accuracy and affecting the user experience of translation functions.
A microphone array, including a first microphone and a second microphone, is used, positioned near and far from the wearer's mouth, respectively. Through silence detection and sound source determination, speech recognition is performed on the audio of the target sound source.
It improves the accuracy of voice separation, reduces the processor performance requirements of wearable devices, and enhances battery life.
Smart Images

Figure CN118942491B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular, to a data processing method, an electronic device, a storage medium, and a computer program product. BACKGROUND
[0002] The application scenarios of wearable devices such as smart glasses are more and more, for example, the smart glasses can support a translation function, speech recognition of speech content of a non-wearer, and translation of the speech content into text content in a target language, and then display of the text content, so that the wearer directly knows the speech content of the non-wearer by watching the displayed text content, thereby improving the communication convenience of the user. However, in the related art, the smart glasses are difficult to effectively distinguish the collected speech in the process of the wearer and the non-wearer speaking at the same time, thereby affecting the accuracy of speech recognition, resulting in a limited translation function, and affecting the user experience. SUMMARY
[0003] In a first aspect, an embodiment of the present application provides a data processing method of a wearable device, the wearable device comprising a microphone array, the microphone array comprising a first microphone and a second microphone, a target sound source of the first microphone being a wearer, and a target sound source of the second microphone being a non-wearer; and the data processing method comprising:
[0004] performing first mute detection on first audio collected by the first microphone to obtain a first mute detection result;
[0005] in response to the first mute detection result being a speech segment, performing second mute detection on second audio collected by the second microphone to obtain a second mute detection result;
[0006] determining a sound source of the first audio and the second audio according to the first mute detection result and the second mute detection result;
[0007] in response to the sound source comprising a target sound source in a current application scenario, performing speech recognition on audio collected by a microphone corresponding to the target sound source in the current application scenario.
[0008] In a second aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, and the processor implements the method of any one of the above when executing the computer program.
[0009] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, and the computer program is executed by a processor to implement the method of any one of the above.
[0010] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method of any one of the above.
[0011] The above description is only a summary of the technical solutions of the present application. In order to enable a more clear understanding of the technical means of the present application, the above description can be implemented according to the content of the description, and in order to enable the above and other purposes, characteristics and advantages of the present application to be more apparent and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0012] In the drawings, like reference numerals refer to same or similar elements throughout the several views. These drawings are not necessarily to scale. It should be understood that these drawings have been simplified for ease of illustration and description.
[0013] Figure 1A An intelligent glasses provided by an embodiment of the present application is shown;
[0014] Figure 1B Another intelligent glasses provided by an embodiment of the present application is shown;
[0015] Figure 2 A flow chart of a data processing method of a wearable device of an embodiment of the present application is shown;
[0016] Figure 3 An application scenario diagram of a data processing method of a wearable device of an embodiment of the present application is shown;
[0017] Figure 4 A block diagram of a data processing apparatus of a wearable device of an embodiment of the present application is shown;
[0018] Figure 5 A block diagram of an electronic device for implementing an embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature rather than limiting.
[0020] In the application of wearable devices such as intelligent glasses in translation scenarios, the collected audio usually includes the sound of the wearer and the non-wearer, and it is necessary to distinguish the audio corresponding to different sound sources in the collected audio and extract the audio of the non-wearer alone for speech processing. In view of the above situation, two schemes are proposed in the related art to distinguish the sound of different sound sources.
[0021] In one solution, a microphone array can be formed by multiple microphones arranged at different positions of the smart glasses, and the sound signals in a specific direction can be enhanced by adjusting the signal phases of the microphones, so as to enhance and extract the sound signals of the non-wearer. However, when the wearer and the non-wearer speak at the same time, it is difficult to effectively distinguish the sound of the wearer and the sound of the non-wearer, thereby affecting the accuracy of the translation function.
[0022] In another solution, machine learning and artificial intelligence algorithms can be used to distinguish the sound of the wearer and the sound of the non-wearer, for example, the collected audio can be processed by a pre-trained convolutional neural network or recurrent neural network. However, these deep learning models have high requirements for computing power, which is limited by the processor performance and power consumption requirements of wearable devices such as smart glasses, and needs to rely on other terminal devices such as mobile phones, computers or cloud servers to run. When the network connection is unstable or there is no network service, the smart glasses cannot realize the translation function alone. It can be seen that this solution has certain limitations and cannot meet the user's use requirements in specific scenarios.
[0023] The above two solutions in the related art have certain defects or limitations in the implementation scenario of the translation function of wearable devices such as smart glasses. In this regard, the embodiment of the present application provides a data processing method of a wearable device, which aims to improve the accuracy of voice separation and is also beneficial to reduce the performance requirements of the wearable device.
[0024] In the data processing method of the wearable device provided by the embodiment of the present application, the wearable device can be provided with a microphone array, and the microphone array can include a first microphone with a wearer sound source as a target sound source and a second microphone with a non-wearer sound source as a target sound source. It can be understood that both microphones have the function of collecting surrounding audio, that is, the microphone with the wearer as the target sound source can also collect the sound emitted by the non-wearer, and the microphone with the non-wearer as the target sound source is the same. Generally, it is necessary to enhance the collection effect of each microphone on the target sound source and reduce the influence of the non-target sound source on the microphone. In the conventional design, the purpose can be achieved by optimizing the position of the microphone, for example, the microphone with the wearer as the target sound source is mostly arranged near the mouth of the wearer, so that it has a good collection effect on the sound emitted by the wearer. The microphone with the non-wearer as the target sound source is mostly arranged away from the mouth of the wearer, which is beneficial to reduce the interference of the wearer's speech on the microphone.
[0025] Figure 1A An intelligent glasses provided by an embodiment of the present application is shown as follows, Figure 1AAs shown, in some examples, the smart glasses can include a frame and two legs, and the two legs are respectively arranged on opposite sides of the frame. The first microphone (MIC1 in the figure) and the second microphone (MIC2 in the figure) can be arranged on one of the legs, and the first microphone is arranged on the side of the leg close to the frame, and the second microphone is arranged on the side of the leg away from the frame. It can be understood that, in the wearing state of the smart glasses, the distance between the first microphone and the mouth of the wearer is less than the distance between the second microphone and the mouth of the wearer.
[0026] Figure 1B As shown, in some examples, the smart glasses can include a frame and two legs, and the two legs are respectively arranged on opposite sides of the frame. The first microphone (MIC1 in the figure) and the second microphone (MIC2 in the figure) can be arranged on one of the legs, and the first microphone is arranged on the side of the leg close to the frame, and the second microphone is arranged on the side of the leg away from the frame. It can be understood that, in the wearing state of the smart glasses, the distance between the first microphone and the mouth of the wearer is less than the distance between the second microphone and the mouth of the wearer. Figure 1B
[0027] The above two examples are only used to illustrate the arrangement of the first microphone and the second microphone. In fact, the arrangement of the first microphone and the second microphone on the smart glasses can be various, as long as the requirement that the first microphone is arranged more close to the mouth of the wearer than the second microphone is met, and the person skilled in the art can arrange flexibly according to the actual situation.
[0028] It should be noted that, the first microphone is arranged more close to the mouth of the wearer than the second microphone, and accordingly, the second microphone is arranged more close to the mouth of the non-wearer than the first microphone. Therefore, the strength of the sound signal of the wearer collected by the first microphone is stronger than the strength of the sound signal of the wearer collected by the second microphone, and vice versa, the strength of the sound signal of the non-wearer collected by the first microphone is weaker than the strength of the sound signal of the non-wearer collected by the second microphone. In addition, for the first microphone, in the case that the volume of the sound source of the wearer and the volume of the sound source of the non-wearer are the same, the strength of the sound signal of the wearer collected by the first microphone is stronger than the strength of the sound signal of the non-wearer; for the second microphone, in the case that the volume of the sound source of the wearer and the volume of the sound source of the non-wearer are the same, the strength of the sound signal of the non-wearer collected by the second microphone is stronger than the strength of the sound signal of the wearer.
[0029] Figure 2 As shown, in some examples, the smart glasses can include a frame and two legs, and the two legs are respectively arranged on opposite sides of the frame. The first microphone (MIC1 in the figure) and the second microphone (MIC2 in the figure) can be arranged on one of the legs, and the first microphone is arranged on the side of the leg close to the frame, and the second microphone is arranged on the side of the leg away from the frame. It can be understood that, in the wearing state of the smart glasses, the distance between the first microphone and the mouth of the wearer is less than the distance between the second microphone and the mouth of the wearer. Figure 2 As shown, in some examples, the smart glasses can include a frame and two legs, and the two legs are respectively arranged on opposite sides of the frame. The first microphone (MIC1 in the figure) and the second microphone (MIC2 in the figure) can be arranged on one of the legs, and the first microphone is arranged on the side of the leg close to the frame, and the second microphone is arranged on the side of the leg away from the frame. It can be understood that, in the wearing state of the smart glasses, the distance between the first microphone and the mouth of the wearer is less than the distance between the second microphone and the mouth of the wearer.
[0030] S101: performing first silence detection on the first audio collected by the first microphone to obtain a first silence detection result;
[0031] S102: in response to the first silence detection result indicating that there is a voice segment, performing second silence detection on the second audio collected by the second microphone to obtain a second silence detection result;
[0032] S103: determining a sound source of the first audio and the second audio according to the first silence detection result and the second silence detection result;
[0033] S104: in response to the sound source including a target sound source in a current application scenario, performing voice recognition on the audio collected by the microphone corresponding to the target sound source in the current application scenario.
[0034] The data processing method of the embodiments of the present application can be applied to various wearable devices with voice processing functions, for example, can be smart glasses, smart watches, smart bracelets, etc., but is not limited thereto. In the following description of the present application, smart glasses are taken as an example for detailed description.
[0035] In the embodiments of the present application, the first audio collected by the first microphone and the second audio collected by the second microphone can be audio collected based on the same time period. Exemplarily, the first microphone and the second microphone can continuously collect audio, and every interval of time, the first audio and the second audio corresponding to the same time period collected are respectively stored in the storage module of the wearable device. The processor of the wearable device reads the first audio and the second audio from the storage module and executes the data processing method provided by the embodiments of the present application. That is, before step S101, the first audio can be obtained from the storage module first; before step S102, the second audio can be obtained from the storage module first.
[0036] In some examples, the VAD (Voice Activity Detection) technology can be used to detect whether the audio collected by the first microphone or the second microphone contains a speech segment, thereby achieving the mute detection. For example, the mute detection process can include the following steps: (1) preprocessing: the audio signal is preprocessed, which can include pre-emphasis, frame division, and windowing processing, etc.; (2) calculating short-time energy: for each audio frame included in the audio signal, the energy value thereof is calculated, which can be calculated by summing the squares of all sample values of the audio; (3) comparing the energy value of each audio frame with a pre-set energy threshold, and determining whether the audio frame contains speech activity according to the comparison result; for example, if the energy value is greater than or equal to the energy threshold, it is determined that the audio frame contains speech activity, and if the energy value is less than the energy threshold, it is determined that the audio frame does not contain speech activity, i.e., the audio frame can be background noise or mute; (4) smoothing processing: in order to avoid too fragmented results of speech detection, smoothing processing can be introduced, for example, only when a plurality of continuous audio frames are detected as containing speech activity, the plurality of audio frames are determined as a speech segment; (5) post-processing: the detected speech segment is further processed, for example, too short speech breakpoints are eliminated, and too short mute frames are merged, etc.
[0037] In addition, in other examples of the present application, more complex machine learning algorithms such as neural networks can also be used to detect whether the audio collected by the first microphone or the second microphone contains a speech segment. For example, the mute detection process can include the following steps: (1) preprocessing of the audio signal: for example, noise removal, echo cancellation, etc. can be included to improve the quality of the audio data; (2) feature extraction of the audio signal: useful features are extracted from the preprocessed audio signal, such as short-time energy, short-time average zero-crossing rate, spectral features (such as Mel-frequency cepstral coefficients, MFCC), etc.; (3) decision making: for example, a certain form of model or algorithm (such as threshold detection, machine learning classifier, etc.) is used to determine whether the current frame contains a speech segment; (4) smoothing processing: in order to avoid too fragmented output of the speech activity signal, some smoothing processing is needed, such as sliding window averaging or using a hidden Markov model (HMM).
[0038] For example, the mute detection result can include whether the audio contains a speech segment. In the case of containing a speech segment, the mute detection result can further include the speech segment after removing the non-speech segment from the audio. The non-speech segment refers to an audio segment composed of continuous audio frames without speech activity.
[0039] In the embodiments of the present application, the first silence detection refers to the silence detection performed on the first audio, and the second silence detection refers to the silence detection performed on the second audio. The processes of the first silence detection and the second silence detection can refer to the foregoing examples, which will not be described here. In the embodiments of the present application, the silence detection is first performed on the first audio, and then the silence detection is performed on the second audio in the case that the first silence detection result indicates that there is a speech segment in the first audio. By using this implementation manner, the number of times of performing the silence detection on the second audio is reduced, and thus the power consumption is reduced. Moreover, the first microphone can be a microphone that takes the wearer sound source as the target sound source, because the power consumption of the silence detection on the wearer is lower than that on the non-wearer, and thus the overall power consumption is further reduced.
[0040] In the embodiments of the present application, the sound source refers to a target object that emits a sound corresponding to a speech segment. In different application scenarios, the target sound source to be recognized is different. Taking the common translation scenario and the voice assistant scenario as examples. In the translation scenario, the process of the wearer talking with the non-wearer is involved, and the sound source corresponding to the first audio or the second audio can include at least one of the wearer and the non-wearer. Generally, in the translation scenario, the speech segment corresponding to the non-wearer is finally subjected to voice processing, and thus the target sound source in the translation scenario is the non-wearer sound source. In the voice assistant scenario, the voice instruction emitted by the wearer needs to be recognized, and thus the target sound source in the voice assistant scenario is the wearer sound source.
[0041] Taking the translation scenario as an example, in the case that the sound source corresponding to the speech segment existing in the first audio or the second audio includes the target sound source, i.e., in the case that the first audio or the second audio includes the speech segment corresponding to the sound emitted by the non-wearer, such as the wearer and the non-wearer talking at the same time or only the non-wearer talking, step S104, i.e., the voice recognition on the second audio collected by the second microphone, is performed. In the case that the first audio and the second audio do not include the speech segment or the sound source corresponding to the existing speech segment does not include the target sound source, such as the wearer and the non-wearer not talking or only the wearer talking, step S104, i.e., the voice recognition on the second audio collected by the second microphone, is not performed.
[0042] Exemplarily, in step S101, the first confidence of the first audio can be obtained by inputting the first audio into the pre-trained silence detection model. For example, the silence detection model can be a model based on the VAD technology described above, and the first confidence can be an energy value obtained after the silence detection model pre-processes the first audio and calculates the short-time energy, or a probability value calculated based on the energy value, which can represent the probability of the presence of voice activity in the first audio. After obtaining the first confidence, a smaller confidence threshold can be used to compare with the first confidence of the first audio. If the first confidence is greater than or equal to the confidence threshold, it is determined that the first silence detection result contains a voice segment.
[0043] It can be understood that, in the case that the first silence detection result contains a voice segment, the voice segment can be emitted by the target sound source or the non-target sound source, and it is not distinguished in step S101. As described above, since the first audio collected by the first microphone can include sounds emitted by the target sound source or the non-target sound source, the energy values of the sounds emitted by the target sound source and the non-target sound source are different, and in order to avoid missing detection, a smaller confidence threshold is needed to compare with the first confidence. For example, in the translation scenario, the confidence threshold used in S101 can be a first threshold, which can be set to be smaller than the confidence output by the silence detection model when the first audio only contains sounds emitted by the non-target sound source in the current application scenario, so as to further reduce the missing detection rate.
[0044] For example, in the translation scenario, if the first silence detection result contains a voice segment, there can be three cases in the time period corresponding to the first audio: the first case is that only the wearer speaks, i.e., the sound source of the first audio only includes the non-target sound source in the current application scenario; the second case is that only the non-wearer speaks, i.e., the sound source of the first audio only includes the target sound source in the current application scenario; and the third case is that both the wearer and the non-wearer speak, i.e., the sound source of the first audio includes the target sound source and the non-target sound source in the current application scenario. Since it cannot be determined whether the sound source of the first audio includes the target sound source in the current application scenario in the case that the first silence detection result contains a voice segment, the second silence detection needs to be performed on the second audio, and whether the sound source of the first audio and the second audio includes the target sound source in the current application scenario is further determined according to the first silence detection result and the second silence detection result. It can be understood that, if the first silence detection result does not contain a voice segment, it means that no one speaks in the time period corresponding to the first audio, i.e., the first microphone collects silence or background noise, etc. In this case, step S102 does not need to be performed, i.e., the second silence detection does not need to be performed on the second audio.
[0045] In step S102, the second silence detection performed on the second audio can input the second audio into a pre-trained silence detection model to obtain a second confidence of the second audio. For example, the silence detection model can be a model based on the aforementioned VAD technology, and the second confidence can be an energy value obtained after the silence detection model pre-processes the second audio and performs short-time energy calculation, or a probability value calculated based on the energy value, which can represent a probability of the presence of voice activity in the second audio. After obtaining the second confidence, a larger confidence threshold can be used to compare with the second confidence of the second audio, and in the case that the second confidence is greater than or equal to the confidence threshold, it is determined that the second detection result contains a voice segment of a non-wearer sound source.
[0046] It can be understood that the purpose of the second silence detection performed on the second audio is to determine whether the second audio contains the sound of the target sound source corresponding to the second microphone. Based on this, taking the translation scenario as an example, the confidence threshold used in step S102 can be a third threshold, which can be equal to or slightly less than the confidence output by the silence detection model when the second audio contains only the sound of the non-wearer, but should be greater than the confidence output by the silence detection model when the second audio contains only the sound of the wearer, so that in the case that the second confidence is greater than the confidence threshold, it can be determined that the second detection result corresponding to the second audio contains a voice segment of a non-wearer sound source.
[0047] In step S103, if the first silence detection result is that the first audio contains a voice segment and the second silence detection result is that the second audio does not contain a voice segment, it can be determined that the sound sources of the first audio and the second audio only include the wearer sound source and do not include the non-wearer sound source.
[0048] If the first silence detection result is that the first audio contains a voice segment and the second silence detection result is that the second audio contains a voice segment, it can be determined that the sound sources of the first audio and the second audio can only include the non-wearer sound source, or can include both the wearer sound source and the non-wearer sound source.
[0049] In step S104, taking the current application scenario as the translation scenario, the target sound source is the non-wearer sound source. If step S103 determines that the sound sources of the first audio and the second audio at least include the non-wearer sound source, then the first audio and the second audio contain voice segments corresponding to the target sound source. Based on this, the second audio collected by the second microphone corresponding to the non-wearer sound source, i.e., the target sound source, can be subjected to voice recognition to obtain a voice recognition result corresponding to the voice segment emitted by the non-wearer.
[0050] In the embodiments of the present application, the speech recognition result can also be translated into target text according to the target language, and displayed on the display interface of the wearable device, such as smart glasses, so that the wearer knows the speech content of the speaker (such as a non-wearer).
[0051] According to the data processing method of the embodiments of the present application, the first audio collected by the first microphone is subjected to first silence detection to obtain a first silence detection result, and in the case that the first silence detection result is that there is a speech segment, the second audio collected by the second microphone is subjected to second silence detection to obtain a second silence detection result, and the sound source of the first audio and the second audio is judged according to the first silence detection result and the second silence detection result. Taking the current application scenario as an example, in the case that the sound source of the second audio includes a target sound source, i.e. a non-wearer sound source, speech recognition is performed on the second audio to obtain the speech recognition result of the sound emitted by the non-wearer sound source. In summary, the data processing method of the embodiments of the present application can realize the recognition of the sound source, specifically, it can recognize the situation that the wearer and the non-wearer speak at the same time, and in this case, the voiceprint separation processing is performed. Based on this, the data processing method of the embodiments of the present application improves the accuracy of the voice separation, thereby improving the application effect of the wearable device in the translation scenario, and also helps to reduce the processor performance of the wearable device, and further improves the endurance of the wearable device.
[0052] In one embodiment, before the first silence detection on the first audio and the second silence detection on the second audio, the first audio and the second audio can be subjected to echo cancellation processing.
[0053] Exemplarily, the first audio or the second audio can be subjected to echo cancellation processing by using the audio signal played by the loudspeaker of the smart glasses. Specifically, the delay of the echo signal can be calculated according to the audio signal played by the loudspeaker and the first audio or the second audio. According to the delay of the echo signal, the echo signal is time-aligned with the first audio and the second audio respectively, and then the cancellation signal which is the same as the echo signal waveform and opposite in phase output by the pre-set filter is used to accumulate with the first audio and the second audio respectively to cancel the echo signal in the first audio and the second audio. The filter can be obtained based on NLMS (Normalized Least Mean Square, an adaptive filtering algorithm).
[0054] In other examples of the present application, a neural network-based echo cancellation algorithm can also be used for echo cancellation processing, and those skilled in the art can choose flexibly according to the actual situation.
[0055] Through the above implementation, the interference of the sound signals played by the loudspeaker collected by the first microphone and the second microphone on subsequent mute detection and voice recognition can be excluded, so as to improve the accuracy of subsequent mute detection and voice recognition.
[0056] In an implementation, the first mute detection result includes a first confidence of the first audio. Wherein, in a case that the first confidence is greater than or equal to a first threshold, the first mute detection result is that there is a voice segment, and step S103 can include:
[0057] In response to the first confidence being greater than or equal to a second threshold, it is determined that the sound source includes the wearer sound source, and the second threshold is greater than or equal to the first threshold.
[0058] In the embodiments of the present application, the first confidence is used to represent the probability that the first audio includes a voice segment, the greater the first confidence, the greater the probability that the first audio includes a voice segment, and vice versa. The first confidence can be obtained by inputting the first audio into a pre-trained mute detection model, for example.
[0059] For example, the first threshold can be less than the second threshold. For the setting of the first threshold, the voice segment corresponding to the wearer sound source in the first audio collected by the first microphone can be input into the mute detection model to obtain a first reference confidence, and the voice segment corresponding to the non-wearer sound source in the first audio collected by the first microphone can be input into the mute detection model to obtain a second reference confidence. Wherein, the first reference confidence is greater than the second reference confidence. The first threshold can be set to be less than the second reference confidence, and the second threshold can be set to be greater than the second reference confidence and less than the first reference confidence.
[0060] In an implementation, the second mute detection result includes a second confidence of the second audio. Wherein, step S103 can further include:
[0061] In response to the second confidence being greater than or equal to a third threshold, it is determined that the sound source includes the non-wearer sound source.
[0062] In the embodiments of the present application, the second confidence is used to represent the probability that the second audio includes a voice segment, the greater the second confidence, the greater the probability that the second audio includes a voice segment, and vice versa. The second confidence can be obtained by inputting the second audio into a pre-trained mute detection model, for example.
[0063] Exemplarily, for setting the third threshold value, the speech segment corresponding to the wearer sound source in the second audio collected by the second microphone can be input into the silence detection model to obtain a third reference confidence, and the speech segment corresponding to the non-wearer sound source in the second audio collected by the second microphone can be input into the silence detection model to obtain a fourth reference confidence. The third reference confidence is less than the fourth reference confidence. The third threshold value can be set to be greater than the third reference confidence and less than the fourth reference confidence.
[0064] Thus, in the case where the second confidence is greater than or equal to the third threshold value, it can be determined that the sound source of the second audio includes the non-wearer sound source.
[0065] It should be noted that, in the case where the first confidence is greater than or equal to the second threshold value and the second confidence is greater than or equal to the third threshold value, it can be determined that the sound source of the first audio and the second audio includes both the wearer sound source and the non-wearer sound source. In the case where the first confidence is less than the second threshold value and the second confidence is greater than or equal to the third threshold value, it can be determined that the sound source of the first audio and the second audio includes only the non-wearer sound source and does not include the wearer sound source. In the case where the first confidence is greater than or equal to the second threshold value and the second confidence is less than the third threshold value, it can be determined that the sound source of the first audio and the second audio includes only the wearer sound source and does not include the non-wearer sound source.
[0066] In an implementation, the step S104 can specifically include the following steps:
[0067] S1401: In response to the target sound source in the current application scenario being a non-wearer sound source and the sound source including a wearer sound source and a non-wearer sound source, performing directional speech enhancement on the second audio to obtain enhanced second audio;
[0068] S1402: Performing speech recognition on the enhanced second audio.
[0069] Exemplarily, in the case where the current application scenario is a translation scenario, the target sound source in the current application scenario is a non-wearer sound source. In the case where the first confidence of the first audio is greater than or equal to the second threshold value and the second confidence of the second audio is greater than or equal to the third threshold value, it is determined that the sound source includes a wearer sound source and a non-wearer sound source.
[0070] Exemplarily, in step S1401, a beamforming technique can be employed to perform directional speech enhancement on the microphone array of the wearable device. Specifically, according to the specific geometry formed by the arrangement of the first microphone and the second microphone, sound signals from different directions can be captured. For example, since the first microphone is arranged more adjacent to the mouth of the wearer relative to the second microphone, and the second microphone is arranged more adjacent to the mouth of the non-wearer relative to the first microphone, the first microphone can be used to collect sound signals from the direction of the wearer sound source, and the second microphone can be used to collect sound signals from the direction of the non-wearer sound source. The time delay of the first audio collected by the first microphone relative to the second audio collected by the second microphone is calculated with the second microphone as the target microphone. According to the time delay, the first audio and the second audio are time-aligned and phase-adjusted so that the sound signals in the direction of the non-wearer sound source in the first audio and the second audio can be coherently superimposed. The superimposed sound signals are enhanced by using a preset beamformer, while sound signals from other directions (for example, the direction of the wearer sound source) are suppressed. The beamformer can be based on the architecture of the General Side-lobe Canceller (GSC) and post-filtering. In this way, the enhanced second audio is obtained, and the enhanced second audio includes the sound segment corresponding to the non-wearer sound source.
[0071] Exemplarily, in step S1402, feature extraction can be performed on the enhanced second audio to obtain corresponding audio features, which can include Mel Frequency Cepstral Coefficients (MFCCs) or Linear Predictive Coding (LPC), etc. The audio features are analyzed by using a pre-trained acoustic model to obtain a corresponding speech recognition result. The acoustic model can be established based on a Hidden Markov Model (HMM), a Deep Neural Network (DNN), or a combination thereof.
[0072] In an embodiment, step S1402 can specifically include the following steps:
[0073] performing voiceprint separation processing on the enhanced second audio to obtain the sound segment emitted by the non-wearer;
[0074] inputting the sound segment emitted by the non-wearer into a speech recognition model to obtain a speech recognition result of the second audio.
[0075] In this embodiment, the enhanced second audio can be processed by using a pre-trained voiceprint separation module. The voiceprint separation module can be obtained based on a neural network model, for example, a convolutional neural network (CNN), a recurrent neural network (RNN), a hybrid model of a convolutional neural network and a recurrent neural network, a Transformer model, or an end-to-end model, etc.
[0076] Exemplarily, the enhanced second audio can be subjected to feature extraction processing to obtain corresponding audio features. The audio features can specifically include time domain features or FBANK features. Then, the audio features are input into the voiceprint separation module to obtain the separated sound segments emitted by the non-wearer. Finally, the sound segments emitted by the non-wearer are input into the speech recognition model to obtain the speech recognition result of the second audio.
[0077] It should be noted that in the embodiments of the present application, the voiceprint separation processing can be performed on the wearable device or on the terminal device. For example, the wearable device can send the enhanced second audio to the terminal device, and perform voiceprint separation processing on the enhanced second audio through the voiceprint separation module pre-set on the terminal device. The wearable device and the terminal device can use Bluetooth transmission or wireless communication network transmission, which is not limited in the embodiments of the present application. Similarly, the speech recognition processing using the speech recognition model can be performed on the wearable device or on the terminal device.
[0078] In an implementation, step S1402 can specifically include the following steps:
[0079] According to the second silence detection result, at least one speech frame with speech activity is intercepted from the enhanced second audio.
[0080] The speech frame with speech activity is input into the speech recognition model to obtain the speech recognition result of the second audio.
[0081] In some examples, for the enhanced second audio, the speech frame with speech activity can be intercepted according to the time domain features of the audio frames included in the second audio. Specifically, the root mean square (RMS) of the time domain features of the audio frames included in the second audio is calculated, and the calculation formula is: {(x1^2+x2^2+x3^2+....+xn^2) / n}^0.5, wherein xn is used to represent the time domain feature value of the nth audio frame. In order to increase the robustness, the length of the sliding window used to calculate the root mean square can be 1 second, that is, the average value of the time domain feature values of all audio frames within 1 second. The audio frames are filtered according to the calculated root mean square, and the audio frames greater than or equal to a third threshold value are determined as the speech frames with speech activity.
[0082] In some examples, the speech frame with the speech activity can be further determined according to a frequency domain energy calculation of the audio frame included in the second audio. Specifically, the frequency domain energy calculation calculates the frequency domain signal of an audio frame using a short-time FFT to obtain a complex value x on N (e.g., 256) frequency domain points, and then calculates the energy of the frequency domain signal. The sliding average is also used for the frequency domain energy calculation, and the sliding window length can be 2 seconds, that is, the average of the frequency domain energy values of all audio frames within 2 seconds. Then, the audio frame is filtered according to the calculated frequency domain energy value, and the audio frame greater than or equal to a third threshold value is determined as the speech frame with the speech activity.
[0083] Through the above implementation, the data processing amount in the speech recognition process can be reduced, and the speech recognition efficiency can be improved.
[0084] Optionally, the speech frame with the speech activity is input into a speech recognition model to obtain a speech recognition result of the second audio, including:
[0085] The speech frame with the speech activity is sent to a terminal device, and the terminal device is configured to run a speech recognition model to obtain a speech recognition result of the second audio.
[0086] The speech recognition result of the second audio fed back by the terminal device is received.
[0087] By deploying the speech recognition model on the terminal device and sending the speech frame with the speech activity to the terminal device to perform speech recognition processing by using the terminal device, the performance requirement of the wearable device can be reduced, and the endurance of the wearable device can be improved.
[0088] The following refers to Figure 3 A data processing method of a wearable device according to an embodiment of the present application is described in a specific application scenario.
[0089] As shown in Figure 3 The data processing method can be implemented by the wearable device in cooperation with a mobile terminal. The wearable device can be smart glasses, and the mobile terminal can be a mobile phone. Figure 3 The data processing method in a translation scenario includes steps S1 to S12, wherein steps S1 to S7 and S12 can be performed by the smart glasses, and steps S8 to S11 can be performed by the mobile phone, as follows.
[0090] S1: recording. The first microphone and the second microphone on the smart glasses continuously record, and the audio data collected in each time (a frame) is transmitted to the wake-up module for processing. The first microphone is arranged more close to the mouth of the wearer than the second microphone.
[0091] S2: echo cancellation processing. In the process of voice interaction, the signal formed after the sound played by the loudspeaker is collected by the microphone is called echo, and the microphone is used to collect the voice of the person. The echo needs to be eliminated, otherwise it will interfere with the collection of the voice of the person and the voice recognition. Specifically, the audio signal played by the loudspeaker can be used to perform echo cancellation processing on the first audio collected by the first microphone and the second audio collected by the second microphone.
[0092] S3: single-path silence detection. In the process of voice interaction, the user does not speak more than half of the time, and more than half of the audio signal collected by the recording is silence. If the voice recognition process is performed on the audio with silence, it will waste computing resources, and if it is transmitted to the cloud for recognition, it will also waste network bandwidth. Therefore, the silent part needs to be detected and deleted before being handed over to the subsequent processing module. Specifically, the first silence detection model can be used to perform first silence detection on the first audio. If the first confidence of the first audio is greater than or equal to the first threshold, the first silence detection result is that the first audio has a voice segment.
[0093] S4: double-path silence detection. If the first silence detection result is that the first audio has a voice segment, the second silence detection model is used to perform second silence detection on the second audio to obtain the second confidence of the second audio.
[0094] S5: double-speech judgment. If the first confidence of the first audio is greater than or equal to the second threshold and the second confidence of the second audio is greater than or equal to the third threshold, it is determined that the sound source of the first audio and the second audio includes both the wearer sound source and the non-wearer sound source. If the first confidence of the first audio is less than the second threshold and the second confidence of the second audio is greater than or equal to the third threshold, it is determined that the sound source of the first audio and the second audio only includes the non-wearer sound source. If the first confidence of the first audio is greater than or equal to the second threshold and the second confidence of the second audio is less than the third threshold, it is determined that the sound source of the first audio and the second audio only includes the wearer sound source.
[0095] If the result of the double-speech judgment is yes, i.e., the sound source of the first audio and the second audio includes both the wearer sound source and the non-wearer sound source, step S6 is entered. If the result of the double-speech judgment is no, i.e., the sound source of the first audio and the second audio only includes the non-wearer sound source, step S7 is entered.
[0096] In addition, the smart glasses end can also transmit the result of the double-speech judgment to the mobile phone end, so that the mobile phone end performs subsequent processing on the second audio transmitted to the mobile phone end according to the result of the double-speech judgment.
[0097] S6: beamforming enhancement. If the result of the double-speech judgment is yes, directional voice enhancement is performed on the second audio using the beamforming technology.
[0098] S7: Transmission. In the case of the double-talk judgment result being yes, the second audio after directional speech enhancement obtained based on step S6 is transmitted to the mobile phone through Bluetooth for the mobile phone to execute step S8. In the case of the double-talk judgment result being no, the second audio is transmitted to the mobile phone through Bluetooth for the mobile phone to execute step S9.
[0099] S8: Voiceprint separation. In the case of the double-talk judgment result being yes, the second audio is subjected to voiceprint separation processing through a voiceprint separation model pre-set on the mobile phone to obtain a sound segment corresponding to the non-wearer sound source.
[0100] S9: Speech recognition. In the case of the double-talk judgment result being yes, the sound segment corresponding to the non-wearer obtained based on step S8 is subjected to speech recognition processing by using a speech recognition model to obtain a speech recognition result. In the case of the double-talk judgment result being no, the second audio obtained based on step S7 is directly subjected to speech recognition processing to obtain a speech recognition result. The step S9 can be executed by the mobile phone or sent to a cloud server by the mobile phone and the speech recognition result is obtained from the cloud server.
[0101] S10: Translation. The speech recognition result is translated according to a preset language to obtain a target text in the preset language. The step S10 can be executed by the mobile phone or sent to a cloud server by the mobile phone and the target text is obtained from the cloud server.
[0102] S11: Transmission of translation result. The mobile phone transmits the target text to the smart glasses through Bluetooth.
[0103] S12: Display. The smart glasses display the target text on a display interface.
[0104] Corresponding to the method provided in the embodiments of the present application, the embodiments of the present application also provide a data processing apparatus of a wearable device. The wearable device includes a microphone array, the microphone array including a first microphone and a second microphone, the target sound source of the first microphone being a wearer, and the target sound source of the second microphone being a non-wearer. As shown in Figure 4 The structure block diagram of the data processing apparatus of the embodiment of the present application can include:
[0105] The first detection module 401 is configured to perform first mute detection on the first audio collected by the first microphone to obtain a first mute detection result.
[0106] The second detection module 402 is configured to, in response to the first mute detection result being that there is a voice segment, perform second mute detection on the second audio collected by the second microphone to obtain a second mute detection result.
[0107] The sound source determination module 403 is configured to determine a sound source of the first audio and the second audio according to the first silence detection result and the second silence detection result.
[0108] The speech recognition module 404 is configured to perform speech recognition on audio collected by a microphone corresponding to the target sound source in the current application scenario in response to the sound source including the target sound source in the current application scenario.
[0109] In an embodiment, the first silence detection result includes a first confidence of the first audio; and in a case where the first confidence is greater than or equal to a first threshold, the first silence detection result is that there is a speech segment; and the sound source determination module is further configured to:
[0110] In response to the first confidence being greater than or equal to a second threshold, determine that the sound source includes the wearer sound source, the second threshold being greater than the first threshold.
[0111] In an embodiment, the second silence detection result includes a second confidence of the second audio; and the sound source determination module is further configured to:
[0112] In response to the second confidence being greater than or equal to a third threshold, determine that the sound source includes the non-wearer sound source, the third threshold being greater than the first threshold and less than the second threshold.
[0113] In an embodiment, the speech recognition module includes:
[0114] The enhancement sub-module is configured to, in response to the target sound source in the current application scenario being the non-wearer sound source and the sound source including the wearer sound source and the non-wearer sound source, perform directional speech enhancement on the second audio to obtain enhanced second audio.
[0115] The recognition sub-module is configured to perform speech recognition on the enhanced second audio.
[0116] In an embodiment, the recognition sub-module includes:
[0117] The voiceprint separation unit is configured to perform voiceprint separation processing on the enhanced second audio to obtain a sound segment emitted by the non-wearer.
[0118] The speech recognition unit is configured to input the sound segment emitted by the non-wearer into a speech recognition model to obtain a speech recognition result of the second audio.
[0119] In an embodiment, the recognition sub-module includes:
[0120] The interception unit is configured to intercept at least one speech frame with speech activity from the enhanced second audio according to the second silence detection result.
[0121] The voice recognition unit is configured to input the voice frame with voice activity into a voice recognition model to obtain a voice recognition result of the second audio.
[0122] In an embodiment, the voice recognition unit is further configured to:
[0123] send the voice frame with voice activity to a terminal device, and the terminal device is configured to run the voice recognition model to obtain the voice recognition result of the second audio;
[0124] receive the voice recognition result of the second audio fed back by the terminal device.
[0125] Figure 5 A block diagram of an electronic device for implementing the embodiments of the present application is shown in FIG. 10. As shown in FIG. 10, the electronic device includes a memory 510 and a processor 520, and the memory 510 stores a computer program executable on the processor 520. The processor 520 implements the method in the above embodiments when executing the computer program. The number of the memory 510 and the processor 520 can be one or more. Figure 5
[0126] The processor 520 can be a system on chip (SOC), and the processor 520 can include a central processing unit (CPU) and can further include other types of processors.
[0127] The processor 520 can include a CPU, a DSP, a microcontroller, or a digital signal processor, and can further include a GPU, an embedded neural network processing unit (NPU), and an image signal processor (ISP). The processor can further include necessary hardware accelerators or logic processing hardware circuits, such as an ASIC, or one or more integrated circuits for controlling the execution of the program of the technical solutions of the present application. In addition, the processor can have the function of operating one or more software programs, and the software programs can be stored in a storage medium.
[0128] The processor 520 can include one or more processing units. For example, the processor can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent components or integrated in one or more processors.
[0129] The memory 510 can include a code storage area and a data storage area. The code storage area can store an operating system. The data storage area can store data created during use of the electronic device, etc. In addition, the memory 510 can include a high-speed random access memory, and can further include a nonvolatile memory such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc.
[0130] The memory 510 can be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions, and can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical storage, a magnetic disk storage medium or other magnetic storage device, or can also be any computer-readable medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer.
[0131] The electronic device further includes:
[0132] The communication interface 530 is configured to communicate with external devices and transmit and receive data.
[0133] If the memory 510, the processor 520 and the communication interface 530 are implemented independently, the memory 510, the processor 520 and the communication interface 530 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 5 Only one thick line is used in the middle, but it does not mean that there is only one bus or one type of bus.
[0134] In some embodiments, the processor 520 can include one or more interfaces. The interface can include an Inter-Integrated Circuit (I2C) interface, an Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device, or can be used to transmit data between the electronic device and the peripheral device.
[0135] The electronic device can also include an external memory interface for connecting an external memory card, such as a Micro SD card, to achieve the expansion of the storage capacity of the electronic device. The external memory card communicates with the processor 520 through the external memory interface to achieve the data storage function. For example, save music, video, etc. Files in the external memory card.
[0136] Optionally, if the memory 510, the processor 520 and the communication interface 530 are integrated on a chip, the memory 510, the processor 520 and the communication interface 530 can complete the communication among each other through an internal interface.
[0137] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method provided in the embodiment of the present application.
[0138] The embodiment of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method of any one of the above.
[0139] The embodiment of the present application also provides a chip, which comprises a processor, and the processor is used to call and run instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.
[0140] The embodiment of the present application also provides a chip, which comprises an input interface, an output interface, a processor and a memory, and the input interface, the output interface, the processor and the memory are connected through an internal connection path, and the processor is used to execute code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiment of the present application.
[0141] It should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It should be noted that the processor can be a processor supporting an advanced RISC machine (ARM) architecture.
[0142] Further, the memory can optionally include a read-only memory and a random access memory. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. The non-volatile memory can include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory, among others. The volatile memory can include a random access memory (RAM), which is used as an external cache. By way of example, and not limitation, many forms of RAM are available. For example, a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a Sync Link DRAM (SLDRAM), and a direct Rambus RAM (DRRAM), among others.
[0143] In the above-described embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium.
[0144] In the description of the application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the application. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples. In addition, different embodiments or examples described in the specification and characteristics of different embodiments or examples can be combined and combined by those skilled in the art without contradiction.
[0145] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0146] Any process or method descriptions in flow charts or described elsewhere herein can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for implementing specific logic functions or steps in the process. And the scope of preferred embodiments of the present application includes additional implementation in which the functions are performed in different orders, in substantially simultaneous fashion, or in reverse order.
[0147] The logic and / or steps represented in flow charts or otherwise described herein, for example, can be considered as a sequence of executable instructions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- based system, or other system that can fetch instructions from a instruction execution system, apparatus, or device and execute the instructions, or in conjunction with such an instruction execution system, apparatus, or device.
[0148] It should be understood that parts of the present application can be realized in hardware, software, firmware or a combination thereof. In the above-described embodiments, a plurality of steps or methods can be realized by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above-described embodiment method can be instructed by a program to complete the relevant hardware, and the program can be stored in a computer-readable storage medium, and the program includes one or a combination of the steps of the method embodiment when executed.
[0149] In addition, each of the function units in each embodiment of the present application can be integrated in one processing module, or each unit can be physically present separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module. When the integrated module is realized in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. The storage medium can be a read-only memory, a magnetic disk or an optical disk, etc.
[0150] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various changes or replacements within the technical scope disclosed in the present application, and these should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data processing method of a wearable device, the wearable device comprising a microphone array, the microphone array comprising a first microphone and a second microphone, a target sound source of the first microphone being a wearer, a target sound source of the second microphone being a non-wearer; wherein the data processing method comprising: performing first silence detection on first audio collected by the first microphone to obtain a first silence detection result, the first silence detection result comprising a first confidence of the first audio; wherein, in a case where the first confidence is greater than or equal to a first threshold, the first silence detection result is that there is a speech segment; in response to the first silence detection result being that there is a speech segment, performing second silence detection on second audio collected by the second microphone to obtain a second silence detection result, the second silence detection result comprising a second confidence of the second audio; determining a sound source of the first audio and the second audio according to the first silence detection result and the second silence detection result; in response to the sound source comprising a target sound source in a current application scenario, performing speech recognition on audio collected by a microphone corresponding to the target sound source in the current application scenario; wherein determining the sound source of the first audio and the second audio according to the first silence detection result and the second silence detection result comprises: in response to the first confidence being greater than or equal to a second threshold, determining that the sound source comprises a wearer sound source, the second threshold being greater than the first threshold; in response to the second confidence being greater than or equal to a third threshold, determining that the sound source comprises a non-wearer sound source. 2.The data processing method of a wearable device according to claim 1, wherein, in response to the sound source comprising a target sound source in a current application scenario, performing speech recognition on audio collected by a microphone corresponding to the target sound source in the current application scenario comprises: in response to the target sound source in the current application scenario being a non-wearer sound source and the sound source comprising a wearer sound source and a non-wearer sound source, performing directional speech enhancement on the second audio to obtain enhanced second audio; performing speech recognition on the enhanced second audio. 3.The data processing method of a wearable device according to claim 2, wherein, performing speech recognition on the enhanced second audio comprises: performing voiceprint separation processing on the enhanced second audio to obtain a sound segment emitted by the non-wearer; inputting the sound segment emitted by the non-wearer into a speech recognition model to obtain a speech recognition result of the second audio. 4.The data processing method of a wearable device according to claim 2, wherein, performing speech recognition on the enhanced second audio comprises: cutting at least one speech frame with speech activity from the enhanced second audio according to the second silence detection result; inputting the speech frame with speech activity into a speech recognition model to obtain a speech recognition result of the second audio. 5.The data processing method of a wearable device according to claim 4, wherein, inputting the speech frame with speech activity into a speech recognition model to obtain a speech recognition result of the second audio comprises: sending the speech frame with speech activity to a terminal device, the terminal device being configured to run the speech recognition model to obtain the speech recognition result of the second audio; receiving the speech recognition result of the second audio fed back by the terminal device.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory, the processor implementing the method of any one of claims 1 to 5 when executing the computer program.
7. A computer readable storage medium having stored therein a computer program, the computer program implementing the method of any one of claims 1 to 5 when executed by a processor.
8. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Wearable sound source positioning tracking system and method
CN105223551A
Audio enhancement method and device, storage medium and wearable equipment
CN111935573A