Voice Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By adding post-processing logic and buffer length limitations in the interception process, the problem of missegment of speech segments is solved, and the robustness of interception of speech segments and the accuracy of awakening word judgment is improved.

CN114792530BActive Publication Date: 2025-07-04MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210450693.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-07-04
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

The prior art is prone to judge the end point of the speech segment too early in a quiet environment, resulting in the complete speech segment being misdivided into multiple segments, affecting the accuracy of the judgment of wake-up words.

Method used

By adding post-processing logic in the interception process of speech segment, increasing the judgment buffer length, limiting the time range when the speech validity detection result is an invalid frame, ensuring the integrity of the speech segment, and using audio intensity values ​​and zero crossing rates to comprehensively judge the effectiveness of speech frames.

Benefits of technology

It improves the robustness of voice clip interception, prevents missegment, and improves the accuracy of wake-up word judgment and the real-time system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792530B_ABST
    Figure CN114792530B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of voice data processing, and provides a voice data processing method, apparatus, electronic device, and storage medium. The method includes: determining that the voice data frame corresponding to the current moment is an invalid frame based on the voice validity detection result at the current moment, and obtaining the voice validity detection results at the first historical moment and the second historical moment; determining that the voice data frame corresponding to the first historical moment is an invalid frame based on the voice validity detection result at the first historical moment, and determining that the voice data frame corresponding to the second historical moment is a valid frame based on the voice validity detection result at the second historical moment, and determining the voice data frame corresponding to the first historical moment as the truncation end point of the target voice segment. By adding post-processing logic during the process of intercepting the valid voice segment, the present application restricts the end condition for intercepting the valid voice segment, thereby preventing the complete voice segment from being erroneously segmented into multiple segments due to short-term silence in the middle, and improving the robustness of intercepting the valid voice segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of voice data processing, and in particular, to a method, apparatus, electronic device, and storage medium for voice data processing. Background Art

[0002] Voice wake-up technology pre-sets a wake-up word in a device or software. When the user issues this voice command, the device is awakened from the sleep state and makes a specified response, greatly improving the efficiency of human-computer interaction. To protect user privacy, voice data cannot be uploaded before the device is awakened. Therefore, voice wake-up often needs to be implemented on a local device.

[0003] Limited by cost, local devices often have limitations in computing power and small memory space. To achieve low-power offline voice wake-up, not all voice signals can be directly subjected to the wake-up word judgment algorithm steps. Instead, after analyzing the voice signals, effective voice segments are extracted for wake-up word judgment.

[0004] By intercepting through effective voice segment detection and analyzing and judging the voice segments separately, not only is the amount of data and the amount of calculation greatly reduced, but it also helps to improve the wake-up rate and reduce the false wake-up rate. Using VAD (Voice Activity Detection) technology to detect the start point and end point of effective voice segments in the input voice signal, the effective voice segments can be intercepted and the voice signal can be analyzed and processed specifically.

[0005] However, when using the existing technology for effective voice segment detection, in some cases (especially in a quiet environment), it is easy to prematurely judge the end point of a voice segment, resulting in mis-segmenting a complete effective voice segment into several segments, affecting the accuracy of obtaining the wake-up word judgment result subsequently. Summary of the Invention

[0006] This application aims to at least solve one of the technical problems existing in the related art. For this purpose, this application proposes a method for voice data processing, which can improve the robustness of voice segment interception.

[0007] This application also proposes a voice data processing apparatus.

[0008] This application also proposes an electronic device.

[0009] This application also proposes a storage medium.

[0010] This application also proposes a computer program product.

[0011] According to an embodiment of the first aspect of this application, the method for voice data processing includes:

[0012] Determine that the speech data frame corresponding to the current moment based on the speech validity detection result of the current moment of the original speech segment is an invalid frame, and obtain the speech validity detection result of the first historical moment and the speech validity detection result of the second historical moment of the original speech segment;

[0013] Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation end point of the target speech segment;

[0014] Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

[0015] According to the speech data processing method of the embodiments of the present application, by adding post-processing logic during the process of intercepting the valid speech segment, the end condition for intercepting the valid speech segment is restricted, thereby preventing the complete speech segment from being erroneously segmented into multiple segments due to short-term silence in the middle, and improving the robustness of intercepting the valid speech segment.

[0016] According to an embodiment of the present application, there is at least one moment between the first historical moment and the current moment.

[0017] According to the speech data processing method of the embodiments of the present application, the determination basis for the truncation end condition of the valid speech segment is further restricted to that the speech validity detection results of the current moment and the first historical moment separated from the current moment by at least one moment are both invalid frames. By increasing the distance between the current moment and the first historical moment, the buffer length of the truncation point of the valid speech segment is increased, further preventing the complete speech segment from being erroneously segmented into multiple segments, and improving the robustness of intercepting the valid speech segment.

[0018] According to an embodiment of the present application, the determining that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determining that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determining the speech data frame corresponding to the first historical moment as the truncation end point of the target speech segment includes:

[0019] Determine, based on the voice validity detection result of the first historical moment, that the voice data frame corresponding to the first historical moment is an invalid frame, and determine, based on the voice validity detection result of the second historical moment, that the voice data frame corresponding to the second historical moment is a valid frame, and obtain at least one voice validity detection result corresponding to at least one moment between the first historical moment and the current moment;

[0020] Based on the at least one speech validity detection result, it is determined that all speech data frames corresponding to the at least one moment are invalid frames, and the speech data frame corresponding to the first historical moment is determined as the truncation endpoint of the target speech segment.

[0021] According to the voice data processing method of the embodiment of the present application, after determining that the voice data frames corresponding to all moments between the current moment and the first historical moment are invalid frames, the voice data frame corresponding to the first historical moment is determined as the truncation endpoint of the target voice segment, thereby further preventing the complete voice segment from being mistakenly divided into multiple segments and improving the robustness of the effective voice segment capture.

[0022] According to an embodiment of the present application, before determining that the voice data frame corresponding to the current moment is an invalid frame based on the voice validity detection result of the original voice segment at the current moment, and obtaining the voice validity detection result of the original voice segment at the first historical moment and the voice validity detection result of the original voice segment at the second historical moment, the method further includes:

[0023] Determine, based on the speech validity detection result at the target moment, a speech data frame corresponding to the target moment as a valid frame, and obtain the speech validity detection result at the third historical moment of the original speech segment;

[0024] Determining the voice data frame corresponding to the third historical moment as an invalid frame based on the voice validity detection result of the third historical moment, and determining the voice data frame corresponding to the target moment as the starting endpoint of the target voice segment;

[0025] The third historical moment is the moment before the target moment.

[0026] According to the speech data processing method of the embodiment of the present application, by determining whether the current moment is the starting point of the target speech segment based on the speech validity detection results at the current moment and the previous moment, the integrity of the starting part of the valid speech segment is ensured, and the robustness of the effective speech segment interception is further improved.

[0027] According to an embodiment of the present application, after determining that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determining that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and then determining the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment, the method further includes:

[0028] Intercepting the target speech segment from the original speech segment based on the start endpoint and the truncation endpoint.

[0029] According to the speech data processing method of the embodiment of the present application, by using the determined start endpoint and truncation endpoint as the basis for intercepting the valid speech segment, the target speech segment is intercepted from the original speech segment, further improving the robustness of intercepting the valid speech segment.

[0030] According to an embodiment of the present application, the number of moments between the first historical moment and the current moment is determined based on the length of the original speech segment; or,

[0031] The number of moments between the first historical moment and the current moment is determined based on the current scene mode of the system.

[0032] According to the speech data processing method of the embodiment of the present application, by adaptively determining the decision buffer length for intercepting the valid speech segment according to the length of the original speech segment or the preset application scenario mode, the robustness and flexibility of intercepting the valid speech segment are improved.

[0033] According to an embodiment of the present application, before determining that the speech data frame corresponding to the current moment of the original speech segment is an invalid frame based on the speech validity detection result of the current moment of the original speech segment, and obtaining the speech validity detection results of the first historical moment and the second historical moment of the original speech segment, the method further includes:

[0034] Determining the audio intensity value and zero-crossing rate of each speech data frame in the original speech segment;

[0035] Determining that the audio intensity value of the speech data frame is greater than the preset intensity threshold and the zero-crossing rate of the speech data frame is less than the preset zero-crossing rate threshold, and determining the speech validity detection result of the speech data frame as a valid frame mark;

[0036] Determining that the audio intensity value of the speech data frame is not greater than the preset intensity threshold or the zero-crossing rate of the speech data frame is not less than the preset zero-crossing rate threshold, and determining the speech validity detection result of the speech data frame as an invalid frame mark.

[0037] According to the voice data processing method of the embodiments of the present application, by calculating the audio intensity value and the zero-crossing rate of each voice data frame as the basis for judging the voice validity of the voice data frame, the accuracy of the validity detection of each voice data frame is improved, thereby further improving the robustness of the effective voice segment interception.

[0038] The voice data processing device according to the second aspect embodiment of the present application includes:

[0039] An acquisition module, configured to determine that the voice data frame corresponding to the current moment is an invalid frame based on the voice validity detection result of the current moment of the original voice segment, and acquire the voice validity detection result of the first historical moment of the original voice segment and the voice validity detection result of the second historical moment of the original voice segment;

[0040] A determination module, configured to determine that the voice data frame corresponding to the first historical moment is an invalid frame based on the voice validity detection result of the first historical moment, and determine that the voice data frame corresponding to the second historical moment is a valid frame based on the voice validity detection result of the second historical moment, and determine the voice data frame corresponding to the first historical moment as the truncation end point of the target voice segment;

[0041] Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target voice segment is one of the voice segments in the original voice segment.

[0042] The electronic device according to the third aspect embodiment of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the voice data processing method as described in any one of the above.

[0043] The non-transitory computer-readable storage medium according to the fourth aspect embodiment of the present application stores a computer program thereon, and when the computer program is executed by a processor, it implements the voice data processing method as described in any one of the above.

[0044] The computer program product according to the fifth aspect embodiment of the present application includes a computer program, and when the computer program is executed by a processor, it implements the voice data processing method as described in any one of the above.

[0045] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:

[0046] By adding post-processing logic during the interception of the effective voice segment and restricting the end condition for intercepting the effective voice segment, it is possible to prevent the complete voice segment from being erroneously segmented into multiple segments due to short-term silence in the middle, thereby improving the robustness of the effective voice segment interception.

[0047] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the accompanying drawings required for use in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 is a schematic flowchart of a voice data processing method provided by an embodiment of the present application;

[0050] Figure 2 is a schematic comparison diagram of original voice and effective voice;

[0051] Figure 3 is a schematic flowchart of a voice wake-up step provided by an embodiment of the present application;

[0052] Figure 4 is a schematic diagram of the processing result of intercepting an effective voice segment of the prior art;

[0053] Figure 5 is a schematic diagram of the processing result of intercepting an effective voice segment provided by an embodiment of the present application;

[0054] Figure 6 is a schematic structural diagram of a voice data processing device provided by an embodiment of the present application;

[0055] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following further describes in detail the embodiments of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0057] Please refer to Figure 1 , an embodiment of the present application provides a voice data processing method, which may include the steps:

[0058] S1. Determine that the speech data frame corresponding to the current moment based on the speech validity detection result at the current moment of the original speech segment is an invalid frame, and obtain the speech validity detection result at the first historical moment and the speech validity detection result at the second historical moment of the original speech segment; wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment;

[0059] It should be noted that in the process of speech processing, for example, the processing process of the speech wake-up technology mainly includes four steps: speech acquisition, VAD, wake-up word judgment, and result acquisition. The speech acquisition step will perform frame division on the acquired original speech signal to form multiple speech data frames, and subsequent modules will process the speech signal in the form of speech data frames. Among them, VAD is crucial for the accuracy of subsequent wake-up word judgment and result acquisition. By using VAD (Voice Activity Detection, also known as voice endpoint detection) technology to detect the start point and end point of the effective speech segment of the input speech signal, the effective speech segment can be intercepted and the speech signal can be analyzed and processed specifically. However, when using the existing technology to detect the effective speech segment, in some cases (especially in a quiet environment), it is easy to prematurely judge the end point of a speech, resulting in an entire effective speech segment being mis-segmented into several segments, affecting the accuracy of obtaining the wake-up word judgment result subsequently.

[0060] The embodiment of the present application improves the VAD module therein. On the basis of detecting the effective speech segment, by adding post-processing logic, the end condition of the speech segment is restricted, so as to avoid easily truncating the speech segment and maintain the integrity of the speech segment.

[0061] It should be noted that each speech data frame has a corresponding speech validity detection result, and this speech validity detection result can be detected by using the existing VAD technology. In the embodiment of the present application, the speech validity detection result of each speech data frame is one of two types: valid frame or invalid frame. Among them, one moment corresponds to one speech data frame, and one moment corresponds to one speech validity detection result.

[0062] It should be noted that during the process of determining the truncation point of the valid speech segment, if it is determined that the corresponding speech data frame is an invalid frame according to the speech validity detection result at the current moment, it is necessary to further obtain the speech validity detection results corresponding to the first historical moment and the second historical moment in the past. Among them, the first historical moment can be the Nth moment before the current moment, where N is a positive integer; the second historical moment is the moment before the first historical moment. For example, when N = 3, if the current moment is t(n), the first historical moment can be t(n - 3), and the second historical moment is t(n - 4). The value of N can be set according to requirements.

[0063] S2. Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment; where the target speech segment is one of the speech segments in the original speech segment.

[0064] In the embodiment of the present application, when it is further determined that the speech data frame corresponding to the first historical moment is an invalid frame, and at the same time it is determined that the speech data frame corresponding to the second historical moment is a valid frame, then the speech data frame corresponding to the first historical moment can be determined as the truncation endpoint of the target speech segment.

[0065] Since it is only possible (it still needs to be further determined according to the detection result of the second historical moment) to determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment only when it is determined that the speech data frames corresponding to the current moment and the first historical moment are both invalid frames, and the truncation is not performed immediately when the first invalid speech frame is detected, which increases the buffer segment before the truncation of the valid speech segment, and avoids the situation that the complete speech segment is mis-segmented into multiple segments due to the existence of short-time silence, effectively improving the robustness of the interception of the valid speech segment.

[0066] It should be noted that after the truncation endpoint is determined, the target speech segment (including the speech data frame corresponding to the truncation endpoint) can be intercepted, and the speech data frame corresponding to the current moment can be discarded and not used as the analysis object of the subsequent speech analysis module.

[0067] According to the speech data processing method of the embodiment of the present application, by adding post-processing logic during the interception of the valid speech segment and restricting the end condition for intercepting the valid speech segment, it is possible to prevent the complete speech segment from being mis-segmented into multiple segments due to the existence of short-time silence in the middle, and improve the robustness of the interception of the valid speech segment.

[0068] In one embodiment, there is at least one moment between the first historical moment and the current moment.

[0069] It should be noted that the first historical moment can be the second moment before the current moment. For example, if the current moment is t(n), the first historical moment is t(n - 2), and there is a moment t(n - 1) between the first historical moment and the current moment.

[0070] According to the voice data processing method of the embodiments of the present application, the determination basis for the truncation end condition of the valid voice segment is further restricted to that the voice validity detection results of the current moment and the first historical moment separated from the current moment by at least one moment are both invalid frames. By increasing the distance between the current moment and the first historical moment, the determination buffer length of the truncation point of the valid voice segment is increased, further preventing a complete voice segment from being mis-segmented into multiple segments and improving the robustness of the interception of the valid voice segment.

[0071] In one embodiment, step S2 may include the steps of:

[0072] S21. Based on the voice validity detection result of the first historical moment, determine that the voice data frame corresponding to the first historical moment is an invalid frame, and based on the voice validity detection result of the second historical moment, determine that the voice data frame corresponding to the second historical moment is a valid frame, and obtain at least one voice validity detection result corresponding to at least one moment between the first historical moment and the current moment;

[0073] S22. Based on the at least one voice validity detection result, determine that all voice data frames corresponding to the at least one moment are invalid frames, and determine the voice data frame corresponding to the first historical moment as the truncation end point of the target voice segment.

[0074] It should be noted that on the premise that there is at least one moment between the first historical moment and the current moment, the following situation may exist: the voice validity detection results of the current moment, the first historical moment, and the second historical moment all meet the determination conditions of the truncation end point of the target voice segment, but there are valid frames in the voice data frames corresponding to at least one moment between the current moment and the first historical moment. Then the current moment is actually very likely to be just a silent stage in a complete voice segment. If the corresponding voice data frame is immediately used as the truncation end point, it is still too early to judge the end of a piece of voice, resulting in mis-segmenting a complete valid voice segment into multiple segments.

[0075] In order to overcome the problems existing in the above situation, in the embodiments of the present application, it is further determined whether all speech data frames corresponding to at least one moment between the first historical moment and the current moment are valid. Only when these speech data frames are all invalid frames, the speech data frame corresponding to the first historical moment is considered as the truncation endpoint of the target speech segment.

[0076] According to the speech data processing method of the embodiments of the present application, only when it is determined that all speech data frames corresponding to all moments between the current moment and the first historical moment are invalid frames, the speech data frame corresponding to the first historical moment is determined as the truncation endpoint of the target speech segment, further preventing a complete speech segment from being mis-segmented into multiple segments and improving the robustness of intercepting valid speech segments.

[0077] In one embodiment, before step S1, the following steps may further be included:

[0078] S11. Based on the speech validity detection result of the target moment, determine that the speech data frame corresponding to the target moment is a valid frame, and obtain the speech validity detection result of the third historical moment of the original speech segment;

[0079] S12. Based on the speech validity detection result of the third historical moment, determine that the speech data frame corresponding to the third historical moment is an invalid frame, and determine the speech data frame corresponding to the target moment as the starting endpoint of the target speech segment;

[0080] Wherein, the third historical moment is the moment before the target moment.

[0081] In the embodiments of the present application, before the starting endpoint of the target speech segment is marked (if the starting endpoint has been marked, the determination process of the truncation endpoint is entered), if it is determined that the speech data frame corresponding to the target moment is a frame that changes from invalid to valid, the speech data frame corresponding to the target moment is directly determined as the starting endpoint of the target speech segment.

[0082] According to the speech data processing method of the embodiments of the present application, by determining whether the target moment is the starting point of the target speech segment according to the speech validity detection results of the target moment and the previous moment, the integrity of the starting part of the valid speech segment is ensured, and the robustness of intercepting the valid speech segment is further improved.

[0083] In one embodiment, after step S2, the following steps may further be included:

[0084] S23. Intercept the target speech segment from the original speech segment based on the starting endpoint and the truncation endpoint.

[0085] It should be noted that after determining the start endpoint and the truncation endpoint, the target speech segment can be intercepted from the original speech segment according to the start endpoint and the truncation endpoint. Based on the determined start endpoint and truncation endpoint, the target speech segment can be intercepted (including the speech data frames corresponding to the start endpoint and the truncation endpoint). It can be understood that the speech data frame corresponding to the current moment can be discarded and not used as the analysis object of the subsequent speech analysis module.

[0086] According to the speech data processing method of the embodiments of the present application, the start endpoint and the truncation endpoint determined by the embodiments of the present application are used to intercept the target speech segment, rather than using the endpoints directly determined by the existing VAD detection technology as the basis for intercepting the target speech segment, effectively preventing the complete speech segment from being mis-segmented into multiple segments and improving the robustness of intercepting the effective speech segment.

[0087] In one embodiment, the number of moments between the first historical moment and the current moment is determined based on the length of the original speech segment; or,

[0088] The number of moments between the first historical moment and the current moment is determined based on the current system scene mode.

[0089] It should be noted that the first historical moment can be the Nth moment before the current moment, where N is a positive integer. For example, when N = 3, if the current moment is t(n), the first historical moment can be t(n - 3). At this time, the number of moments between the first historical moment and the current moment is 2, that is, two moments t(n - 1) and t(n - 2) are separated. It can be understood that the corresponding relationship between the original speech length and the number of separated moments can be pre-configured, or the corresponding relationship between the system scene mode and the number of separated moments can be pre-configured; before intercepting the effective speech segment, the number of separated moments between the first historical moment and the current moment can be adaptively determined by obtaining the length of the original speech or the current system scene mode this time, without manual intervention.

[0090] According to the speech data processing method of the embodiments of the present application, by adaptively determining the determination buffer length for intercepting the effective speech segment this time according to the length of the original speech segment or the preset application scenario mode, the robustness and flexibility of intercepting the effective speech segment are improved.

[0091] In one embodiment, before step S1, the following steps may further be included:

[0092] S13. Determine the audio intensity value and the zero-crossing rate of each speech data frame in the original speech segment;

[0093] S14. Determine that the audio intensity value of the voice data frame is greater than a preset intensity threshold and the zero-crossing rate of the voice data frame is less than a preset zero-crossing rate threshold, and determine the voice validity detection result of the voice data frame as a valid frame flag;

[0094] S15. Determine that the audio intensity value of the voice data frame is not greater than the preset intensity threshold, or the zero-crossing rate of the voice data frame is not less than the preset zero-crossing rate threshold, and determine the voice validity detection result of the voice data frame as an invalid frame flag.

[0095] In the embodiments of the present application, for each voice data frame, by calculating the audio intensity value and the zero-crossing rate of the frame and comparing them with the preset thresholds respectively, the voice validity detection result of the frame is determined. Compared with other voice validity detection methods (such as only determining the validity of the corresponding voice frame by the audio intensity value), in the embodiments of the present application, by using both the audio intensity value and the zero-crossing rate as the basis for judging the voice validity of the voice data frame, the accuracy of the validity detection of each voice data frame is improved, thereby further improving the robustness of the effective voice segment extraction.

[0096] Please refer to Figures 2 - 5 , based on the above solution, to better understand the voice data processing method provided by the embodiments of the present application, the following takes the processing of voice wake-up data as an example for detailed description:

[0097] It should be noted that voice wake-up often needs to be implemented on a local device. However, due to cost limitations, local devices often have limitations such as insufficient computing power and small memory space. In order to achieve low-power offline voice wake-up, not all voice signals can be directly subjected to the wake-up word judgment algorithm steps. Instead, after analyzing the voice signals, the effective voice segments are extracted for wake-up word judgment.

[0098] As Figure 2 shown, this figure is a waveform diagram of a voice signal. The larger the amplitude, the louder the volume. The upper part of the figure is the original voice, including voice segments and silent segments; the lower part of the figure is the intercepted effective voice segment. It can be understood that the amount of original voice data is large, but a large amount of the data does not contain valid voice data. These data do not contain voice semantic information and are usually environmental noise without any valid information such as wake-up words. By intercepting through effective voice segment detection and analyzing and judging the voice segments separately, not only the amount of data and the amount of calculation are greatly reduced, but also it helps to improve the wake-up rate and reduce the false wake-up rate.

[0099] As Figure 3 shown, the voice wake-up technology mainly includes four steps: voice acquisition, VAD, wake-up word judgment, and result acquisition. Among them, VAD is crucial for the accuracy of subsequent wake-up word judgment and result acquisition.

[0100] The VAD (Voice Activity Detection) technology is used to detect the start point and end point of the effective speech segment of the input speech signal, so that the effective speech segment can be intercepted, and the speech signal can be analyzed and processed specifically to obtain a more accurate wake-up result.

[0101] The effective speech segment detection scheme often comprehensively judges the start point and end point of the speech segment through the energy of the speech signal (audio intensity value) and the zero-crossing rate of the speech signal. For example, for a frame of speech signal (such as 32 milliseconds), when the calculated audio intensity value is greater than the preset intensity threshold and the zero-crossing rate is less than the preset zero-crossing rate threshold at the same time, it is considered that the speech frame is in the effective speech segment.

[0102] However, in actual applications, when using the above method to detect the effective speech segment, in some cases (especially in a quiet environment), it is easy to prematurely judge the end of a speech segment, resulting in splitting an effective speech segment into several segments.

[0103] Such as Figure 4 shown, this figure shows the situation where a complete speech is intercepted into two segments. In the figure, t0 to t32 represent the speech segments obtained by the system at each moment. For example, at the t0 moment, it represents the speech segment from the 0th to 31st ms (a speech data frame), and at the t1 moment, it represents the speech segment from the 32nd to 63rd ms. When intercepting the effective speech segment, first perform VAD judgment on the speech segment. The judgment result of 1 indicates that the speech frame is an effective frame, otherwise it is an invalid frame.

[0104] Based on the current moment t(n) and the previous moment t(n - 1), the start and end of the speech segment are judged.

[0105] For example, at the t4 moment, the VAD result is 1, and at the t3 moment, the VAD result is 0. This is a change from 0 to 1, which is considered a sign of the start of the speech segment.

[0106] For example, at the t18 moment, the VAD result is 0, and at the t17 moment, the VAD result is 1. This is a change from 1 to 0, which is considered a sign of the end of the speech segment.

[0107] In this way, the start and end of the effective speech segment can be judged through the change of the VAD result.

[0108] When using the above-mentioned existing technology for effective speech segment interception, since the VAD detection results of the two frames of speech at t18 and t19 are invalid speech, the overall piece of speech is split into two segments (t4 - t17 and t20 - t26). The voice wake-up model will then analyze the two segments of speech separately. For example, the original "Xiaomei Xiaomei" is truncated into "Xiaomei Xiao" and "Mei", which affects the subsequent wake-up word judgment and the accuracy of the finally obtained wake-up result.

[0109] The purpose of the embodiment of this application is to solve the problem that a complete piece of speech is easily split during the process of effective speech segment detection. Without affecting the real-time performance of the entire system, by adding a judgment condition for the end of the speech segment, the robustness of effective speech segment interception is improved, and thus the wake-up effect is enhanced.

[0110] The speech data processing method provided by the embodiment of this application restricts the end condition of the speech segment. It judges the end of the speech segment through time t(n), time t(n - t) and time t(n - t - 1) (in this embodiment, t is taken as 3).

[0111] The end of the speech segment needs to meet three conditions:

[0112] 1. The VAD result at the current time t(n) is 0;

[0113] 2. The VAD result at time t(n - 3) is 0;

[0114] 3. The VAD result at time t(n - 4) is 1.

[0115] In the embodiment of this application, a VAD result of 0 represents that the speech validity detection result is an invalid frame, and a VAD result of 1 represents that the speech validity detection result is a valid frame. In other embodiments, other speech validity detection methods other than VAD can also be used, and other validity representation methods other than 0 and 1 can also be used.

[0116] As Figure 5 shown, the VAD result at time t4 is 1, and the VAD result at time t3 is 0, then the current speech segment is in the start state, and the frame data is sent to the wake-up word judgment module;

[0117] ……

[0118] The VAD result at time t18 is 0, and the VAD result at time t15 is 1, which does not meet condition 2, so the current speech segment is in the active state, and the frame data is sent to the wake-up word judgment module;

[0119] The VAD result at time t19 is 0, and the VAD result at time t16 is 1, which does not meet condition 2, so the current speech segment is in the active state, and the frame data is sent to the wake-up word judgment module;

[0120] At time t20, the VAD result is 1, which does not meet Condition 1. Then the current speech segment is in the active state, and the frame data is sent to the wake word judgment module.

[0121] ……

[0122] At time t30, the VAD result is 0, at time t27 the VAD result is 0, and at time t26 the VAD result is 1, meeting 3 conditions. Then the current speech segment is in the end state, and the frame data is discarded.

[0123] At time t31, the VAD result is 0, at time t28 the VAD result is 0, and at time t27 the VAD result is 0, which does not meet Condition 3. Then the current speech segment is in the non-active state, and the frame data is discarded.

[0124] As can be seen from the above, the effective speech segment intercepted this time is the speech frames corresponding to t4 - t29. By imposing constraints on the end condition, the speech analysis module (wake word judgment module) can analyze the overall data from t4 to t29, thus avoiding the previous situation of splitting a complete speech into two parts.

[0125] Combining the above operations, the optimizations we have made can be summarized as follows:

[0126] Assume that the speech segment being processed at the current moment is t(n). After analysis by VAD, t(n)=0 indicates that this segment of speech is invalid speech, and t(n)=1 indicates that this segment of speech is valid speech. Then the corresponding pseudo-code can be:

[0127] if t(n) = 1

[0128] if t(n - 1) = 0

[0129] then the current speech segment is the start point, and the frame data is sent to the wake word judgment module

[0130] else the current speech segment is in the active state, and the frame data is sent to the wake word judgment module

[0131] else if t(n) = 0

[0132] if t(n - t) = 0

[0133] if t(n - t - 1) = 1

[0134] then the current speech segment is the end point, and the frame data is discarded

[0135] else

[0136] the current speech segment is in the non-active state, and the frame data is discarded

[0137] else

[0138] The current voice segment is in an active state, and the frame data is sent to the wake word judgment module.

[0139] Compared with the prior art, by implementing the embodiments of the present application, on the basis of being able to perform effective voice segment detection, the robustness of intercepting effective voice segments is improved, and further the accuracy of subsequent processing is improved. Such an advantage comes from that on the basis of performing voice validity detection, by adding post-processing logic, the end condition for intercepting the target voice segment is restricted, so as to avoid easily truncating the voice segment and being able to maintain the integrity of the voice segment.

[0140] In addition, it should be noted that the voice data processing method of the embodiments of the present application does not introduce additional time delay and computational complexity, that is, it can complete the restriction of the end of the voice segment without generating frame delay, thus not affecting the real-time performance of the entire system. Such an advantage comes from sacrificing the cost of adding data at the end of the voice segment in exchange for not generating frame delay, because real-time performance is more important for the system.

[0141] Reference Figure 6 , Figure 6 is a schematic diagram of the modules of the voice data processing device provided by the embodiments of the present application. The voice data processing device provided by the embodiments of the present application includes:

[0142] An acquisition module 1, configured to determine that the voice data frame corresponding to the current moment is an invalid frame based on the voice validity detection result of the current moment of the original voice segment, and acquire the voice validity detection result of the first historical moment of the original voice segment and the voice validity detection result of the second historical moment of the original voice segment;

[0143] A determination module 2, configured to determine that the voice data frame corresponding to the first historical moment is an invalid frame based on the voice validity detection result of the first historical moment, and determine that the voice data frame corresponding to the second historical moment is a valid frame based on the voice validity detection result of the second historical moment, and determine the voice data frame corresponding to the first historical moment as the truncation endpoint of the target voice segment;

[0144] Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target voice segment is one of the voice segments in the original voice segment.

[0145] In one embodiment, there is at least one moment between the first historical moment and the current moment.

[0146] In one embodiment, the determination module 2 is specifically configured to:

[0147] Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and obtain at least one speech validity detection result corresponding to at least one moment between the first historical moment and the current moment;

[0148] Based on the at least one speech validity detection result, determine that all speech data frames corresponding to the at least one moment are invalid frames, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment.

[0149] In one embodiment, the speech data processing device further includes a start determination module, which is configured to:

[0150] Determine that the speech data frame corresponding to the target moment is a valid frame based on the speech validity detection result of the target moment, and obtain the speech validity detection result of the third historical moment of the original speech segment;

[0151] Based on the speech validity detection result of the third historical moment, determine that the speech data frame corresponding to the third historical moment is an invalid frame, and determine the speech data frame corresponding to the target moment as the start endpoint of the target speech segment;

[0152] Wherein, the third historical moment is the previous moment of the target moment.

[0153] In one embodiment, the speech data processing device further includes a truncation module, which is configured to:

[0154] Obtain the target speech segment by truncating from the original speech segment based on the start endpoint and the truncation endpoint.

[0155] In one embodiment, the speech data processing device further includes a detection module, which is configured to:

[0156] Determine the audio intensity value and the zero-crossing rate of each speech data frame in the original speech segment;

[0157] Determine that the audio intensity value of the speech data frame is greater than a preset intensity threshold, and the zero-crossing rate of the speech data frame is less than a preset zero-crossing rate threshold, and determine the speech validity detection result of the speech data frame as a valid frame mark;

[0158] Determine that the audio intensity value of the speech data frame is not greater than the preset intensity threshold, or the zero-crossing rate of the speech data frame is not less than the preset zero-crossing rate threshold, and determine the speech validity detection result of the speech data frame as an invalid frame mark.

[0159] Figure 7Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the following method:

[0160] S1. Determine that the speech data frame corresponding to the current moment of the original speech segment is an invalid frame based on the speech validity detection result at the current moment, and obtain the speech validity detection result at the first historical moment and the speech validity detection result at the second historical moment of the original speech segment;

[0161] S2. Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result at the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result at the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation end point of the target speech segment;

[0162] Among them, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

[0163] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.

[0164] On the other hand, an embodiment of the present application discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments, for example, including:

[0165] S1. Determine that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result of the current moment of the original speech segment, and obtain the speech validity detection result of the first historical moment and the speech validity detection result of the second historical moment of the original speech segment;

[0166] S2. Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment;

[0167] Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

[0168] On another aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the transmission methods provided in the above-mentioned embodiments, for example, including:

[0169] S1. Determine that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result of the current moment of the original speech segment, and obtain the speech validity detection result of the first historical moment and the speech validity detection result of the second historical moment of the original speech segment;

[0170] S2. Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment;

[0171] Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application, and should all be covered by the scope of the claims of the present application.

Claims

1. A method for processing voice data, characterized in that, Including: Determine that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result of the current moment of the original speech segment, and obtain the speech validity detection result of the first historical moment and the speech validity detection result of the second historical moment of the original speech segment; Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment; Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

2. The voice data processing method according to claim 1, characterized in that , there is at least one moment between the first historical moment and the current moment.

3. The voice data processing method according to claim 2, wherein The determining that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determining that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and determining the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment includes: Determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result of the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result of the second historical moment, and obtain at least one speech validity detection result corresponding to at least one moment between the first historical moment and the current moment; Determine that all speech data frames corresponding to the at least one moment are invalid frames based on the at least one speech validity detection result, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment.

4. The voice data processing method according to claim 1, wherein Before determining that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result of the current moment of the original speech segment, and obtaining the speech validity detection result of the first historical moment and the speech validity detection result of the second historical moment of the original speech segment, it further includes: Determine that the speech data frame corresponding to the target moment is a valid frame based on the speech validity detection result of the target moment, and obtain the speech validity detection result of the third historical moment of the original speech segment; Determine that the speech data frame corresponding to the third historical moment is an invalid frame based on the speech validity detection result of the third historical moment, and determine the speech data frame corresponding to the target moment as the starting endpoint of the target speech segment; Wherein, the third historical moment is the moment before the target moment.

5. The voice data processing method according to claim 4, wherein After determining that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result at the first historical moment, and determining that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result at the second historical moment, and then determining the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment, the following steps are further included: Intercepting the target speech segment from the original speech segment based on the start endpoint and the truncation endpoint.

6. The voice data processing method according to claim 1, wherein The number of moments between the first historical moment and the current moment is determined based on the length of the original speech segment; or, The number of moments between the first historical moment and the current moment is determined based on the current scene mode of the system.

7. The voice data processing method according to any one of claims 1-6, characterized in that Before determining that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result at the current moment of the original speech segment, and obtaining the speech validity detection results at the first historical moment and the second historical moment of the original speech segment, the following steps are further included: Determining the audio intensity value and zero-crossing rate of each speech data frame in the original speech segment; Determining that the audio intensity value of the speech data frame is greater than the preset intensity threshold, and the zero-crossing rate of the speech data frame is less than the preset zero-crossing rate threshold, and determining the speech validity detection result of the speech data frame as a valid frame mark; Determining that the audio intensity value of the speech data frame is not greater than the preset intensity threshold, or the zero-crossing rate of the speech data frame is not less than the preset zero-crossing rate threshold, and determining the speech validity detection result of the speech data frame as an invalid frame mark.

8. A voice data processing device, characterized in that, Including: An acquisition module, configured to determine that the speech data frame corresponding to the current moment is an invalid frame based on the speech validity detection result at the current moment of the original speech segment, and obtain the speech validity detection results at the first historical moment and the second historical moment of the original speech segment; A determination module, configured to determine that the speech data frame corresponding to the first historical moment is an invalid frame based on the speech validity detection result at the first historical moment, and determine that the speech data frame corresponding to the second historical moment is a valid frame based on the speech validity detection result at the second historical moment, and determine the speech data frame corresponding to the first historical moment as the truncation endpoint of the target speech segment; Wherein, the first historical moment is a certain moment before the current moment, and the second historical moment is the moment before the first historical moment; the target speech segment is one of the speech segments in the original speech segment.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the speech data processing method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the speech data processing method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice signal detection method and device

    CN106887241A

  • Voice recognition method and device for smart home

    CN110853631A

  • Voice wake-up method and device, chip, electronic equipment and storage medium

    CN112951243A