Voice endpoint detection methods and related apparatuses, devices, and media
Patent Information
- Application Number
- CN202511413518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-09-29
AI Technical Summary
但是,现有技术在面对复杂环境时语音端点检测的准确性并不理想
[0009]The above scheme uses an endpoint detection model to detect the audio to be detected, obtaining the probability value of each audio frame in the audio belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The endpoint detection model detects each first target frame in the first target audio and obtains the predicted probability that it belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. These sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and these sub-losses are based on the alignment information and the predicted probability. The probability is measured, and based on this, the final frame type of the audio frame is determined from several preset frame types. Since the endpoint detection model combines alignment information and predicted probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish sub-losses for these categories in the training loss of the endpoint detection model. Furthermore, by minimizing the training loss during training, the endpoint detection model can improve detection accuracy for all categories, including difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps improve the endpoint detection model's ability to detect difficult samples in complex environments such as those with complex, high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments.
Smart Images

Figure CN121438875B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech endpoint detection method and related apparatus, equipment and medium. Background Technology
[0002] In the field of intelligent speech processing, effective speech detection (Voice Activity Detection, also known as speech endpoint detection, speech activity detection, etc.) is usually required to determine the start and end points of effective speech from an audio signal in order to distinguish human voices from non-human voices.
[0003] Existing VAD (Voice Endpoint Detection) technologies include correlation methods based on temporal features, correlation methods based on transform domains, correlation methods based on template matching, and correlation methods based on neural network models, among others. However, the accuracy of existing technologies in speech endpoint detection is not ideal when facing complex environments. Therefore, improving the accuracy of speech endpoint detection in complex environments has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide a voice endpoint detection method and related apparatus, equipment and medium, which can improve the accuracy of voice endpoint detection in complex environments.
[0005] To address the aforementioned technical problems, the first aspect of this application provides a speech endpoint detection method, comprising: detecting audio to be detected based on an endpoint detection model to obtain probability values of audio frames in the audio to be detected belonging to a target frame type; wherein, the target frame type is at least one of several preset frame types, the several preset frame types include human voice and non-human voice, the endpoint detection model is trained based on a first target audio, the first target audio is labeled with frame-level alignment information between itself and the corresponding first audio text, the first target audio is detected by the endpoint detection model to obtain the predicted probability that each first target frame in the first target audio belongs to the target frame type, the first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability, the several sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, the training loss of the endpoint detection model includes sub-losses of each sample set, and the sub-losses are measured based on the alignment information and the predicted probability; and determining the final frame type of the audio frame among the several preset frame types based on the probability values of the audio frame belonging to the target frame type.
[0006] To address the aforementioned technical problems, a second aspect of this application provides a speech endpoint detection device, comprising: an endpoint detection module and a type determination module. The endpoint detection module is used to detect audio to be detected based on an endpoint detection model, obtaining probability values of audio frames in the audio to be detected belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is annotated with frame-level alignment information with a corresponding first audio text. The first target audio is detected by the endpoint detection model to obtain the predicted probability that each first target frame in the first target audio belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. The several sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and the sub-losses are measured based on the alignment information and the predicted probability. The type determination module is used to determine the final frame type of the audio frame among several preset frame types based on the probability values of the audio frame belonging to the target frame type.
[0007] To address the aforementioned technical problems, a third aspect of this application provides an electronic device comprising at least a memory and a processor coupled to each other, wherein the memory stores at least program instructions, and the processor executes the program instructions to implement the voice endpoint detection method of the first aspect described above.
[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being used to implement the voice endpoint detection method of the first aspect described above.
[0009] The above scheme uses an endpoint detection model to detect the audio to be detected, obtaining the probability value of each audio frame in the audio belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The endpoint detection model detects each first target frame in the first target audio and obtains the predicted probability that it belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. These sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and these sub-losses are based on the alignment information and the predicted probability. The probability is measured, and based on this, the final frame type of the audio frame is determined from several preset frame types. Since the endpoint detection model combines alignment information and predicted probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish sub-losses for these categories in the training loss of the endpoint detection model. Furthermore, by minimizing the training loss during training, the endpoint detection model can improve detection accuracy for all categories, including difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps improve the endpoint detection model's ability to detect difficult samples in complex environments such as those with complex, high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating an embodiment of the voice endpoint detection method of this application; Figure 2 This is a schematic diagram of the framework of an embodiment of the voice endpoint detection device of this application; Figure 3 This is a schematic diagram of the framework of an embodiment of the electronic device of this application; Figure 4 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0011] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0012] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0013] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the slash " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper indicates two or more objects.
[0014] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the voice endpoint detection method of this application. Specifically, it may include the following steps: Step S11: Detect the audio to be detected based on the endpoint detection model and obtain the probability value of the audio frame in the audio to be detected belonging to the target frame type.
[0015] In this embodiment of the disclosure, the target frame type can be at least one of several preset frame types, which may include human voice and non-human voice. For example, the target frame type can be human voice; or, for another example, the target frame type can also be non-human voice; or, for yet another example, the target frame type can include both human voice and non-human voice. The specific type of the target frame type is not limited here. It should be noted that an audio frame with a human voice type indicates that the audio frame contains valid speech information (i.e., the speaker's voice), while an audio frame with a non-human voice type indicates that the audio frame is a silence frame, a noise frame (e.g., background noise, white noise), etc.
[0016] In this embodiment, the endpoint detection model can be trained based on a first target audio, which can be labeled with frame-level alignment information between itself and the corresponding first audio text. That is, the first target audio can be labeled with the start and end frames of each word in the first audio text within the first target audio. Taking the first target audio "The weather is nice today" as an example, its corresponding first audio text is the identical text data "The weather is nice today," where "today" starts at frame i and ends at frame i+N in the first target audio; "weather" starts at frame j and ends at frame j+M; and "nice" starts at frame k and ends at frame k+P. Of course, the above example is merely one possible instance of alignment information in practical applications; other possible scenarios will not be listed here. Furthermore, the endpoint detection model detects the first target audio, yielding a predicted probability that each first target frame in the audio belongs to a target frame type. The first target frame can be categorized into several sample sets based on alignment information and prediction probabilities. These sample sets can include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. Specifically, the sample sets can be categorized into four sets: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The "difficult human voice samples" set encompasses first target frames that belong to human voices and are difficult to identify as human voices; the "difficult non-human voice samples" set encompasses first target frames that belong to non-human voices and are difficult to identify as non-human voices; the "easy human voice samples" set encompasses first target frames that belong to human voices and are relatively easy to identify as human voices; and the "easy non-human voice samples" set encompasses first target frames that belong to non-human voices and are relatively easy to identify as non-human voices. The training loss of the endpoint detection model, trained based on the first target audio, can include sub-losses for each sample set, and these sub-losses can be measured based on alignment information and prediction probabilities.
[0017] In one implementation scenario, the endpoint detection model can be applied to the target business scenario. The first target audio can be obtained by superimposing the target non-human voice audio segment of the target business scenario onto the first sample audio with the target signal-to-noise ratio. Furthermore, the first sample audio can be labeled with alignment information. This eliminates the need for time-consuming and laborious collection of noisy labeled audio data for the target business scenario. Instead, conventional (e.g., open-source) speech recognition data can be collected, and then noisy labeled audio data for the target business scenario can be added using the target non-human voice audio segment. This simulates the noisy labeled audio data of the target business scenario, significantly reducing the collection and production costs of training data while ensuring the quality of the training data as much as possible. Moreover, it should be noted that the target business scenario can be set according to actual application needs. For example, in the field of intelligent vehicles, the target business scenario can be a vehicle cabin scenario; in the field of smart homes, it can be a home scenario; in the field of smart classrooms, it can be a classroom scenario; or in the field of smart offices, it can be a meeting scenario. Of course, the above examples are only a few possible examples of target business scenarios in practical applications; other scenarios will not be listed here.
[0018] In a specific implementation scenario, the first sample audio can be specifically labeled with corresponding first audio text. For example, the first sample audio "The weather is so nice today" can be labeled with the same first audio text "The weather is so nice today". Based on this, forced alignment (FA) can be performed on the first sample audio with the above-mentioned annotation information to obtain word-level labeled time segments (i.e., the alignment information mentioned above). Taking the aforementioned first sample audio as an example, the word-level labeled time segments can specifically include: the start and end frames of "today" in the first sample audio, the start and end frames of "weather" in the first sample audio, and the start and end frames of "so nice" in the first sample audio. Other cases can be deduced similarly, and will not be listed one by one here. It should be noted that forced alignment can be achieved through speech recognition models such as HMM-DNN-ASR (Hidden Markov Model-Deep Neural Network-Automatic Speech Recognition) and HMM-BiLSTM-ASR (Hidden Markov Model-Bi Long Short Term Memory-Automatic Speech Recognition). The technical details of the aforementioned speech recognition models will not be elaborated here.
[0019] In a specific implementation scenario, to obtain the target non-human voice audio segment, as a possible implementation example, it can be directly segmented based on the target non-human voice audio of the target business scenario to obtain the target non-human voice audio segment. For example, the target non-human voice audio can be randomly segmented, such as into short segments of 10 to 20 seconds, which can then be used as the target non-human voice audio segment. It should be noted that the target non-human voice audio can be collected within the target business scenario. For example, a small amount of non-human voice audio (e.g., 20 hours) can be collected within the target business scenario to serve as the target non-human voice audio. Alternatively, as another possible implementation example, to obtain the target non-human voice audio segment, it can be segmented based on the target non-human voice audio of the target business scenario to obtain several candidate non-human voice audio segments. Then, based on a first energy threshold, the temporal energy of each non-human voice frame in the candidate non-human voice audio segments is statistically analyzed to obtain the target frame count. For example, the total number of non-human voice frames in the candidate non-human voice audio segments with a temporal energy lower than the first energy threshold can be counted as the target frame count. For ease of description, the first energy threshold can be denoted as energy_thred. The total number of non-human voice frames in the candidate non-human voice audio segments whose temporal energy is lower than the first energy threshold energy_thred, i.e., the target frame count, can be denoted as lower_num. Based on this, the ratio between the target frame count and the audio segment duration of the candidate non-human voice audio segments can be obtained as the average occurrence count. In other words, the average occurrence count can be expressed as: lower_per_sec=lower_num / wav_len In the above formula, lower_per_sec represents the average number of occurrences, and wav_len represents the duration of the audio segment in the candidate non-human voice frequency band. After obtaining the average number of occurrences, the average number of occurrences can be compared with an occurrence number threshold to determine whether to select the candidate non-human voice frequency band as the target non-human voice frequency band. For example, if the average number of occurrences is less than the occurrence number threshold, the candidate non-human voice frequency band can be selected as the target non-human voice frequency band; otherwise, the candidate non-human voice frequency band may not be selected as the target non-human voice frequency band. Still taking the foregoing example as an example, for ease of description, the occurrence number threshold may be denoted as appear_num_thred. If lower_per_sec < appear_num_thred, the candidate non-human voice frequency band can be selected as the target non-human voice frequency band. In the above method, a series of operations including audio segmentation, energy statistics and occurrence statistics are used to determine whether to select the candidate non-human voice frequency band as the target non-human voice frequency band for subsequent noise addition, which can ensure as much as possible that the audio attribute of the original human voice in the first sample audio will not be changed due to an excessive noise addition coefficient during subsequent noise addition. For ease of understanding, taking two candidate non-human voice frequency bands as an example, wherein the average number of occurrences of the first candidate non-human voice frequency band is significantly lower than that of the second candidate non-human voice frequency band, then after using the two candidate non-human voice frequency bands to respectively superimpose the same first sample audio with the same signal-to-noise ratio, the superimposed noise coefficient value of the former may be significantly lower than the corresponding value of the latter. Since the noise superposition coefficient of the latter may be significantly too high, when the value of its real high-noise interval is multiplied by the superposition noise coefficient and then superposed with the original audio data, the original human voice attribute in the corresponding interval of the original audio data may be annihilated due to excessive noise. It can thus be seen that the above selection operation of the target non-human voice frequency band can help ensure as much as possible that the audio attribute of the original human voice in the first sample audio will not be changed due to an excessive noise addition coefficient during subsequent noise addition.
[0020] In a specific implementation scenario, after obtaining the target non-human voice frequency band, the target non-human voice frequency band can be used to add signal-to-noise ratio noise (i.e., target signal-to-noise ratio) of a magnitude similar to that of the target service scenario to the first sample audio, so as to obtain the first target audio.
[0021] In one implementation scenario, after obtaining the first target audio for training the endpoint detection model, the model can be trained based on this first target audio. Specifically, based on alignment information, the sample frame type of the first target frame can be determined from several preset frame types. Based on the sample frame type and prediction probability of the first target frame, it can be assigned to any of several sample sets. Thus, based on whether the sample set corresponds to human voice or non-human voice, the positive or negative cross-entropy term in the binary cross-entropy loss function can be selected to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set, obtaining the sub-loss of the sample set. Based on whether the sample set corresponds to difficult or easy samples, a first weight or a second weight can be selected as the loss weight of the sample set. Then, based on the sub-loss and loss weight of each sample set, a weighted training loss can be obtained, and the network parameters of the endpoint detection model can be adjusted based on the training loss. The above method first combines alignment information and prediction probability to classify the first target frame into one of four sample sets: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. Then, based on whether the sample set corresponds to non-human voice or human voice, the positive or negative cross-entropy term in the binary cross-entropy loss function is selected for loss measurement to obtain the sub-loss of the sample set. Based on whether the sample set corresponds to difficult or easy samples, the first or second weight is selected as the loss weight of the sample set. Finally, the training loss is obtained by weighting the sub-loss and loss weight of each sample set. This method can force the endpoint detection model to improve detection accuracy for different situations of difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples by minimizing the training loss during the training process. This helps to improve the endpoint detection model's ability to detect difficult samples in complex environments such as complex high-noise backgrounds.
[0022] In a specific implementation scenario, as mentioned above, the alignment information may include the start and end frames of each word in the first audio text in the first target audio. According to the alignment information, if the first target frame in the first target audio is located between the start and end frames of a certain word in the first target audio, it can be determined that the sample frame type of this first target frame is human voice; otherwise, it can be determined that the sample frame type of this first target frame is non-human voice.
[0023] In a specific implementation scenario, taking the target frame type being speech as an example, for a first target frame whose sample frame type is speech, if the prediction probability that the first target frame belongs to the target frame type (i.e., speech) is not higher than a first probability threshold, the first target frame can be classified as a hard speech sample; if the prediction probability that the first target frame belongs to the target frame type (i.e., speech) is higher than the first probability threshold, the first target frame can be classified as an easy speech sample. For ease of description, the first probability threshold for distinguishing hard / easy speech samples may be denoted as thred1, and the prediction probability that the i-th first target frame belongs to the target frame type (i.e., speech) may be denoted as p(y i ), then when p(y i )>thred1, the first target frame can be classified as an easy speech sample, and when p(y i )≤thred1, the first target frame can be classified as a hard speech sample. That is, a hard speech sample is usually a first target frame that actually belongs to speech but has a low prediction probability of belonging to speech, such that it is difficult to correctly classify it as speech based on its prediction probability of belonging to speech, while an easy speech sample is usually a first target frame that actually belongs to speech and has a high prediction probability of belonging to speech, such that it can be easily correctly classified as speech based on its prediction probability of belonging to speech. In this way, first target frames with the sample frame type of speech can be as accurately as possible distinguished into hard speech samples or easy speech samples according to their prediction probabilities of belonging to speech.
[0024] In a specific implementation scenario, taking the target frame type being speech as an example, for a first target frame whose sample frame type is non-speech, if the prediction probability that the first target frame belongs to the target frame type is not lower than a second probability threshold, the first target frame can be classified as a hard non-speech sample; if the prediction probability that the first target frame belongs to the target frame type is lower than the second probability threshold, the first target frame can be classified as an easy non-speech sample. For ease of description, the second probability threshold for distinguishing hard / easy non-speech samples may be denoted as thred2, and the prediction probability that the i-th first target frame belongs to the target frame type (i.e., speech) may be denoted as p(y i ), then when p(y i )<thred2, the first target frame can be classified as an easy non-speech sample, and when p(y iWhen the probability of a non-human voice sample is greater than or equal to 2, the first target frame can be classified as a difficult non-human voice sample. In other words, difficult non-human voice samples typically exhibit a low predicted probability of the first target frame, which is actually non-human voice, belonging to human voice. Therefore, based on the predicted probability of it belonging to human voice, it can be easily and correctly classified as non-human voice. Conversely, easy non-human voice samples typically exhibit a high predicted probability of the first target frame, which is actually non-human voice, belonging to human voice. Therefore, based on the predicted probability of it belonging to human voice, it is difficult to correctly classify it as non-human voice. Thus, the first target frame of non-human voice type can be classified as either a difficult non-human voice sample or an easy non-human voice sample as accurately as possible based on its predicted probability of belonging to human voice.
[0025] It should be noted that in the examples above, where the target frame type is human voice, the first probability threshold and the second probability threshold can be negatively correlated. For instance, the sum of the first and second probability thresholds can be 1. For example, the second probability threshold can be obtained by subtracting the first probability threshold from 1, i.e., threshold2 = 1 - threshold1. Of course, the above example is merely one possible conversion example when the first and second probability thresholds are negatively correlated; other possible negative correlation scenarios will not be listed here. Furthermore, when the target frame type is non-human voice, the distinction between difficult human voice samples, easy human voice samples, difficult non-human voice samples, and easy non-human voice samples can be deduced similarly, and will not be listed here.
[0026] In a specific implementation scenario, to facilitate understanding of the subsequent selection of positive or negative cross-entropy terms based on the sample set corresponding to human or non-human voice, the binary cross-entropy loss function will be briefly explained below.
[0027]
[0028] In the above formula, BCELoss represents the binary classification cross-entropy loss function, y i log(p(y i )) represents the positive class cross-entropy term in the binary cross-entropy loss function, (1-y i )log(1-p(y i )) represents the negative class cross-entropy term in the binary cross-entropy loss function. Furthermore, y i The binary label (e.g., 0 or 1) can represent human voice or non-human voice in this embodiment of the disclosure (e.g., 1 represents human voice label, and 0 represents non-human voice label), p(y iThe value ) represents the predicted probability of belonging to the binary label, which in this embodiment can represent the predicted probability of belonging to the target frame type. N represents the number of samples, which in this embodiment can represent the total number of the first target frames. It should be noted that the binary cross-entropy loss function is typically used to evaluate the quality of a binary classification model's prediction results. That is, for a binary label of 1, if the predicted probability approaches 1, the loss value approaches 0; conversely, if the predicted probability approaches 0, the loss value increases sharply. Similarly, for a binary label of 0, if the predicted probability approaches 1, the loss value increases sharply; conversely, if the predicted probability approaches 0, the loss value approaches 0. From the above analysis, it can be seen that if the binary cross-entropy loss function is used to measure the loss, the contribution of each frame signal to the loss value is the same. Although the binary cross-entropy loss function can effectively guide the binary classification model to distinguish between human and non-human voice signals in normal scenarios, its guiding effect decreases sharply when facing complex scenarios such as high-noise backgrounds, because the contribution of each frame signal is the same. Therefore, by distinguishing between difficult human voice samples, easy human voice samples, difficult non-human voice samples, and easy non-human voice samples when measuring loss, the endpoint detection model can be forced to learn how to distinguish between human voices and non-human voices in complex scenarios.
[0029] In a specific implementation scenario, taking the target frame type as human voice as an example, when the sample set corresponds to human voice, that is, when the sample set consists of difficult or easy human voice samples, the positive class cross-entropy term (i.e., y in the aforementioned formula) can be selected. i log(p(y i The loss is measured by the sample frame type and prediction probability of the first target frame in the sample set to obtain the sub-loss of the sample set. For example, for any first target frame in either the difficult human voice sample set or the easy human voice sample set, the sub-loss for the difficult human voice sample set and the sub-loss for the easy human voice sample set can be obtained by substituting its sample frame type (i.e., human voice, which can be represented by the number "1") and prediction probability (i.e., the probability value predicted as belonging to human voice) into the above positive class cross-entropy term. Similarly, still taking the target frame type as human voice as an example, when the sample set corresponds to non-human voice, that is, when the sample set is either a difficult or easy human voice sample, a negative class cross-entropy term (i.e., (1-y) in the aforementioned formula) can be selected. i )log(1-p(y iThe loss is measured by the sample frame type and prediction probability of the first target frame in the sample set to obtain the sub-loss of the sample set. For example, for any first target frame in either the non-human voice difficult sample set or the non-human voice easy sample set, the sub-loss of the non-human voice difficult sample set and the sub-loss of the non-human voice easy sample set can be obtained by substituting its sample frame type (i.e., non-human voice, which can be represented by the number "0") and prediction probability (i.e., the probability value predicted as belonging to human voice) into the aforementioned negative class cross-entropy term. This method, by selecting the positive class cross-entropy term for loss measurement when the sample set corresponds to human voice and the negative class cross-entropy term for loss measurement when the sample set corresponds to non-human voice, allows for the selection of the appropriate cross-entropy term for loss measurement based on whether the sample set corresponds to human voice or non-human voice.
[0030] In a specific implementation scenario, after measuring the sub-loss of each sample set, the loss weight of each sample set can be determined. Specifically, taking the target frame type as human voice as an example, when the sample set corresponds to a hard sample, the first weight can be selected as the loss weight of the sample set, and when the sample set corresponds to an easy sample, the second weight can be selected as the loss weight of the sample set. It should be noted that the first weight and the second weight can be negatively correlated, and the first weight is positively correlated with the classification requirement for hard samples. That is, the higher the classification requirement for hard samples, the larger the first weight can be, and vice versa. For example, the sum of the first weight and the second weight can be 1. For ease of description, the first weight can be denoted as α, and the second weight can be denoted as 1-α. Then, the training loss of the endpoint detection model can be expressed as:
[0031] In the above formula, thred1 represents the first probability threshold, and thred2 represents the second probability threshold. For details, please refer to the aforementioned description. p(y i ) represents the predicted probability that the i-th first target frame belongs to the target frame type "human voice". The sample set is referred to as "difficult human voice samples". This indicates that the sample set is "non-human voice difficult sample". The sample set is called "Human Voice Easy Sample". Let α represent the sample set "Non-human voice easy samples", α represent the first weight (i.e., the loss weight of the sub-loss of the sample sets "human voice difficult samples" and "non-human voice difficult samples"), 1-α represent the second weight (i.e., the loss weight of the sub-loss of the sample sets "human voice easy samples" and "non-human voice easy samples"), and N represent the total number of samples, i.e., the total number of first target frames in each sample set. After obtaining the training loss, the network parameters of the endpoint detection model can be adjusted based on the training loss. This iterative training is repeated until the training loss no longer decreases or the maximum number of iterations is reached, resulting in an endpoint detection model suitable for complex scenes such as high-noise backgrounds.
[0032] In one implementation scenario, as another possible example, after the endpoint detection model is trained based on the first target audio, it can also be retrained based on the first target audio and the second target audio. For example, after performing speech detection on the unlabeled second sample audio, the endpoint detection model segments the second sample audio into sample human voice audio segments and sample non-human voice audio segments based on the detection results. These sample human voice audio segments and sample non-human voice audio segments are then recognized by a speech recognition model, and the recognition results determine whether they are used as the second target audio. Furthermore, the second sample audio can be collected extensively in the target business scenario (e.g., no less than 1000 hours). The specific meaning of the target business scenario can be found in the aforementioned descriptions and will not be repeated here. The above method, after training the endpoint detection model based on the first target audio, further trains the endpoint detection model on the first and second target audio. After detecting the unlabeled second sample audio, the endpoint detection model segments the second sample audio into human voice segments and non-human voice segments based on the detection results. These human voice and non-human voice segments are then recognized by a speech recognition model, and their inclusion as second target audio is determined based on the recognition results. Therefore, during the secondary training, the recognition results of the speech recognition model can be used to filter and supplement a large number of training samples, which helps to further improve the training effect of the endpoint detection model. Of course, the above implementation is only one possible example in practical applications and does not mean that the endpoint detection model must be trained on the first and second target audios after initial training; it is also possible to train the endpoint detection model only once based on the first target audio.
[0033] In a specific implementation scenario, as mentioned earlier, the endpoint detection model can be applied to the target business scenario. The first target audio can be obtained by superimposing the target non-human audio segment of the target business scenario onto the first sample audio with the target signal-to-noise ratio. The first sample audio can be labeled with alignment information, and the speech recognition model can be trained based on the first target audio. For example, the speech recognition model can be used to recognize the first target audio to obtain the first recognized text of the first target audio. Then, based on the difference between the first recognized text of the first target audio and the first audio text, the network parameters of the speech recognition model can be adjusted. In other words, the speech recognition model can reuse noisy labeled audio data for training to improve the recognition performance of the speech recognition model in complex scenarios such as high-noise backgrounds, thereby helping to filter out invalid training data.
[0034] In a specific implementation scenario, as mentioned earlier, after the endpoint detection model is trained based on the first target audio, it can first detect the unlabeled second sample audio in the target business scenario. Based on the detection results, the second sample audio can be segmented to obtain sample human voice audio segments and sample non-human voice audio segments. For example, by detecting the second sample audio, the endpoint detection model can obtain the probability value of each audio frame in the second sample audio belonging to the target frame type (e.g., human voice). If the probability value of belonging to the target frame type (e.g., human voice) is higher than the probability threshold, it can be determined that it belongs to the target frame type (e.g., human voice); otherwise, it does not belong to the target frame type (e.g., non-human voice). This is how the sample human voice audio segments and sample non-human voice audio segments are separated.
[0035] In a specific implementation scenario, for a sample human voice audio segment, the recognition results of the sample human voice audio segment can be statistically analyzed based on a speech recognition model to obtain the number of characters in the sample human voice audio segment per unit time. Furthermore, the confusion level of the sample human voice audio segment at the individual word level can be obtained by analyzing the recognition results based on the speech recognition model. For example, the speech recognition model recognizes the sample human voice audio segment, obtaining the recognized text (i.e., the recognition result). Based on this, the ratio between the total number of characters in the recognized text and the total duration of the sample human voice audio segment can be used as the number of characters in the sample human voice audio segment per unit time. For ease of description, the total number of characters in the recognized text can be denoted as rec_word_num, and the total duration of the sample human voice audio segment can be denoted as time_wav. Therefore, the number of characters in the sample human voice audio segment per unit time can be expressed as: rec_num_per_sec=rec_word_num / time_wav In the above formula, rec_num_per_sec represents the number of characters in the sample human voice audio segment per unit time. Furthermore, the recognized text of the sample human voice audio segment can be analyzed based on a language model to obtain the confusion level of the recognized text. Then, based on the ratio between the confusion level of the recognized text and the total number of characters in the recognized text, the confusion level of the sample human voice audio segment in a single word can be obtained. For ease of description, the confusion level of the recognized text can be denoted as PPL, and the confusion level of the sample human voice audio segment in a single word can be expressed as: PPL_avg=PPL / rec_word_num In the above formula, PPL_avg represents the confusion level of the sample human voice audio segment in a single word. After obtaining the number of characters in the sample human voice audio segment per unit time and the confusion level of the human voice audio segment in a single word, it is possible to determine whether to select the sample human voice audio segment as the second target audio based on these two factors. As a possible implementation example, it is possible to detect whether the number of characters in the sample human voice audio segment per unit time is within the target value range and whether the confusion level of the sample human voice audio segment in a single word is not higher than the confusion level threshold. If the above conditions are met, the sample human voice audio segment can be considered normal, and it can be retained as the second target audio. Conversely, if the conditions are not met, the sample human voice audio segment can be considered abnormal and it can be removed, i.e., it is not selected as the second target audio. In other words, for a sample human voice audio segment, it can be checked whether it meets any of the following criteria: the number of characters in the sample human voice audio segment per unit time is lower than the lower limit of the target numerical range rec_num_thred1; the number of characters in the sample human voice audio segment per unit time is higher than the upper limit of the target numerical range rec_num_thred2; or the confusion level of the sample human voice audio segment for a single word is higher than the confusion level threshold PPL_thred. If any of the above criteria are met, the sample human voice audio segment can be considered abnormal. Similarly, for sample non-human voice audio segments, the non-human voice audio segment can be identified based on a speech recognition model to obtain the identified text of the non-human voice audio segment (i.e., the recognition result). If the identified text is empty (or the total number of characters is less than the lower limit), the sample non-human voice audio segment can be considered normal and can be retained as the second target audio; otherwise, the sample non-human voice audio segment can be considered abnormal and can be discarded, i.e., not selected as the second target audio. The above method, by performing statistics and analysis based on the recognition results, obtains the number of characters in the sample human voice audio segment per unit time and the confusion level of individual words, and determines whether to select the sample human voice audio segment as the second target audio. This method can improve the sample quality of the second target audio as much as possible when increasing the number of samples for secondary training.
[0036] In a specific implementation scenario, after obtaining the second target audio, the endpoint detection model can be retrained based on both the first and second target audio. It should be noted that the specific method for loss measurement during retraining can be found in the detailed process of retraining the endpoint detection model based on the first target audio. For example, the aforementioned BCELoss_new method can be used to distinguish between difficult human voice samples, easy human voice samples, and difficult and easy non-human voice samples for loss measurement, which will not be elaborated upon here. This process is repeated iteratively to train the endpoint detection model during retraining until the training loss no longer decreases or the maximum number of iterations is reached, thus obtaining an endpoint detection model suitable for complex scenarios such as high-noise backgrounds.
[0037] Step S12: Based on the probability value of the audio frame belonging to the target frame type, determine the final frame type of the audio frame from several preset frame types.
[0038] In one implementation scenario, as a possible example, the probability value of an audio frame belonging to the target frame type can be compared with a probability threshold used to distinguish between human voices and non-human voices. If the probability value of the audio frame belonging to the target frame type is not lower than (or higher than) the probability threshold, then the final frame type of the audio frame can be determined to be the target frame type; otherwise, the final frame type of the audio frame can be determined to be non-human voices. Taking the target frame type as human voice as an example, if the probability value of the audio frame belonging to human voice is not lower than (or higher than) the probability threshold, then the final frame type of the audio frame can be determined to be human voices; otherwise, the final frame type of the audio frame can be determined to be non-human voices.
[0039] In another implementation scenario, unlike the aforementioned implementation, as another possible example, a comparison can be made first between the second energy threshold and the temporal energy of the audio frame to determine whether to directly determine non-human voice as the final frame type among several preset frame types. Based on this, if non-human voice is not directly determined as the final frame type among several preset frame types, the final frame type can be determined among several preset frame types based on the probability value of the audio frame belonging to the target frame type. Of course, if non-human voice is directly determined as the final frame type among several preset frame types, the probability value can be disregarded. The above method, by further considering temporal energy when determining the final frame type, helps improve the accuracy of distinguishing between human voice and non-human voice. While temporal energy can be considered in conjunction with the determination of the final frame type in practical applications, this does not mean that temporal energy must be referenced in practice. As in the aforementioned implementation example, only the probability value can be referenced to determine the final frame type; this is not a limitation here.
[0040] In a specific implementation scenario, when the temporal energy of an audio frame is lower than (or not higher than) a second energy threshold, non-human voice can be directly identified as the final frame type among several preset frame types. For ease of description, the second energy threshold can be denoted as energy_thred. When the temporal energy of an audio frame is lower than (or not higher than) the second energy threshold energy_thred, the final frame type of the audio frame can be directly determined to be non-human voice. Otherwise, the final frame type of the audio frame can be further determined based on the probability value of the audio frame belonging to the target frame type.
[0041] In a specific implementation scenario, when it is necessary to determine the final frame type of an audio frame based on the probability value of the audio belonging to the target frame type, please refer to the aforementioned implementation example, which will not be repeated here.
[0042] In one implementation scenario, after obtaining the final frame type of each audio frame in the audio to be detected, the audio to be detected can be further divided into human voice segments and non-human voice segments based on the final frame type of each audio frame, so as to facilitate the next stage of downstream tasks (such as speech recognition). This will not be elaborated on here.
[0043] The above scheme uses an endpoint detection model to detect the audio to be detected, obtaining the probability value of each audio frame in the audio belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The endpoint detection model detects each first target frame in the first target audio and obtains the predicted probability that it belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. These sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and these sub-losses are based on the alignment information and the predicted probability. The probability is measured, and based on this, the final frame type of the audio frame is determined from several preset frame types. Since the endpoint detection model combines alignment information and predicted probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish sub-losses for these categories in the training loss of the endpoint detection model. Furthermore, by minimizing the training loss during training, the endpoint detection model can improve detection accuracy for all categories, including difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps improve the endpoint detection model's ability to detect difficult samples in complex environments such as those with complex, high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments.
[0044] Please see Figure 2 , Figure 2This is a schematic diagram of the framework of an embodiment of the speech endpoint detection device of this application. The speech endpoint detection device 20 includes: an endpoint detection module 21 and a type determination module 22. The endpoint detection module 21 is used to detect the audio to be detected based on the endpoint detection model and obtain the probability value of the audio frame in the audio to be detected belonging to the target frame type; wherein, the target frame type is at least one of several preset frame types, the several preset frame types include human voice and non-human voice, the endpoint detection model is trained based on a first target audio, the first target audio is labeled with frame-level alignment information between it and the corresponding first audio text, the first target audio is detected by the endpoint detection model to obtain the predicted probability of each first target frame in the first target audio belonging to the target frame type, the first target frame is classified into any of several sample sets based on the alignment information and the predicted probability, the several sample sets include human voice difficult samples, non-human voice difficult samples, human voice easy samples and non-human voice easy samples, the training loss of the endpoint detection model includes the sub-loss of each sample set, and the sub-loss is obtained by measuring based on the alignment information and the predicted probability; the type determination module 22 is used to determine the final frame type of the audio frame in the several preset frame types based on the probability value of the audio frame belonging to the target frame type.
[0045] In the above scheme, the speech endpoint detection device 20 detects the audio to be detected based on an endpoint detection model, obtaining the probability value of audio frames in the audio to be detected belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The first target audio is detected by the endpoint detection model to obtain the predicted probability that each first target frame in the first target audio belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. The sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and the sub-losses are based on the training loss of the target audio. The alignment information and prediction probability are measured to obtain the final frame type of the audio frame. Based on this, and then based on the probability value of the audio frame belonging to the target frame type, the final frame type of the audio frame is determined from several preset frame types. Since the endpoint detection model combines alignment information and prediction probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish the sub-losses of difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples in the training loss of the endpoint detection model. Therefore, by minimizing the training loss during training, the endpoint detection model can be forced to improve the detection accuracy for different cases of difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps to improve the endpoint detection model's ability to detect difficult samples in complex environments such as complex high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments.
[0046] In some disclosed embodiments, the speech endpoint detection device 20 includes a sample determination module for determining the sample frame type of a first target frame from several preset frame types based on alignment information; the speech endpoint detection device 20 includes a sample assignment module for assigning the first target frame to any of several sample sets based on the sample frame type and prediction probability of the first target frame; the speech endpoint detection device 20 includes an item selection module for selecting the positive or negative cross-entropy term in the binary cross-entropy loss function to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set based on whether the sample set corresponds to human voice or non-human voice, thereby obtaining a sub-loss of the sample set; the speech endpoint detection device 20 includes a weight determination module for selecting a first weight or a second weight as the loss weight of the sample set based on whether the sample set corresponds to difficult or easy samples; and the speech endpoint detection device 20 includes a loss weighting module for weighting the training loss based on the sub-loss and loss weight of each sample set.
[0047] In some disclosed embodiments, the sample classification module includes a first classification submodule, configured to, when the target frame type is human voice, classify the first target frame as a difficult human voice sample if the predicted probability of the first target frame belonging to the target frame type is not higher than a first probability threshold, and classify the first target frame as an easy human voice sample if the predicted probability of the first target frame belonging to the target frame type is higher than the first probability threshold. The sample classification module also includes a second classification submodule, configured to, when the target frame type is human voice, classify the first target frame as a difficult non-human voice sample if the predicted probability of the first target frame belonging to the target frame type is not lower than a second probability threshold, and classify the first target frame as an easy non-human voice sample if the predicted probability of the first target frame belonging to the target frame type is lower than the second probability threshold.
[0048] In some disclosed embodiments, the item selection module includes a first selection submodule, used to select a positive class cross-entropy term to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set when the target frame type is human voice and when the sample set corresponds to human voice, thereby obtaining a sub-loss of the sample set. The item selection module also includes a second selection submodule, used to select a negative class cross-entropy term to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set when the target frame type is human voice and when the sample set corresponds to non-human voice, thereby obtaining a sub-loss of the sample set.
[0049] In some disclosed embodiments, the weight determination module is specifically used to select a first weight as the loss weight of the sample set when the target frame type is human voice and the sample set corresponds to a difficult sample, and to select a second weight as the loss weight of the sample set when the sample set corresponds to an easy sample; wherein the first weight and the second weight are negatively correlated, and the first weight is positively correlated with the classification requirements for difficult samples.
[0050] In some disclosed embodiments, the endpoint detection model is applicable to the target business scenario. The first target audio is obtained by superimposing the target non-human voice audio segment of the target business scenario on the first sample audio with the target signal-to-noise ratio. The first sample audio is labeled with alignment information. The voice endpoint detection device 20 includes an audio segmentation module for segmenting the target non-human voice audio based on the target business scenario to obtain several candidate non-human voice audio segments. The voice endpoint detection device 20 includes a frame count module for counting the temporal energy of each non-human voice frame in the candidate non-human voice audio segments based on a first energy threshold to obtain the target frame number. The voice endpoint detection device 20 includes an occurrence calculation module for obtaining the ratio between the target frame number and the audio segment duration of the candidate non-human voice audio segments as the average occurrence frequency. The voice endpoint detection device 20 includes a sample screening module for comparing the occurrence frequency threshold with the average occurrence frequency to determine whether to select the candidate non-human voice audio segment as the target non-human voice audio segment.
[0051] In some publicly disclosed embodiments, the endpoint detection model is trained based on the first target audio and then trained again based on the first target audio and the second target audio. After the endpoint detection model detects the unlabeled second sample audio, it segments the second sample audio according to the detection results to obtain sample human voice audio segments and sample non-human voice audio segments. After the sample human voice audio segments and sample non-human voice audio segments are recognized by the speech recognition model, it is determined whether they are used as the second target audio according to the recognition results.
[0052] In some disclosed embodiments, the voice endpoint detection device 20 includes a statistical analysis module for performing statistics based on the recognition results to obtain the number of characters in the sample human voice audio segment per unit time, and performing analysis based on the recognition results to obtain the confusion degree of the sample human voice audio segment in a single word; the voice endpoint detection device 20 includes a selection determination module for determining whether to select the sample human voice audio segment as the second target audio based on the number of characters and the confusion degree.
[0053] In some publicly disclosed embodiments, the endpoint detection model is applicable to the target business scenario. The first target audio is obtained by superimposing the target non-human voice audio segment of the target business scenario on the first sample audio with the target signal-to-noise ratio. The first sample audio is labeled with alignment information, and the speech recognition model is trained based on the first target audio.
[0054] In some disclosed embodiments, the voice endpoint detection device 20 includes an energy comparison module, which is used to compare the temporal energy of the audio frame with a second energy threshold to determine whether to directly determine non-human voice as the final frame type of the audio frame among a number of preset frame types; the type determination module 22 is specifically used to determine the final frame type of the audio frame among a number of preset frame types in response to the fact that non-human voice is not directly determined as the final frame type of the audio frame among a number of preset frame types, based on the probability value of the audio frame belonging to the target frame type.
[0055] Please see Figure 3 , Figure 3 This is a schematic diagram of a framework of an embodiment of the electronic device of this application. The electronic device 30 includes at least a memory 31 and a processor 32 coupled to each other. The memory 31 stores at least program instructions, and the processor 32 is used to execute the program instructions to implement the steps in any of the above-described embodiments of the voice endpoint detection method. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here. The electronic device 30 may include, but is not limited to, smartphones, tablets, learning machines, translators, notebooks, in-vehicle systems, smart screens, servers, etc. The specific type of the electronic device 30 is not limited here.
[0056] Specifically, processor 32 controls itself and memory 31 to implement the steps in any of the above-described voice endpoint detection method embodiments. Processor 32 can also be referred to as a CPU (Central Processing Unit). Processor 32 may be an integrated circuit chip with signal processing capabilities. Processor 32 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 32 can be implemented using integrated circuit chips.
[0057] In the above scheme, the electronic device 30 detects the audio to be detected based on an endpoint detection model, obtaining the probability value of audio frames in the audio to be detected belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The endpoint detection model detects the first target audio and obtains the predicted probability that each first target frame in the first target audio belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. The sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and the sub-losses are based on the alignment information. The information and prediction probability are measured, and based on this, the final frame type of the audio frame is determined from several preset frame types based on the probability value of the audio frame belonging to the target frame type. Since the endpoint detection model combines alignment information and prediction probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish sub-losses for difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples in the training loss of the endpoint detection model. Therefore, by minimizing the training loss during training, the endpoint detection model can improve detection accuracy for different cases of difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps improve the endpoint detection model's ability to detect difficult samples in complex environments such as complex high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments.
[0058] Please see Figure 4 , Figure 4 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 40 stores program instructions 41 that can be executed by a processor. The program instructions 41 are used to implement the steps in any of the above-described embodiments of the voice endpoint detection method.
[0059] In the above scheme, the computer-readable storage medium 40 detects the audio to be detected based on an endpoint detection model, obtaining the probability value of audio frames in the audio to be detected belonging to a target frame type. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The endpoint detection model detects the first target audio and obtains the predicted probability that each first target frame in the first target audio belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. The sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes sub-losses for each sample set, and the sub-losses are based on... Alignment information and prediction probability are measured to obtain the final frame type of the audio frame. Based on this, and then based on the probability value of the audio frame belonging to the target frame type, the final frame type of the audio frame is determined from several preset frame types. Since the endpoint detection model combines alignment information and prediction probability during training to classify the first target frame in the first target audio into one of the following categories: difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples, it can distinguish sub-losses for difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples in the training loss of the endpoint detection model. Therefore, by minimizing the training loss during training, the endpoint detection model can improve detection accuracy for different scenarios of difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. This helps improve the endpoint detection model's ability to detect difficult samples in complex environments such as complex high-noise backgrounds. Therefore, it can improve the accuracy of speech endpoint detection in complex environments.
[0060] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0061] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0062] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0063] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0064] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0065] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0066] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A method for detecting speech endpoints, characterized in that, include: The endpoint detection model is used to detect the audio to be detected, and the probability value of the audio frame in the audio to be detected belonging to the target frame type is obtained. The target frame type is at least one of several preset frame types, including human voice and non-human voice. The endpoint detection model is trained based on a first target audio, which is labeled with frame-level alignment information between itself and the corresponding first audio text. The first target audio is detected by the endpoint detection model to obtain the predicted probability that each first target frame in the first target audio belongs to the target frame type. The first target frame is assigned to any of several sample sets based on the alignment information and the predicted probability. The several sample sets include difficult human voice samples, difficult non-human voice samples, easy human voice samples, and easy non-human voice samples. The training loss of the endpoint detection model includes the sub-loss of each of the sample sets, and the sub-loss is measured based on the alignment information and the predicted probability. Based on the probability value of the audio frame belonging to the target frame type, the final frame type of the audio frame is determined from the plurality of preset frame types; The training loss measurement step includes: determining the sample frame type of the first target frame from among the several preset frame types based on the alignment information; classifying the first target frame into any of the several sample sets based on the sample frame type and prediction probability of the first target frame; selecting either the positive or negative cross-entropy term in the binary cross-entropy loss function to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set, based on whether the sample set corresponds to human voice or non-human voice, to obtain the sub-loss of the sample set; and selecting either a first weight or a second weight as the loss weight of the sample set based on whether the sample set corresponds to difficult or easy samples; and weighting the sub-loss and loss weight of each sample set to obtain the training loss. The endpoint detection model is applicable to the target business scenario. The first target audio is obtained by superimposing the target non-human voice audio segment of the target business scenario on the first sample audio with the target signal-to-noise ratio. The first sample audio is labeled with the alignment information. The steps for obtaining the target non-human voice audio segment include: segmenting the target non-human voice audio of the target business scenario to obtain several candidate non-human voice audio segments; statistically analyzing the temporal energy of each non-human voice frame in the candidate non-human voice audio segments based on a first energy threshold to obtain the target frame number; obtaining the ratio between the target frame number and the audio segment duration of the candidate non-human voice audio segments as the average occurrence frequency; and comparing the occurrence frequency threshold with the average occurrence frequency to determine whether to select the candidate non-human voice audio segment as the target non-human voice audio segment.
2. The method according to claim 1, characterized in that, When the target frame type is human voice, classifying the first target frame into any of the plurality of sample sets based on the sample frame type and predicted probability of the first target frame includes: For the first target frame of the sample frame type being human voice: if the predicted probability of the first target frame belonging to the target frame type is not higher than the first probability threshold, the first target frame is classified as the difficult human voice sample; if the predicted probability of the first target frame belonging to the target frame type is higher than the first probability threshold, the first target frame is classified as the easy human voice sample. For the first target frame whose sample frame type is non-human voice: if the predicted probability of the first target frame belonging to the target frame type is not lower than the second probability threshold, the first target frame is classified into the non-human voice difficult sample; if the predicted probability of the first target frame belonging to the target frame type is lower than the second probability threshold, the first target frame is classified into the non-human voice easy sample.
3. The method according to claim 1, characterized in that, When the target frame type is human voice, the step of selecting either the positive or negative cross-entropy term in the binary cross-entropy loss function based on whether the sample set corresponds to human voice or non-human voice to perform loss measurement on the sample frame type and prediction probability of the first target frame in the sample set, based on whether the sample set corresponds to human voice or non-human voice, to obtain the sub-loss of the sample set, includes: When the sample set corresponds to human voice, the positive class cross-entropy term is selected to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set, thus obtaining the sub-loss of the sample set; and, When the sample set corresponds to non-human voice, the negative class cross-entropy term is selected to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set, so as to obtain the sub-loss of the sample set.
4. The method according to claim 1, characterized in that, When the target frame type is human voice, the step of selecting a first weight or a second weight as the loss weight for the sample set based on whether the sample set corresponds to a difficult or easy sample includes: When the sample set corresponds to a difficult sample, the first weight is selected as the loss weight of the sample set; when the sample set corresponds to an easy sample, the second weight is selected as the loss weight of the sample set. The first weight and the second weight are negatively correlated, and the first weight is positively correlated with the classification requirements for difficult samples.
5. The method according to claim 1, characterized in that, The endpoint detection model is trained based on the first target audio and then trained again based on the first target audio and the second target audio. After detecting the unlabeled second sample audio, the endpoint detection model segments the second sample audio into sample human voice audio segments and sample non-human voice audio segments according to the detection results. After the sample human voice audio segments and the sample non-human voice audio segments are recognized by the speech recognition model, it is determined whether they are used as the second target audio according to the recognition results.
6. The method according to claim 5, characterized in that, The step of determining whether to select the sample human voice audio segment as the second target audio includes: Based on the recognition results, statistics are performed to obtain the number of characters in the sample human voice audio segment per unit time, and based on the recognition results, analysis is performed to obtain the confusion degree of the sample human voice audio segment in a single word. Based on the number of characters and the level of confusion, determine whether to select the sample human voice audio segment as the second target audio.
7. The method according to claim 5, characterized in that, The endpoint detection model is applicable to the target business scenario. The first target audio is obtained by superimposing the target non-human voice audio segment of the target business scenario on the first sample audio with the target signal-to-noise ratio. The first sample audio is labeled with the alignment information. The speech recognition model is trained based on the first target audio.
8. The method according to claim 1, characterized in that, Before determining the final frame type of the audio frame from among several preset frame types based on the probability value of the audio frame belonging to the target frame type, the method further includes: Based on the comparison between the second energy threshold and the temporal energy of the audio frame, it is determined whether to directly determine non-human voice among the several preset frame types as the final frame type of the audio frame; The step of determining the final frame type of the audio frame from among several preset frame types based on the probability value of the audio frame belonging to the target frame type includes: Since it is not directly determined that non-human voice is the final frame type of the audio frame among the several preset frame types, the final frame type of the audio frame is determined among the several preset frame types based on the probability value of the audio frame belonging to the target frame type.
9. A voice endpoint detection device, characterized in that, include: An endpoint detection module is used to detect audio to be detected based on an endpoint detection model, and obtain the probability value of audio frames in the audio to be detected belonging to a target frame type; wherein, the target frame type is at least one of several preset frame types, the several preset frame types include human voice and non-human voice, the endpoint detection model is trained based on a first target audio, the first target audio is labeled with frame-level alignment information between itself and the corresponding first audio text, the first target audio is detected by the endpoint detection model to obtain the predicted probability that each first target frame in the first target audio belongs to the target frame type, the first target frame is classified into any one of several sample sets based on the alignment information and the predicted probability, the several sample sets include human voice difficult samples, non-human voice difficult samples, human voice easy samples and non-human voice easy samples, the training loss of the endpoint detection model includes the sub-loss of each of the sample sets, and the sub-loss is obtained by measuring based on the alignment information and the predicted probability; A type determination module is used to determine the final frame type of the audio frame from among several preset frame types based on the probability value of the audio frame belonging to the target frame type; The training loss is obtained through the following operational metrics: based on the alignment information, the sample frame type of the first target frame is determined from among the several preset frame types; based on the sample frame type and prediction probability of the first target frame, the first target frame is assigned to any of the several sample sets; based on whether the sample set corresponds to human voice or non-human voice, the positive or negative cross-entropy term in the binary cross-entropy loss function is selected to measure the loss of the sample frame type and prediction probability of the first target frame in the sample set, thereby obtaining the sub-loss of the sample set; and based on whether the sample set corresponds to difficult or easy samples, a first weight or a second weight is selected as the loss weight of the sample set; the training loss is obtained by weighting the sub-losses and loss weights of each sample set. The endpoint detection model is applicable to the target business scenario. The first target audio is obtained by superimposing the target non-human voice audio segment of the target business scenario on the first sample audio with the target signal-to-noise ratio. The first sample audio is labeled with the alignment information. The target non-human voice audio segment is obtained through the following operations: the target non-human voice audio of the target business scenario is segmented to obtain several candidate non-human voice audio segments; the temporal energy of each non-human voice frame in the candidate non-human voice audio segment is statistically analyzed based on a first energy threshold to obtain the target frame number; the ratio between the target frame number and the audio segment duration of the candidate non-human voice audio segment is obtained as the average occurrence frequency; the occurrence frequency threshold is compared with the average occurrence frequency to determine whether to select the candidate non-human voice audio segment as the target non-human voice audio segment.
10. An electronic device, characterized in that, It includes at least a memory and a processor coupled to each other, wherein the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the voice endpoint detection method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the voice endpoint detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Sound quality detection model training method, sound quality detection method, electronic equipment and medium
CN114694678A
Voice endpoint detection method and system based on text perception
CN116092535A