Phoneme positioning method for target text and related device
By detecting the target single-state phoneme sequence in a speech segment and combining single-state and multi-state phoneme localization methods, the problem of insufficient localization accuracy and efficiency in the existing technology is solved, and efficient and accurate phoneme localization is achieved in complex speech environments.
Patent Information
- Application Number
- CN202510848816.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing phoneme localization methods for target texts either lack sufficient localization accuracy or insufficient localization efficiency.
The method first detects whether there is a target single-state phoneme sequence corresponding to the target text in the collected speech segments. If it exists, the phoneme state-level position is accurately located through multi-state phoneme localization. This method combines the advantages of single-state and multi-state phoneme localization methods, taking into account both localization efficiency and accuracy.
It achieves efficient detection of target text in complex speech environments and accurate location of phoneme state-level positions, improving robustness, reducing localization errors, and balancing localization efficiency and accuracy.
Smart Images

Figure CN120375803B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, in particular to a target text phoneme positioning method, a target text phoneme positioning device, an electronic device and a computer readable storage medium. BACKGROUND
[0002] The target text phoneme positioning technology aims to position the phoneme position of the target text in the speech, which is widely used in voice devices such as smart home devices and intelligent voice assistants. For example, in the application scenario of determining the position of the target text in the speech, the first phoneme and the last phoneme of the target text in the speech are determined by the phoneme positioning technology, so as to determine the position of the target text in the speech. For another example, in the speech segmentation application scenario, the position of the last phoneme of the target text in the speech is determined by the phoneme positioning technology, and the position of the last phoneme in the speech is taken as a segmentation point to segment the speech.
[0003] However, the existing target text phoneme positioning method either lacks positioning accuracy or lacks positioning efficiency. SUMMARY
[0004] The present application provides a target text phoneme positioning method, a target text phoneme positioning device, an electronic device and a computer readable storage medium, which can solve the problem that the existing target text phoneme positioning method either lacks positioning accuracy or lacks positioning efficiency.
[0005] The present application provides a target text phoneme positioning method, comprising: detecting whether a target single-state phoneme sequence corresponding to a target text exists in a collected speech segment; in response to the existence of the target single-state phoneme sequence, determining a preset state of a preset phoneme in a target multi-state phoneme sequence corresponding to the target text, and determining a preset speech frame corresponding to the preset state of the preset phoneme in the collected speech segment.
[0006] The present application provides a target text phoneme positioning device, comprising: a detection module and a determination module. Wherein the detection module is used to detect whether a target single-state phoneme sequence corresponding to a target text exists in a collected speech segment; the determination module is used to determine a preset state of a preset phoneme in a target multi-state phoneme sequence corresponding to the target text in response to the existence of the target single-state phoneme sequence, and to determine a preset speech frame corresponding to the preset state of the preset phoneme in the collected speech segment.
[0007] The present application provides an electronic device, comprising a memory and a processor, the processor is used to execute program instructions stored in the memory to realize the above method.
[0008] The application provides a computer readable storage medium, which stores program instructions, and the program instructions are executed by a processor to implement the method.
[0009] In the above scheme, the application first detects whether the target text corresponding to the target single-state phoneme sequence exists in the collected speech segment. In the case where the target single-state phoneme sequence exists, it means that the target text exists, and the preset state of the preset phoneme of the target text is repositioned in the corresponding preset speech frame in the collected speech segment. That is, the single-state phoneme positioning method is used to efficiently detect whether the target text exists, and in the case where the target text exists, the multi-state phoneme positioning method is used to accurately position the phoneme state level position, so as to balance the positioning efficiency and the positioning accuracy by comprehensively using the advantages of the single-state phoneme positioning method and the multi-state phoneme positioning method. Moreover, the single-state phoneme positioning method can efficiently detect whether the target text exists in a complex speech environment, and is complementary to the accurate positioning of the multi-state phoneme positioning method, thereby improving the robustness.
[0010] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, rather than limiting the application. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the application and, together with the specification, serve to explain the technical solutions of the application.
[0012] Figure 1 is a flowchart of an embodiment of the phoneme positioning method of the target text provided by the application;
[0013] Figure 2 is a flowchart of another embodiment of the phoneme positioning method of the target text provided by the application;
[0014] Figure 3 is a flowchart of still another embodiment of the phoneme positioning method of the target text provided by the application;
[0015] Figure 4 is a flowchart of an embodiment of the phoneme positioning model training method provided by the application;
[0016] Figure 5 is a flowchart of another embodiment of the phoneme positioning model training method provided by the application;
[0017] Figure 6 is a flowchart of another embodiment of the phoneme positioning model training method provided by the application;
[0018] Figure 7is a flowchart of an embodiment of a method for locating phonemes of target text provided by the present application;
[0019] Figure 8 is a structural diagram of a phoneme locating model of the present application;
[0020] Figure 9 is a diagram of a specific example of one-stage training of a phoneme locating model of the present application;
[0021] Figure 10 is a diagram of a specific example of two-stage training of a phoneme locating model of the present application;
[0022] Figure 11 is a diagram of a specific example of an inference stage of a phoneme locating model;
[0023] Figure 12 is a comparison diagram of two modeling methods of single-state phonemes and multi-state phonemes of the present application;
[0024] Figure 13 is a diagram of separating a wake-up instruction from a voice instruction of the present application;
[0025] Figure 14 is a structural diagram of an embodiment of a phoneme locating device for target text of the present application;
[0026] Figure 15 is a structural diagram of an embodiment of an electronic device of the present application;
[0027] Figure 16 is a structural diagram of an embodiment of a computer readable storage medium of the present application. DETAILED DESCRIPTION
[0028] The scheme of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0029] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to thoroughly understand the present application, but are not intended to limit the present application.
[0030] The term "and / or" in this paper is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" in this paper means two or more than two. In addition, the term "at least one" in this paper means any one of multiple or any combination of at least two of multiple, for example, including at least one of A, B and C, which can mean including any one or more elements selected from the set consisting of A, B and C.
[0031] The phoneme positioning in the related art is divided into single-state phoneme positioning and multi-state phoneme positioning. The single-state phoneme positioning refers to modeling the phonemes in the speech as single states, or modeling the phonemes as a whole without dividing the states, and determining the positions of the phonemes of the target text in the speech. The multi-state phoneme positioning refers to modeling the phonemes in the speech as multiple states, and determining the positions of the phoneme states of the target text in the speech.
[0032] The present inventors have found through long-term research that:
[0033] In the single-state phoneme positioning mode, the phonemes are modeled as a whole, and the states of the phonemes do not need to be divided, so the positioning complexity is low, and the positioning efficiency is high. Moreover, since the positioning unit is the phoneme, better performance can be exhibited in a complex speech environment. For example, since the phonemes are not easily covered by noise, better robustness can be exhibited in a low signal-to-noise ratio speech. The speech phenomena such as fast speech speed and swallowing are also highly acceptable. However, the phoneme-level position is positioned, and the phoneme-level position accuracy is low.
[0034] In the multi-state phoneme positioning mode, the phonemes are modeled as multiple states (for example, 3 states, respectively, the start state, the stable state, and the end state), so the phoneme state-level position can be positioned, and the phoneme state-level position accuracy is high. However, the positioning complexity is high, and the positioning efficiency is low.
[0035] In addition, the single-state phoneme positioning is implemented based on a single-state phoneme positioning model, and the multi-state phoneme positioning is implemented based on a multi-state phoneme positioning model. The single-state phoneme positioning model has a simple structure. The multi-state phoneme positioning model has a complex structure, and a large amount of calculation resources and time resources are required for training before the multi-state phoneme positioning model is applied to phoneme state positioning. Moreover, the multi-state phoneme positioning model depends on high-quality and high-quantity training speech, and is difficult to be applied to a resource-limited scene.
[0036] To solve at least part of the above technical problems, the present application provides a phoneme positioning method of a target text. The embodiments of the phoneme positioning method of the target text are introduced as follows.
[0037] Figure 1 is a flowchart of an embodiment of the phoneme positioning method of the target text provided by the present application. As shown in Figure 1 In the embodiment, the phoneme positioning method of the target text can include the following steps:
[0038] S110: detecting whether a target single-state phoneme sequence corresponding to the target text exists in the collected speech segment.
[0039] The collected speech segment can be a real-time speech stream or a complete speech that is not real-time. In the case of a real-time speech stream, the collected speech segment includes a current speech segment and a historical speech segment located before the current speech segment. The current speech segment and the historical speech segment belong to the same source speech, which can be a speech instruction, a narrative speech, etc.
[0040] The target text can be a command word, a proper noun, a number and a date, a professional term, a phrase or a sentence, a dialect word, a sensitive word, a custom-defined text, etc. The command word can be a wake-up word or an action word, etc. The wake-up word is, for example, "Hello, Harry", and the action word is, for example, "Turn on the TV". The target text can be a text in any language such as Chinese or English.
[0041] The target text includes a plurality of phonemes, and the phonemes of the target text can be single phonemes or tri-phonemes, etc. The single phoneme refers to the current phoneme itself, which is regarded as an individual without the context information (the previous adjacent phoneme and the next adjacent phoneme) of the current phoneme. The tri-phoneme refers to the previous adjacent phoneme-current phoneme-next adjacent phoneme, which includes the context information of the current phoneme. For example, the single phoneme composition of the target text "Hello Harry" is n, i3, h, ao3, h, a1, l, i4, and the tri-phoneme composition is sil-n-i3, n-i3-h, i3-h-ao3, h-ao3-h, ao3-h-a1, h-a1-l, a1-l-i4, l-i4-sil, where the numbers represent tones, 01234 respectively represent light tone, yin ping, yang ping, shang tone, and qu tone.
[0042] In the target single-state phoneme sequence, each phoneme of the target text is represented as a state, or each phoneme is regarded as a whole without state division. For example, the target single-state phoneme sequence is {phoneme 1, phoneme 2, phoneme 3, phoneme 4, phoneme 5}.
[0043] Since it is detected whether the target single-state phoneme sequence exists, in S110, the collected speech segment is modeled by a single-state phoneme to determine whether the target single-state phoneme sequence corresponding to the target text exists, so that the detection process is efficient.
[0044] It can be understood that, in the case that the target single-state phoneme sequence exists in the collected speech segment, it means that the target text exists in the collected speech segment.
[0045] S120: In response to the existence of the target single-state phoneme sequence, determining a preset state of a preset phoneme in a target multi-state phoneme sequence corresponding to the target text, and determining a preset speech frame corresponding to the preset state of the preset phoneme in the collected speech segment.
[0046] In the target multi-state phoneme sequence, each phoneme of the target text is divided into multiple states, for example, each phoneme is represented as three states, which are the start state, the stable state and the end state. The target multi-state phoneme sequence is {start state of phoneme 1, stable state of phoneme 1, end state of phoneme 1, start state of phoneme 2, stable state of phoneme 2, end state of phoneme 2, …}.
[0047] The preset phoneme can be any phoneme in the target multi-state phoneme sequence. In some embodiments, the preset phoneme is the first phoneme. In some embodiments, the preset phoneme is the last phoneme. For example, the preset phoneme is the first phoneme sil-n-i3 and the last phoneme l-i4-sil of the target text “Hello Harry”. For another example, the preset phoneme is the last phoneme l-i4-sil of the target text “Hello Harry”.
[0048] The preset state can be any state of the preset phoneme. In some embodiments, the preset state is the start state. In some embodiments, the preset state is the stable state. In some embodiments, the preset state is the end state.
[0049] One phoneme state corresponds to one phoneme frame in the collected speech segment. One phoneme state corresponds to a plurality of phoneme frames in the collected speech segment.
[0050] The preset phoneme frame can be any phoneme frame corresponding to the preset state of the preset phoneme. In some embodiments, the preset phoneme frame is all phoneme frames corresponding to the preset state of the preset phoneme. In some embodiments, the preset phoneme frame is part of the phoneme frames corresponding to the preset state of the preset phoneme. In some embodiments, the preset phoneme frame is the first phoneme frame or the last phoneme frame corresponding to the preset state of the preset phoneme.
[0051] In a specific example, the preset phoneme frame corresponding to the preset state of the preset phoneme is the first phoneme frame of the end state of the last phoneme in the target multi-state phoneme sequence. That is, the transition point from the previous state to the last state of the last phoneme.
[0052] The preset phoneme frame corresponding to the preset state of the preset phoneme in the collected speech segment is the phoneme state level position. Therefore, in S120, the collected speech segment is modeled as a multi-state phoneme, and the phoneme state level position is determined, so that accurate positioning can be achieved.
[0053] Through implementation of the embodiment, the application first detects whether the target text corresponding to the target single-state phoneme sequence exists in the collected speech segment. In the case where the target single-state phoneme sequence exists, it means that the target text exists, and the preset state of the preset phoneme of the target text is repositioned in the corresponding preset speech frame in the collected speech segment. That is, the single-state phoneme positioning method is used to efficiently detect whether the target text exists, and in the case where the target text exists, the multi-state phoneme positioning method is used to accurately position the phoneme state level position, so as to comprehensively utilize the advantages of the single-state phoneme positioning method and the multi-state phoneme positioning method, balance the positioning efficiency and the positioning accuracy, and improve the robustness.
[0054] In addition, in the case where the collected speech segment is real-time, once the collected speech segment contains the target text, the collected speech segment is regarded as a speech segment related to the target text for multi-state phoneme modeling, and it is not necessary to perform multi-state phoneme modeling on a new speech segment collected after the collected speech segment, that is, only the in-set state of the target text is subjected to multi-state phoneme modeling, and the out-set state of the target text is not subjected to multi-state phoneme modeling, thereby effectively compressing the range of multi-state phoneme modeling and greatly reducing the calculation amount required for multi-state phoneme modeling.
[0055] In addition, experiments have verified that the single-state phoneme positioning method is used to efficiently detect whether the target text exists, and in the case where the target text exists, the multi-state phoneme positioning method is used to accurately position the phoneme state level position, and the positioning error can be reduced by more than 30% compared with direct positioning by the multi-state phoneme positioning method.
[0056] Further, in some embodiments, the target single-state phoneme sequence is a target single-state monophone sequence, the target multi-state phoneme sequence is a target multi-state monophone sequence, each phoneme of the target text is in the form of a monophone in the target single-state monophone sequence and the target multi-state monophone sequence, and each phoneme of the target text is divided into a plurality of states in the target multi-state monophone sequence.
[0057] In some embodiments, the target single-state phoneme sequence is a target single-state triphone sequence, the target multi-state phoneme sequence is a target multi-state triphone sequence, each phoneme of the target text is in the form of a triphone in the target single-state triphone sequence and the target multi-state triphone sequence, and each phoneme of the target text is divided into a plurality of states in the target multi-state triphone sequence.
[0058] The multi-state can be a three-state, a four-state, a five-state, or the like. The three-state can be a start state, a stable state, and an end state.
[0059] In a specific example, the target multi-state phoneme sequence is a target triphone sequence, the target single-state phoneme sequence is a target monophone sequence, each phoneme of the target text is in the form of a triphone in the target monophone sequence and the target triphone sequence, and each phoneme of the target text is divided into a start state, a stable state and an end state in the target triphone sequence.
[0060] In some embodiments, in the case of real-time speech flow, after S110, there can further include: in response to the existence of the target single-state phoneme sequence, continuing to collect a new speech segment to update the collected speech segment, and cyclically executing S110-S120 until the existence of the target single-state phoneme sequence in the collected speech segment is detected.
[0061] In some embodiments, in the case of non-real-time, after S110, there can further include: in response to the non-existence of the target single-state phoneme sequence, ending the positioning.
[0062] Figure 2 is a flow diagram of another embodiment of the phoneme positioning method of the target text provided by the present application. The present embodiment is a further extension of S110, and the collected speech segment includes a current speech segment and a historical speech segment located before the current speech segment. As shown in Figure 2 S110 can include the following steps:
[0063] S111: obtaining a first probability that each speech frame sequence in the current speech segment belongs to each candidate phoneme class.
[0064] The speech frame sequence includes a plurality of continuous speech frames.
[0065] All pronunciations correspond to a first number of candidate phonemes. A candidate phoneme class corresponds to a candidate phoneme.
[0066] In some embodiments, different candidate phonemes correspond to different candidate phoneme classes, in which case there are a first number of candidate phoneme classes.
[0067] In some embodiments, candidate phonemes with pronunciation similarity satisfying the similarity requirement correspond to the same candidate phoneme class. The first number of candidate phonemes can be clustered according to the similarity of the pronunciations to obtain a second number of candidate phoneme sets, the second number being less than the first number, the pronunciation similarity between different candidate phonemes in the same candidate phoneme set satisfying the similarity requirement, a candidate phoneme set being a candidate phoneme class, and different candidate phoneme sets being different candidate phoneme classes. In this case, there are a second number of candidate phoneme classes. It can be understood that reducing the candidate phoneme classes from the first number to the second number can reduce the difficulty of prediction.
[0068] The mapping relationship between the candidate phoneme classes and the candidate phonemes can be stored through a phoneme ID mapping table. Each phoneme ID in the phoneme ID mapping table represents a candidate phoneme class, and the candidate phoneme corresponding to the phoneme ID represents a candidate phoneme belonging to the candidate phoneme class. For example, the phoneme ID table includes 3000 phoneme IDs, a total of 3000 candidate phoneme classes, and each phoneme ID corresponds to at least one candidate phoneme.
[0069] In some embodiments, the candidate phoneme class with the maximum first probability corresponding to the speech frame can be taken as the phoneme class to which the speech frame belongs.
[0070] A speech frame sequence can include only one speech frame or a plurality of continuous speech frames. It can be understood that the speech frame sequence is taken as a basic unit for obtaining the first probability, and the number of speech frames contained in the speech frame sequence is positively correlated with the efficiency of obtaining the first probability. Therefore, the efficiency of obtaining the first probability can be improved by increasing the number of speech frames contained in the speech frame sequence.
[0071] In some embodiments, feature extraction can be performed on the current speech segment to obtain the features of each speech frame sequence in the current speech segment; and phoneme class prediction can be performed on the features of each speech frame sequence in the current speech segment to obtain the first probability corresponding to each speech frame sequence in the current speech segment. The features of each speech frame sequence in the current speech segment can be obtained by sequentially performing frequency domain feature extraction and CNN network feature extraction on each speech frame sequence in the current speech segment. Alternatively, the features of each speech frame in the current speech segment can be obtained by sequentially performing frequency domain feature extraction and CNN network feature extraction on each speech frame in the current speech segment; and the features of each speech frame sequence in the current speech segment can be obtained by combining the features of each speech frame in the current speech segment.
[0072] S112: Jointly determining the first probability corresponding to each speech frame sequence in the current speech segment and the historical speech segment, to determine whether a target single-state phoneme sequence exists in the collected speech segment.
[0073] The first probability corresponding to each speech frame sequence in the historical speech segment is obtained in the same manner as the first probability corresponding to each speech frame sequence in the current speech segment.
[0074] In some embodiments, the first probability corresponding to each speech frame sequence in the current speech segment and the historical speech segment can be jointly decoded to obtain a score of the existence of the target single-state phoneme sequence in the collected speech segment. In the case that the score reaches a score threshold, it is determined that the target single-state phoneme sequence exists, and in the case that the score is lower than a low score threshold, it is determined that the target single-state phoneme sequence does not exist. The decoding manner can be, but is not limited to, Viterbi decoding.
[0075] Unlike the previous embodiment, in this embodiment, when a current speech segment arrives in a real-time speech stream, single-state phoneme modeling is performed on the current segment. The first probabilities corresponding to each speech frame sequence in the current segment are then determined. The first probabilities corresponding to each speech frame sequence in the current segment and previous segments are then combined to determine whether the target single-state phoneme sequence exists. This results in improved real-time performance, and because single-state phoneme modeling is used before determining whether the target single-state phoneme sequence exists, the system is highly efficient.
[0076] In some embodiments, different from S111-S112, when the collected voice segment is a non-real-time voice segment, S110 includes: obtaining a first probability that each voice frame sequence in the current collected voice segment belongs to each candidate phoneme category; based on obtaining the first probability corresponding to each voice frame sequence in the current collected voice segment, judging whether there is a target single-state phoneme sequence in the collected voice segment.
[0077] Figure 3 This is a flow chart of another embodiment of the method for locating the phonemes of the target text provided by this application. This embodiment is a further extension of S120, and the collected voice segments include the current voice segment and the historical voice segments before the current voice segment. Figure 3 As shown, in this embodiment, S120 may include the following steps:
[0078] S121: Obtain a second probability that each speech frame in the current speech segment belongs to each candidate state category of a preset phoneme and a second probability that each speech frame belongs to other state categories.
[0079] The second probability of belonging to other state categories is the sum of the second probabilities of each candidate state category belonging to other phonemes, and other phonemes are each phoneme other than the preset phonemes in the candidate phonemes.
[0080] In some embodiments, each candidate state of each candidate phoneme can be considered a different candidate state category. In this case, the number of candidate state categories is the total number of candidate states for each candidate phoneme (the third number). For example, if there are 6,000 candidate phonemes and each candidate phoneme has 3 candidate states, there are 6,000 × 3 = 18,000 candidate states in total, and the number of candidate state categories is 18,000.
[0081] In some embodiments, the candidate states of each candidate phoneme can be clustered according to the similarity of pronunciations, to obtain a fourth number of candidate state sets. The fourth number is less than the third number, the pronunciations of different candidate states in a same candidate state set satisfy the similarity requirement, a candidate state set is a candidate state category, and different candidate state sets are different candidate state categories. In this case, there are a fourth number of candidate state categories. For example, there are 18000 candidate states, and clustering the candidate states obtains 3000 candidate state sets, and there are 3000 candidate state categories. It can be understood that reducing the number of candidate phoneme categories from the third number to the fourth number can reduce the difficulty of prediction.
[0082] In some embodiments, the candidate phonemes can be clustered according to the similarity of pronunciations, to obtain a fifth number of candidate phoneme sets, the fifth number being less than the fourth number, the pronunciations of different candidate phonemes in a same candidate phoneme set satisfying the similarity requirement, a candidate phoneme set being a candidate phoneme category, and different candidate phoneme sets being different candidate phoneme categories. Each candidate state of each candidate phoneme category is taken as a candidate state category. For example, there are 6000 candidate phonemes, each candidate phoneme having 3 candidate states, clustering the candidate phonemes obtains 1000 candidate phoneme sets, and there are 1000*3=3000 candidate state categories.
[0083] The mapping relationship between the candidate state categories and the candidate states of the candidate phonemes can be stored through a phoneme state ID mapping table. Each phoneme state ID in the phoneme state ID mapping table represents a candidate state category, and the candidate state corresponding to the phoneme state ID represents a candidate state belonging to the candidate state category. For example, the phoneme state ID table includes 3000 phoneme state IDs, and there are 3000 candidate state categories, each phoneme state ID corresponding to at least one candidate state.
[0084] The candidate state category of the preset phoneme refers to the candidate state category corresponding to the candidate state of the preset phoneme. The other state category refers to the candidate state category corresponding to the candidate state of the other phoneme. For example, there are 3000 candidate state categories, the number of candidate state categories of the preset phoneme is 3 (the start state category, the stable state category, and the end state category), the number of other state categories is 3000-3=2997, and the second probability of the other state category is the sum of the second probabilities of the 2997 candidate phoneme categories.
[0085] In some embodiments, the candidate phoneme category with the maximum second probability corresponding to the speech frame can be taken as the state category to which the speech frame belongs.
[0086] In some embodiments, feature extraction can be performed on the current speech segment to obtain features of each speech frame in the current speech segment; and state category prediction can be performed on the features of each speech frame in the current speech segment to obtain the second probability corresponding to each speech frame in the current speech segment.
[0087] In S122, the second probability corresponding to each speech frame in the current speech segment and the historical speech segment is combined to determine the preset speech frame in the collected speech segment.
[0088] The second probability corresponding to each speech frame in the historical speech segment is obtained in the same way as the second probability corresponding to each speech frame in the current speech segment.
[0089] The second probability corresponding to each speech frame in the current speech segment and the historical speech segment can be combined for decoding to obtain the preset speech frame in the collected speech segment. The decoding method can be, but is not limited to, dynamic programming decoding.
[0090] Different from the foregoing embodiments, in this embodiment, in the real-time speech stream scenario, when the current speech segment arrives, the current speech segment is subjected to multi-state phoneme modeling to obtain the second probability corresponding to each speech frame in the current speech segment, and the second probability corresponding to each speech frame in the current speech segment and the historical speech segment is combined to determine the preset speech frame in the collected speech segment. Therefore, the real-time performance is good, and since it is multi-state phoneme modeling and repositioning, the positioning accuracy is high.
[0091] Further, the phoneme positioning method of the target text provided in the present application can be implemented based on a phoneme positioning model.
[0092] In some embodiments, the phoneme positioning model includes a single-state phoneme branch and a multi-state phoneme branch, and the single-state phoneme branch and the multi-state phoneme branch are independent of each other. The single-state phoneme branch can be used to perform S110, and the multi-state phoneme branch can be used to perform S120. The single-state phoneme branch includes a first feature extraction module, a phoneme category prediction module, and a first decoding module. The first feature extraction module is configured to perform feature extraction to obtain features of a speech frame sequence, the phoneme category prediction module is configured to perform phoneme category prediction based on the features of the speech frame sequence to obtain a first probability, and the first decoding module is configured to determine whether a target single-state phoneme sequence exists based on the first probability. The multi-state phoneme branch includes a second feature extraction module, a state category prediction module, and a second decoding module. The second feature extraction module is configured to perform feature extraction to obtain features of a speech frame, the state category prediction module is configured to perform state category prediction based on the features of the speech frame to obtain a second probability, and the second decoding module is configured to determine a preset speech frame corresponding to a preset state of a preset phoneme category based on the second probability.
[0093] In some embodiments, the phoneme positioning model comprises a feature extraction network, a single-state phoneme branch, and a multi-state phoneme branch. The single-state phoneme branch comprises a phoneme class prediction module and a first decoding module. The multi-state phoneme branch comprises a state class prediction module and a second decoding module. The single-state phoneme branch and the multi-state phoneme branch share the feature extraction network. The feature extraction network, the single-state phoneme branch are used to jointly perform S110, and the feature extraction network, the multi-state phoneme branch are used to jointly perform S120. Among them, the feature extraction network is used for feature extraction to obtain the features of the speech frames, the single-state phoneme branch is used for processing the features of the speech frame sequence obtained by combining the features of each speech frame, to determine whether there is a target single-state phoneme sequence. The multi-state phoneme branch is used for processing the features of the speech frame sequence obtained by combining the features of each speech frame, to determine the preset speech frame corresponding to the preset state of the preset phoneme class.
[0094] Before the phoneme positioning model is used to implement the phoneme positioning method of the target text, the phoneme positioning model needs to be trained.
[0095] Figure 4 is a flowchart of an embodiment of the phoneme positioning model training method provided by the present application. As shown in Figure 4 In the present embodiment, the training steps of the phoneme positioning model include:
[0096] S210: Obtain a first training speech and a second training speech.
[0097] Among them, the first training speech has a first label and a second label, the first label includes the true probability of each speech frame sequence in the first training speech belonging to each candidate phoneme class, and the second label includes the true probability of each speech frame in the first training speech belonging to each candidate state class of each candidate phoneme. The second training speech has a third label, and the third label includes the true probability of each speech frame in the second training speech belonging to each candidate state class of a preset phoneme and the true probability of belonging to other state classes.
[0098] The first training set includes a plurality of first training speeches, and the second training set includes a plurality of second training speeches. The plurality of first training speeches in the first training set will all participate in training, and the plurality of second training speeches in the second training set will all participate in training. In order to simplify the description, the training of the phoneme positioning model is only described by taking one first training speech and one second training speech as an example.
[0099] In some embodiments, the speech recognition complexity of the first training set and the second training set is the same.
[0100] In some embodiments, the speech recognition complexity of the first training set is higher than the speech recognition complexity of the second training set. For example, the first training set includes a rich variety of first training speech with low signal-to-noise ratio, different speech speeds, different accents, different ages and genders, existence of coarticulation, etc. Training with the rich variety of first training speech can enable the phoneme positioning model to maintain high positioning accuracy in complex speech environments.
[0101] The first label can represent the real phoneme category to which each sequence of speech frames in the first training speech belongs. For each candidate phoneme category, if the sequence of speech frames belongs to the candidate phoneme category, the real probability of the speech frame belonging to the candidate phoneme category is 1; if the sequence of speech frames does not belong to the candidate phoneme category, the real probability of the speech frame belonging to the candidate phoneme category is 0. Thus, the candidate phoneme category with a real probability of 1 is the real phoneme category to which the sequence of speech frames belongs.
[0102] The second label can represent the real state category to which each speech frame in the first training speech belongs. For each candidate state category, if the speech frame belongs to the candidate state category, the real probability of the speech frame belonging to the candidate state category is 1; if the sequence of speech frames does not belong to the candidate state category, the real probability of the speech frame belonging to the candidate state category is 0. Thus, the candidate state category with a real probability of 1 is the real state category to which the sequence of speech frames belongs.
[0103] The third label can represent whether each speech frame in the second training speech belongs to a preset phoneme and which candidate state category of the preset phoneme the speech frame belongs to. For each candidate state category of the preset phoneme and other state categories, if the speech frame belongs to the candidate state category / other state category, the real probability of the speech frame belonging to the candidate state category / other state category is 1; if the speech frame does not belong to the candidate state category / other state category, the real probability of the speech frame belonging to the candidate state category / other state category is 0. Thus, the candidate state category / other state category with a real probability of 1 is the real state category of the preset phoneme to which the sequence of speech frames belongs.
[0104] S220: training the single-state phoneme branch using the first training speech and the first label, and first-stage training the multi-state phoneme branch using the first training speech and the second label.
[0105] S230: second-stage training the multi-state phoneme branch using the second training speech and the third label.
[0106] Different from the foregoing embodiments, the present embodiment divides the training of the phoneme positioning model into two stages. In the first stage, the single-state phoneme branch is trained using the first training speech and the first label, so that the single-state phoneme branch has the ability to distinguish different candidate phoneme categories (general phoneme category classification ability). The multi-state phoneme branch is trained using the first training speech and the second label, so that the multi-state phoneme branch has the ability to distinguish different candidate state categories of different candidate phonemes (general phoneme state classification ability). In the second stage, the multi-state phoneme branch is trained using the second training speech and the third label, so as to strengthen the ability of the multi-state phoneme branch to distinguish the different candidate state categories of the preset phoneme and other state categories (preset phoneme state classification ability).
[0107] Further, in some embodiments, for the first label, the second label and the third label, they can be obtained by manual annotation, or can be obtained by a speech recognition model using forced alignment technology (FA) for automatic annotation, or can be obtained by a voice activity detection model (VAD model) for automatic annotation. The FA and VAD methods can obtain high-quality and high-quantity first labels, second labels and third labels without manual annotation, thereby reducing the annotation cost and improving the annotation efficiency.
[0108] Figure 5 is a flowchart of another embodiment of the phoneme positioning model training method provided by the present application. The present embodiment is a further extension of S220, and the phoneme positioning model further includes a feature extraction network, and the single-state phoneme branch and the multi-state phoneme branch share the feature extraction network. As shown in Figure 5 S220 can include the following steps in the present embodiment:
[0109] S221: performing feature extraction on the first training speech using the feature extraction network to obtain the features of each speech frame in the first training speech; and combining the features of each speech frame in the first training speech into the features of a plurality of speech frame sequences.
[0110] For example, the features of 10 speech frames in the first training speech are obtained, the features of the first 5 speech frames are combined into the features of one speech frame sequence, and the features of the last 5 speech frames are combined into the features of another speech frame sequence.
[0111] S222: obtaining, by the single-state phoneme branch, the first probability that each speech frame sequence in the first training speech belongs to each candidate phoneme category based on the features of each speech frame sequence in the first training speech; and obtaining, by the multi-state phoneme branch, the third probability that each speech frame in the first training speech belongs to each candidate state category of each candidate phoneme based on the features of each speech frame in the first training speech.
[0112] It can be understood that the single-state phoneme branch predicts in the unit of the sequence of speech frames, which can improve the efficiency. The multi-state phoneme branch predicts in the unit of speech frames, which can improve the accuracy.
[0113] S223: Adjust the parameters of the single-state phoneme branch and the feature extraction network based on the difference between the first probability corresponding to each sequence of speech frames in the first training speech and the first label, and adjust the parameters of the multi-state phoneme branch and the feature extraction network based on the difference between the second probability corresponding to each speech frame in the second training speech and the second label.
[0114] The first loss function can be constructed based on the difference between the first probability corresponding to each sequence of speech frames in the first training speech and the first label, and the parameters of the single-state phoneme branch and the feature extraction network are adjusted based on the first loss function.
[0115] The second loss function can be constructed based on the difference between the second probability corresponding to each speech frame in the second training speech and the second label, and the parameters of the multi-state phoneme branch and the feature extraction network are adjusted based on the second loss function.
[0116] The first loss function and the second loss function can be, but are not limited to, cross-entropy loss functions.
[0117] It can be understood that the training of the single-state phoneme branch and the training of the multi-state phoneme branch both involve adjusting the parameters of the feature extraction network, which can complement and promote the training of the single-state phoneme branch and the multi-state phoneme branch.
[0118] For other detailed descriptions of the present embodiment, please refer to the related descriptions of the previous embodiments, which will not be repeated here.
[0119] Figure 6 is a flowchart of another embodiment of the phoneme positioning model training method provided by the present application. The present embodiment is a further extension of S230, and the phoneme positioning model further includes a feature extraction network, and the single-state phoneme branch and the multi-state phoneme branch share the feature extraction network. As shown in Figure 6 S230 can include the following steps in the present embodiment:
[0120] S231: Extract features of the second training speech by using the feature extraction network to obtain the features of each speech frame in the second training speech.
[0121] S232: Obtain the second probability of each candidate state category of the preset phoneme and the second probability of other state categories of each speech frame in the second training speech by using the multi-state phoneme branch based on the features of each speech frame in the second training speech.
[0122] S233: adjusting parameters of the multi-state phoneme branch based on a difference between the second probability corresponding to each speech frame in the second training speech and the third label.
[0123] The third loss function can be constructed based on a difference between the second probability corresponding to each speech frame in the second training speech and the third label, and the parameters of the multi-state phoneme branch are adjusted based on the third loss function. The third loss function can be, but is not limited to, a cross-entropy loss function.
[0124] The multi-state phoneme branch includes a plurality of nodes, and different nodes are used to obtain probabilities of a speech frame belonging to different candidate state categories of different candidate phonemes.
[0125] In the first stage of training, the multi-state phoneme branch is a multi-state full-space phoneme branch, and the parameters of the corresponding nodes are adjusted based on the third probability of a speech frame belonging to different candidate state categories of different candidate phonemes.
[0126] In the second stage of training, the multi-state full-space phoneme branch nodes collapse into a multi-state collapsed phoneme branch. The so-called node collapse refers to the merging of nodes of other state categories into one node. At this time, the nodes of the multi-state collapsed phoneme branch include nodes corresponding to each candidate state category of a preset phoneme and the merged node. The output of the merged node is the second probability of a speech frame belonging to other state categories. The output of the node corresponding to the candidate state category of the preset phoneme is the second probability of a speech frame belonging to the candidate state category of the preset phoneme. The parameters of the corresponding nodes are adjusted based on the second probability of the candidate state category of the preset phoneme and the probability of belonging to other state categories.
[0127] Further, the phoneme positioning method of the target text provided by the present application can be applied to a speech segmentation scene. One of the speech segmentation scenes can be the separation between a wake-up instruction and an action instruction in a voice instruction. The voice instruction is used to control the device. The wake-up instruction is an instruction containing a wake-up word, which is used to control the device to switch from a sleep state to a wake-up state. The action instruction is an instruction containing an action word, which is used to control the device to perform a corresponding action.
[0128] It can be understood that there are two ways for a user to issue a voice instruction to a device. The first way is to issue a wake-up instruction and an action instruction as two independent voice instructions, that is, the wake-up instruction is issued first, and after the device switches from a sleep state to a wake-up state in response to the wake-up instruction, the action instruction is issued to make the device perform an action corresponding to the action instruction. For example, the wake-up instruction containing "Hello Harry" is issued first, and after the device switches to the wake-up state in response to the wake-up instruction, the action instruction containing "turn on the TV" is issued, and the device performs the action of turning on the TV in response to the action instruction. The second way is to issue the wake-up instruction and the action instruction in one voice instruction, and the device switches from the sleep state to the wake-up state in response to the voice instruction and further performs a corresponding action.
[0129] The first way described above does not interfere with the recognition of the action instruction, but the user needs to wait for the device to respond to the wake-up instruction before issuing the action instruction, and the waiting time affects the user experience. The second way described above is free of waiting, but the wake-up instruction can interfere with the recognition of the action instruction, resulting in a decrease in the recognition accuracy of the action instruction.
[0130] In order to be free of waiting while avoiding the recognition accuracy of the action instruction from being affected, the wake-up instruction in the voice instruction can be separated by using the phoneme positioning method for a target text provided in the present application. Specifically as follows:
[0131] Figure 7 is a flowchart of another embodiment of the phoneme positioning method for a target text provided in the present application. The present embodiment is a further extension of the above-mentioned embodiment, and the target text is a wake-up word. As shown in Figure 7 S120 can further include the following after S120:
[0132] S130: regarding the preset voice frame and the voice frame before the preset voice frame as the wake-up instruction, separating the voice frame corresponding to the wake-up instruction from the collected voice segment, and continuing to collect a new voice segment, and using the new voice segment and the separated collected voice segment to form an action instruction.
[0133] In the present embodiment, the voice instruction is issued in real time, and the collected voice segment is the part collected in the voice instruction. It can be determined by the foregoing S110-S120 that the collected voice segment contains the wake-up word, and the preset voice frame corresponding to the wake-up word. The preset voice frame corresponding to the wake-up word can be any voice frame in which the last state of the last phoneme of the wake-up word is located, for example, the first voice frame in which the last state of the last phoneme of the wake-up word is located.
[0134] S140: controlling the device in the wake-up state to perform a corresponding action in response to the action instruction.
[0135] The device in the wake-up state is woken up based on the wake-up instruction.
[0136] In some embodiments, the device in the wake-up state is woken up at the first speech frame at which the last phoneme of the wake-up word occurs.
[0137] In some embodiments, the device in the wake-up state is woken up at any speech frame (such as the first speech frame) at which the last state of the last phoneme of the wake-up word occurs.
[0138] Through implementation of the embodiment, the wake-up instruction can be accurately separated from the speech instruction, avoiding interference of the wake-up instruction on recognition of the action instruction, and improving recognition accuracy of the action instruction on the basis of no waiting.
[0139] In order to facilitate understanding of the present application, the application of the phoneme positioning method of the target text provided by the present application to the wake-up instruction separation scene in the speech instruction is described in the form of a specific example as follows.
[0140] The target text is "Hello Harry". Figure 8 is a structural schematic diagram of the phoneme positioning model of the present application. As shown in Figure 8 , the phoneme positioning model includes a feature extraction network, a single-state phoneme branch, and a multi-state phoneme branch. The single-state phoneme branch includes a phoneme class prediction module and a first decoding module. The multi-state phoneme branch includes a state class prediction module and a second decoding module.
[0141] I. Model training phase
[0142] Figure 9 is a specific example schematic diagram of one-stage training of the phoneme positioning model of the present application. As shown in Figure 9 , the one-stage training process includes:
[0143] 1. Obtain a first training speech, the first training speech having a first label and a second label.
[0144] 2. Extract features of each speech frame in the first training speech by using the feature extraction network, and combine the features of each speech frame into features of a plurality of speech frame sequences.
[0145] 3. Perform phoneme class prediction on the features of each speech frame sequence in the first training speech by using the phoneme class prediction module, to obtain a first probability matrix (size T / 4*P, T / 4 represents the number of speech frame sequences, and P represents the number of candidate phoneme classes) belonging to each candidate phoneme class. The phoneme class prediction module and the feature extraction network are trained by using the difference between the first probability and the first label.
[0146] 4. The state category prediction module is used to predict the state category of each feature sequence in the first training speech, and a third probability matrix (size T*K, T represents the number of speech frames, and K represents the total number of each candidate state category of each candidate phoneme) of each candidate state category belonging to each candidate phoneme is obtained. The state category prediction module and the feature extraction network are trained by using the difference between the third probability and the second label.
[0147] Figure 10 is a specific example of the two-stage training of the phoneme positioning model. As shown in Figure 10 , the two-stage training process includes:
[0148] 5. Obtain the second training speech, and the second training speech has a third label.
[0149] 6. The state category prediction module is used to predict the state category of each feature sequence in the second training speech, and a second probability of each candidate state category belonging to the last phoneme of "Hello Harry" and a second probability of other state categories are obtained, which constitute a second probability matrix (size T*4, T represents the number of speech frames, and 4 represents the number of predicted state categories, including 3 candidate state categories of the last phoneme of "Hello Harry" and other state categories). The state category prediction module is trained by using the difference between the second probability and the third label.
[0150] II. Model application inference stage
[0151] Figure 11 is a specific example of the inference stage of the phoneme positioning model. As shown in Figure 11 , the inference stage process includes:
[0152] 1. The current speech segment i is collected.
[0153] 2. The feature extraction network is used to extract the features of the speech frames i1-i8 in the speech segment i.
[0154] 3. The features of the speech frames i1-i4 and the features of the speech frames i5-i8 are combined into the features of the speech frame sequence 1 and the features of the speech frame sequence 2, respectively. Each speech frame sequence includes 4 consecutive speech frames.
[0155] 4. The phoneme category prediction module is used to predict the phoneme category of the features of the speech frame sequence 1 and the features of the speech frame sequence 2, and a first probability matrix (size 2*P, 2 represents the number of speech frame sequences, and P represents the number of candidate phoneme categories) of each candidate phoneme category belonging to the speech frame sequence 1 and the speech frame sequence 2 is obtained.
[0156] 5. Utilize the first decoding module to jointly the first probability matrix of each speech frame sequence in the collected speech segment 1-i belonging to each candidate phoneme category to perform Viterbi decoding to determine the score of the target single-state phoneme sequence existing in the collected speech segment 1-i; and determine that the target single-state phoneme sequence exists in the case that the score is greater than the score threshold.
[0157] 6. In response to the target single-state phoneme sequence existing, utilize the state category prediction module to perform state category prediction on the features of the speech frames i1-i8 to obtain the second probability of each candidate state category of the speech frames i1-i8 belonging to the last phoneme of “Hello Harry”, the second probability of belonging to other state categories, and form a second probability matrix (size 8*4, 8 representing the number of speech frames, and 4 representing the number of predicted state categories, including 3 candidate state categories of the last phoneme of “Hello Harry” and other state categories).
[0158] 7. Utilize the second decoding module to jointly the second probability of each speech frame in the collected speech segment 1-i belonging to each candidate state category of the preset phoneme and the second probability of belonging to other state categories to perform dynamic programming decoding to determine the last state of the last phoneme of “Hello Harry”, and the corresponding first speech frame in the collected speech segment 1-i is speech frame i7.
[0159] 8. Take the speech frame 11-speech frame i7 as a wake-up instruction to separate from the collected speech segment 1-i, and the separated collected speech segment 1-i is speech frame i8.
[0160] 9. Continue to collect new speech segments i+1-N, and utilize the speech frame i8 and the new speech segments i+1-N to form an action instruction “turn on the TV”.
[0161] 10. In response to the action instruction, control the device in the wake-up state to perform the action of “turning on the TV”.
[0162] Figure 12 is a comparison diagram of the single-state phoneme and the multi-state phoneme modeling modes of the present application. As shown in Figure 12 , for the above collected speech segment 1-i, the single-state phoneme modeling mode is modeled as 8 phonemes. The multi-state phoneme modeling mode is modeled as 24 phoneme states, and 8 phonemes are modeled as 3 states, i.e., the start state, the stable state, and the end state. The first speech frame of the appearance of the 8th phoneme is taken as the wake-up point, and the first speech frame (speech frame i7) of the appearance of the end state of the 8th phoneme is taken as the segmentation point.
[0163] Figure 13 is a schematic diagram of the present application separating the wake-up instruction from the speech instruction. As shown in Figure 13As shown, the voice instruction includes a wake-up instruction (voice segments 1-i) and an action instruction (illustrated as a recognition instruction, containing voice segments i+1-N), the first voice frame corresponding to the last phoneme of "Hello Harry" in the wake-up instruction is taken as a wake-up point, the device is controlled to switch from a sleep state to a wake-up state at the wake-up point, the first voice frame (voice frame i7) corresponding to the last state of the last phoneme of the wake-up word is taken as a cutting point, and the voice frame before the cutting point is separated from the voice instruction to obtain a separated voice instruction (illustrated as a recognition audio after cutting).
[0164] Figure 14 is a structural schematic diagram of an embodiment of the phoneme positioning apparatus of the target text. As shown in the figure, Figure 14 The phoneme positioning apparatus 30 of the target text includes a detection module 31 and a determination module 32. Wherein:
[0165] The detection module 31 is configured to detect whether a target single-state phoneme sequence corresponding to the target text exists in the collected voice segment.
[0166] The determination module 32 is configured to determine a preset state of a preset phoneme in a target multi-state phoneme sequence corresponding to the target text and determine a preset voice frame corresponding to the preset state of the preset phoneme in the collected voice segment in response to the existence of the target single-state phoneme sequence.
[0167] For other detailed descriptions of the phoneme positioning apparatus 30 of the target text, please refer to the previous embodiments, which will not be repeated here.
[0168] Figure 15 is a structural schematic diagram of an embodiment of the electronic device. As shown in the figure, Figure 15 The electronic device 40 includes a memory 41 and a processor 42, and the processor 42 is configured to execute program instructions stored in the memory 41 to implement the steps in any of the method embodiments described above. In one specific implementation scenario, the electronic device 40 can include but is not limited to a microcomputer, a server, in addition, the electronic device 40 can also include a notebook computer, a tablet computer and other carrying devices, which are not limited here.
[0169] In particular, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above method embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 can be an integrated circuit chip with signals. The processor 42 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 42 can be implemented by an integrated circuit chip together.
[0170] Referring to Figure 16 , Figure 16 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 50 stores program instructions 51, and the program instructions 51 are executed by the processor to implement the steps in any of the above method embodiments.
[0171] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.
[0172] The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be mutually referred to. For brevity, details are not repeated here.
[0173] In several embodiments provided in the present application, it should be understood that the disclosed method and device can be implemented by other ways. For example, the above-described device implementation is only schematic, for example, the division of the module or unit is only a logical function division, and actual implementation can have another division manner, for example, the unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In another image position, the coupling or direct coupling or communication connection between the displayed or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0174] In addition, each of the function units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit. When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
Claims
1. A method for phoneme localization of a target text, characterized in that: include: Detecting whether there is a target single-state phoneme sequence corresponding to a target text in a collected speech segment, wherein the collected speech segment includes a current speech segment and a historical speech segment located before the current speech segment, including: obtaining a first probability that each speech frame sequence in the current speech segment belongs to each candidate phoneme category, wherein the speech frame sequence includes a plurality of consecutive speech frames; combining the first probabilities corresponding to each speech frame sequence in the current speech segment and the historical speech segment, and judging whether the target single-state phoneme sequence exists in the collected speech segment, wherein the first probability corresponding to each speech frame sequence in the historical speech segment is obtained in the same manner as the first probability corresponding to each speech frame sequence in the current speech segment, all pronunciations correspond to a first number of candidate phonemes, different candidate phonemes correspond to different candidate phoneme categories, or candidate phonemes whose pronunciation similarity meets the similarity requirement correspond to the same candidate phoneme category; In response to the existence of the target single-state phoneme sequence, the preset state of the preset phoneme in the target multi-state phoneme sequence corresponding to the target text is determined, and the preset speech frame corresponding to the preset state of the preset phoneme in the collected speech segment is determined, including: obtaining the second probability that each speech frame in the current speech segment belongs to each candidate state category of the preset phoneme and the second probability that it belongs to other state categories, wherein the second probability of belonging to other state categories is the sum of the second probabilities of each candidate state category of other phonemes, and the other phonemes are each phoneme other than the preset phoneme in the candidate phonemes; combining the second probabilities corresponding to each speech frame in the current speech segment and the historical speech segment, to determine the preset speech frame in the collected speech segment, wherein the second probability corresponding to each speech frame in the historical speech segment is obtained in the same way as the second probability corresponding to each speech frame in the current speech segment.
2. The method according to claim 1, characterized in that The obtaining of first probabilities that each speech frame sequence in the current speech segment belongs to each candidate phoneme category includes: Extracting features of the current speech segment to obtain features of each speech frame sequence in the current speech segment; Phoneme category prediction is performed on features of each speech frame sequence in the current speech segment to obtain a first probability corresponding to each speech frame sequence in the current speech segment.
3. The method according to claim 1, characterized in that The obtaining of the second probability that each speech frame in the current speech segment belongs to each candidate state category of the preset phoneme and the second probability that each speech frame belongs to other state categories includes: Extracting features of the current speech segment to obtain features of each speech frame in the current speech segment; A state category prediction is performed on features of each speech frame in the current speech segment to obtain a second probability corresponding to each speech frame in the current speech segment.
4. The method according to claim 1, wherein The target text phoneme localization method is implemented based on a phoneme localization model, wherein the phoneme localization model includes a single-state phoneme branch and a multi-state phoneme branch. Before the phoneme localization model is used to implement the target text phoneme localization method, the method further includes a training step of the phoneme localization model: Obtaining a first training speech and a second training speech, wherein the first training speech has a first label and a second label, the first label including a true probability of each speech frame sequence in the first training speech belonging to each candidate phoneme category, the second label including a true probability of each speech frame in the first training speech belonging to each candidate state category of each candidate phoneme, and the second training speech has a third label including a true probability of each speech frame in the second training speech belonging to each candidate state category of a preset phoneme and a true probability of belonging to other state categories; Training the single-state phoneme branch using the first training speech and the first label, and performing a first-stage training on the multi-state phoneme branch using the first training speech and the second label; The second training speech and the third label are used to perform second-stage training on the multi-state phoneme branch.
5. The method according to claim 4, characterized in that The phoneme localization model further includes a feature extraction network, wherein the single-state phoneme branch is trained using the first training speech and the first label, and the multi-state phoneme branch is trained in a first phase using the first training speech and the second label, including: Using the feature extraction network to perform feature extraction on the first training speech to obtain features of each speech frame in the first training speech; and combining the features of each speech frame in the first training speech into features of a plurality of speech frame sequences; Obtaining, using the single-state phoneme branch based on features of each speech frame sequence in the first training speech, a first probability that each speech frame sequence in the first training speech belongs to each candidate phoneme category; and obtaining, using the multi-state phoneme branch based on features of each speech frame in the first training speech, a third probability that each speech frame in the first training speech belongs to each candidate state category of each candidate phoneme; Adjusting parameters of the single-state phoneme branch and the feature extraction network based on a difference between a first probability corresponding to each speech frame sequence in the first training speech and the first label; and adjusting parameters of the multi-state phoneme branch and the feature extraction network based on a difference between a second probability corresponding to each speech frame in the first training speech and the second label; And / or, performing a second stage of training on the multi-state phoneme branch using the second training speech and the third label includes: Performing feature extraction on the second training speech using the feature extraction network to obtain features of each speech frame in the second training speech; Obtaining, using the multi-state phoneme branch based on features of each speech frame in the second training speech, a second probability that each speech frame in the second training speech belongs to each candidate state category of a preset phoneme and a second probability that each speech frame in the second training speech belongs to other state categories; Adjust the parameters of the multi-state phoneme branch based on the difference between the second probability corresponding to each speech frame in the second training speech and the third label.
6. The method according to claim 1, characterized in that The target text is a wake-up word; After determining the preset state of the preset phoneme corresponding to the preset speech frame in the collected speech segment, the method further includes: The preset voice frame and the voice frame before the preset voice frame are regarded as a wake-up instruction, the voice frame corresponding to the wake-up instruction is separated from the collected voice segments, and a new voice segment is continuously collected, and the action instruction is formed by using the new voice segment and the separated collected voice segment; In response to the action instruction, the device in the awake state is controlled to perform a corresponding action, and the device in the awake state is awakened based on the awakening instruction.
7. The method according to claim 1, characterized in that The preset speech frame corresponding to the preset state of the preset phoneme is the first speech frame in which the ending state of the last phoneme in the target multi-state phoneme sequence is located; and / or The target multi-state phoneme sequence is a target three-state three-phoneme sequence, the target single-state phoneme sequence is a target single-state three-phoneme sequence, the individual phonemes of the target text are in the form of three phones in the target single-state phoneme sequence and the target three-state three-phoneme sequence, and the individual phonemes of the target text are divided into a starting state, a stable state and an ending state in the target three-state three-phoneme sequence.
8. A device for locating phonemes of a target text, characterized in that: include: A detection module is used to detect whether a target single-state phoneme sequence corresponding to a target text exists in a collected speech segment, wherein the collected speech segment includes a current speech segment and a historical speech segment located before the current speech segment, including: obtaining a first probability that each speech frame sequence in the current speech segment belongs to each candidate phoneme category, wherein the speech frame sequence includes a plurality of consecutive speech frames; combining the first probabilities corresponding to each speech frame sequence in the current speech segment and the historical speech segment, and judging whether the target single-state phoneme sequence exists in the collected speech segment, wherein the first probability corresponding to each speech frame sequence in the historical speech segment is obtained in the same manner as the first probability corresponding to each speech frame sequence in the current speech segment, all pronunciations correspond to a first number of candidate phonemes, different candidate phonemes correspond to different candidate phoneme categories, or candidate phonemes whose pronunciation similarity meets the similarity requirement correspond to the same candidate phoneme category; A determination module is used to determine the preset state of the preset phoneme in the target multi-state phoneme sequence corresponding to the target text in response to the existence of the target single-state phoneme sequence, and determine the preset speech frame corresponding to the preset state of the preset phoneme in the collected speech segment, including: obtaining the second probability that each speech frame in the current speech segment belongs to each candidate state category of the preset phoneme and the second probability that it belongs to other state categories, wherein the second probability of belonging to other state categories is the sum of the second probabilities of each candidate state category of other phonemes, and the other phonemes are each phoneme other than the preset phoneme in the candidate phonemes; combining the second probabilities corresponding to each speech frame in the current speech segment and the historical speech segment, to determine the preset speech frame in the collected speech segment, wherein the second probability corresponding to each speech frame in the historical speech segment is obtained in the same way as the second probability corresponding to each speech frame in the current speech segment.
9. An electronic device, characterized in that: The system comprises a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Voice processing method and equipment
CN109994106A
Training method of voice wake-up model, wake-up word detection method and related equipment
CN113963688A
Speech recognition method and device
CN115019779A