Phoneme sequence labeling method and device, equipment and storage medium

By using phoneme sequence labeling models and multiple topological structures to represent different types of phonemes in the field of singing synthesis, the problem of low phoneme-level labeling efficiency in audio data in the prior art is solved, and more efficient and accurate phoneme sequence labeling is achieved.

CN120148480APending Publication Date: 2025-06-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311694942.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has low efficiency in the field of singing synthesis, which is mainly due to the lack of accuracy due to the dependence on manual annotation and the existence of polyphonic characters.

Method used

By obtaining the target audio text pair and inputting it into the trained phoneme sequence annotation model, different types of phonemes are represented using multiple topological structures, including phonemes that do not carry sound characteristics, phonemes that carry short sound characteristics, and phonemes that carry long sound characteristics, thereby realizing automatic annotation of phoneme sequences.

Benefits of technology

It improves the efficiency and accuracy of phoneme sequence labeling, can express the phoneme drag characteristics more accurately, and reduces the dependence on manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148480A_ABST
    Figure CN120148480A_ABST
Patent Text Reader

Abstract

The invention provides a phoneme sequence labeling method and device, equipment and a storage medium, and the method comprises the steps: firstly, obtaining a target audio text pair, then inputting a target audio segment and a target text segment in a target audio text into a trained phoneme sequence labeling model, and carrying out the processing of the phoneme sequence labeling model, thereby obtaining a phoneme sequence labeling result. And obtaining a target phoneme sequence of the target audio text pair. Wherein the target phoneme sequence comprises phonemes with a time sequence relationship, a first phoneme in the target phoneme sequence is represented by adopting a first topological structure in the phoneme sequence labeling model, and the first topological structure comprises a preset first number of states with the time sequence relationship; each state is configured with a time parameter for representing the duration of pronunciation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular, to a method, apparatus, device, and storage medium for phoneme sequence annotation. Background Art

[0002] A phoneme sequence usually includes multiple phonemes with a time-sequential relationship. Among them, phonemes generally include initials and finals in Chinese character pronunciation. In the field of singing synthesis technology, how to perform phoneme-level annotation on a segment of audio data has attracted increasing attention.

[0003] Currently, generally, the phoneme-level annotation of audio data is achieved through manual annotation. Specifically, for a segment of audio data, first, the corresponding text data (i.e., lyrics) is determined, then the text data is pinyin-annotated, and then the phoneme sequence annotation is performed based on the pinyin annotation result. It can be seen that the above-mentioned phoneme sequence annotation method has low efficiency. Summary of the Invention

[0004] To solve the above technical problems, an embodiment of the present disclosure provides a method for phoneme sequence annotation.

[0005] In a first aspect, the present disclosure provides a method for phoneme sequence annotation, the method including:

[0006] Obtain a target audio-text pair; wherein, the target audio-text pair includes a target audio segment and a target text segment with a corresponding relationship;

[0007] Input the target audio-text pair into a phoneme sequence annotation model, and after being processed by the phoneme sequence annotation model, obtain a target phoneme sequence of the target audio-text pair; wherein, the target phoneme sequence includes phonemes with a time-sequential relationship, the first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model, the first topological structure includes a preset first number of states with a time-sequential relationship, and the state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence annotation model is trained by using audio segment samples and phoneme sequence samples with a corresponding relationship.

[0008] In an optional implementation manner, the second phoneme in the target phoneme sequence is represented by a second topological structure in the phoneme sequence annotation model, the first phoneme is a phoneme without a trailing sound feature, the second phoneme is a phoneme with a first trailing sound feature, and the second topological structure includes a preset second number of states with a time-sequential relationship, and the preset second number is greater than the preset first number.

[0009] In an alternative embodiment, the third phoneme in the target phoneme sequence is represented by a third topological structure in the phoneme sequence annotation model. The third phoneme is a phoneme carrying a second drawl feature, the first drawl feature is a short drawl feature, the second drawl feature is a long drawl feature, and the third topological structure includes a preset third number of states having a chronological order relationship, and the preset third number is greater than the preset second number.

[0010] In an alternative embodiment, before inputting the target audio text pair into the phoneme sequence annotation model and obtaining the target phoneme sequence corresponding to the target audio text pair after being processed by the phoneme sequence annotation model, it further includes:

[0011] Obtain audio segment samples and phoneme sequence samples with corresponding relationships;

[0012] Model the phonemes in the phoneme sequence sample by using at least one of the first topological structure, the second topological structure, and the third topological structure to obtain a modeling result corresponding to the phoneme sequence sample;

[0013] Train the model by using the modeling results corresponding to the audio segment samples and phoneme sequence samples to obtain the phoneme sequence annotation model.

[0014] In an alternative embodiment, the first topological structure belongs to a three-state structure, and the second topological structure belongs to a six-state structure or a nine-state structure.

[0015] In an alternative embodiment, the first topological structure belongs to a three-state structure, the second topological structure belongs to a six-state structure, and the third topological structure belongs to a nine-state structure.

[0016] In an alternative embodiment, the state is further configured with fields for storing phoneme identifiers and state identifiers.

[0017] In a second aspect, an embodiment of the present disclosure provides a phoneme sequence annotation device, and the device includes:

[0018] A first acquisition module, configured to acquire a target audio text pair; wherein, the target audio text pair includes a target audio segment and a target text segment with a corresponding relationship;

[0019] A labeling processing module, configured to input the target audio text pair into a phoneme sequence labeling model. After being processed by the phoneme sequence labeling model, a target phoneme sequence of the target audio text pair is obtained. Wherein, the target phoneme sequence includes phonemes with a chronological relationship, the first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence labeling model, the first topological structure includes a preset first number of states with a chronological relationship, and the state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence labeling model is trained by using corresponding audio segment samples and phoneme sequence samples.

[0020] In a third aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device implements the above method.

[0021] In a fourth aspect, the present disclosure provides a phoneme sequence labeling device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.

[0022] In a fifth aspect, the present disclosure provides a computer program product, which includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the above method is implemented.

[0023] The technical solutions provided in the embodiments of the present disclosure have at least the following advantages compared with the prior art:

[0024] The embodiments of the present disclosure provide a phoneme sequence labeling method. First, a target audio text pair is obtained. Then, the target audio segment and the target text segment in the target audio text are input into a trained phoneme sequence labeling model. After being processed by the phoneme sequence labeling model, a target phoneme sequence of the target audio text pair is obtained. Wherein, the target phoneme sequence includes phonemes with a chronological relationship, the first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence labeling model, the first topological structure includes a preset first number of states with a chronological relationship, and each state is configured with a time parameter for characterizing the pronunciation duration. It can be seen that the embodiments of the present disclosure use a trained phoneme sequence labeling model to implement the labeling of the phoneme sequence, which can improve the efficiency of phoneme sequence labeling. Description of the Drawings

[0025] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0026] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 Flowchart of a phoneme sequence annotation method provided by an embodiment of the present disclosure;

[0028] Figure 2 Schematic diagram of the topological structure of a phoneme provided by an embodiment of the present disclosure;

[0029] Figure 3 Schematic diagram of another topological structure of a phoneme provided by an embodiment of the present disclosure;

[0030] Figure 4 Schematic diagram of yet another topological structure of a phoneme provided by an embodiment of the present disclosure;

[0031] Figure 5 Schematic diagram of the structure of a phoneme sequence annotation device provided by an embodiment of the present disclosure;

[0032] Figure 6 Schematic diagram of the structure of a phoneme sequence annotation device provided by an embodiment of the present disclosure. Detailed implementation manners

[0033] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0034] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0035] A phoneme sequence includes a plurality of phonemes having a chronological relationship. A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one pronunciation action constitutes one phoneme, and phonemes are divided into two major categories: vowels and consonants. In Chinese, it generally includes initials and finals in the pronunciation of characters.

[0036] In one implementation, the phoneme sequence is labeled only based on the text segment corresponding to the audio segment. Due to the existence of polyphonic characters, the accuracy of the result of labeling the phoneme sequence using this implementation is insufficient. Therefore, it is necessary to determine the pronunciation of polyphonic characters by combining audio information, so as to label the phoneme sequence more accurately.

[0037] In another implementation, the phoneme sequence is labeled only based on the audio segment. Since there may be pronunciation errors in the audio, the accuracy of the result of labeling the phoneme sequence using this implementation is insufficient. Therefore, it is necessary to correct the wrong pronunciation by combining text information, so as to label the phoneme sequence more accurately.

[0038] In addition, there may be a feature of prolonged sounds in audio segments in the form of songs, etc. However, this feature is not labeled in the above-mentioned implementations, resulting in the fact that the labeled phoneme sequence cannot accurately express the pronunciation characteristics of the audio segment.

[0039] For this reason, the present disclosure provides a method for labeling a phoneme sequence. First, a target audio-text pair is obtained. Then, the target audio segment and the target text segment in the target audio-text pair are input into a trained phoneme sequence labeling model. After being processed by the phoneme sequence labeling model, the target phoneme sequence of the target audio-text pair is obtained. Among them, the target phoneme sequence includes phonemes with a time-sequence relationship. The first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence labeling model. The first topological structure includes a preset first number of states with a time-sequence relationship, and each state is respectively configured with a time parameter for characterizing the pronunciation duration. It can be seen that the embodiments of the present disclosure use a trained phoneme sequence labeling model to implement the labeling of the phoneme sequence, which can improve the efficiency of phoneme sequence labeling.

[0040] In addition, the embodiments of the present disclosure can also improve the accuracy of phoneme sequence labeling by combining audio segments and text segments to implement phoneme sequence labeling.

[0041] Based on this, the embodiments of the present disclosure provide a method for labeling a phoneme sequence. Refer to Figure 1 , which is a flowchart of a method for labeling a phoneme sequence provided by the embodiments of the present disclosure, and specifically includes:

[0042] S101: Obtain a target audio-text pair.

[0043] Among them, the target audio-text pair includes a target audio segment and a target text segment with a corresponding relationship.

[0044] In the embodiments of the present disclosure, there is a corresponding relationship between the target text segment and the target audio segment. The target text segment and the target audio segment belong to different resource carriers for the same or similar content. For example, the target audio segment belongs to a segment of audio data in a certain song, and the target text segment is the lyric text corresponding to the audio data; the target audio segment belongs to a segment of audio data in a certain movie, and the target text segment is the line text corresponding to the audio data.

[0045] S102: Input the target audio-text pair into the phoneme sequence annotation model. After being processed by the phoneme sequence annotation model, obtain the target phoneme sequence of the target audio-text pair.

[0046] Among them, the target phoneme sequence includes phonemes with a chronological order. The first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model. The first topological structure includes a preset first number of states with a chronological order. The state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence annotation model is trained using corresponding audio segment samples and phoneme sequence samples.

[0047] In the embodiments of the present disclosure, after obtaining the target audio-text pair, input the target audio segment and the target text segment in the target audio-text pair into the trained phoneme sequence annotation model. After being processed by the phoneme sequence annotation model, output the phoneme sequence as the target phoneme sequence of the target audio-text pair.

[0048] Among them, the target phoneme sequence includes the phoneme sequence corresponding to the target audio segment in the target audio-text pair, including phonemes with a chronological order.

[0049] In the embodiments of the present disclosure, phonemes can be represented by a first topological structure during the training stage and the application stage (also known as the inference stage) of the phoneme sequence annotation model. Specifically, the first topological structure includes a preset first number of states with a chronological order.

[0050] In an optional implementation manner, the first topological structure can be a three-state structure, that is, it includes three states. Specifically, each phoneme in the phoneme sequence annotation model can be represented by a three-state structure. As Figure 2 shown, it is a schematic diagram of a topological structure of a phoneme provided by the embodiments of the present disclosure. Among them, each state in the first topological structure has two outgoing arcs. Represented by the arrowed arcs in Figure 2 , it can only jump to itself or the next state. Among them, any state has corresponding jump probabilities when jumping to itself or the next state. As Figure 2As shown, the transition probability P1 of jumping to itself is 0.75, and the transition probability P2 of jumping to the next state is 0.25. It should be noted that the above are only exemplary transition probabilities and do not constitute a limitation on the transition probability setting in the embodiments of the present disclosure. In addition, the transition probabilities corresponding to different states jumping to themselves or the next state can be the same or different. Each state is configured with a time parameter, which is used to characterize the pronunciation duration of each state of the phoneme. Specifically, it can be achieved by setting a time value for each outgoing arc. The time value set for each outgoing arc can be a fixed duration, such as 10 ms. By increasing the number of outgoing arcs, the characterized pronunciation duration can be increased.

[0051] In addition, the topological structure of different phonemes has corresponding phoneme identifiers. For example, the topological structure of phoneme n has the phoneme identifier "n". Each state in the topological structure of phoneme n can have a state identifier respectively. For example, 1, 2, and 3 are used to represent three different states. Using n1, n2, and n3 to represent the three different states of phoneme n reflects both the phoneme identifier n and the state identifiers 1, 2, and 3.

[0052] In the phoneme sequence annotation method provided by the embodiments of the present disclosure, first, a target audio text pair is obtained. Then, the target audio segment and the target text segment in the target audio text are input into a trained phoneme sequence annotation model. After being processed by the phoneme sequence annotation model, the target phoneme sequence of the target audio text pair is obtained. Among them, the target phoneme sequence includes phonemes with a time sequence relationship. The first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model. The first topological structure includes a preset first number of states with a time sequence relationship, and each state is respectively configured with a time parameter for characterizing the pronunciation duration. It can be seen that the embodiments of the present disclosure use a trained phoneme sequence annotation model to implement the annotation of the phoneme sequence, which can improve the efficiency of phoneme sequence annotation.

[0053] In addition, the embodiments of the present disclosure can also improve the accuracy of phoneme sequence annotation by combining the audio segment and the text segment.

[0054] In an alternative embodiment, during the training stage and the application stage (also called the inference stage) of the phoneme sequence annotation model, the first phoneme can be represented by a first topological structure, and the second phoneme can be represented by a second topological structure. That is to say, different types of phonemes can be represented by different topological structures respectively. Specifically, the first phoneme is a phoneme without a trailing sound feature, that is, a phoneme whose pronunciation does not have a trailing sound. Among them, the trailing sound feature refers to a relatively long pronunciation duration of the corresponding phoneme. And the second phoneme is a phoneme with a trailing sound feature, that is, a phoneme whose pronunciation has a trailing sound.

[0055] Among them, the first topological structure includes a preset first number of states with a chronological relationship, while the second topological structure includes a preset second number of states with a chronological relationship. Specifically, the preset second number is greater than the preset first number, that is, a phoneme with a drawn-out pronunciation needs to be represented by a topological result with a larger number of states to reflect its drawn-out feature.

[0056] As Figure 3 shown, an embodiment of the present disclosure provides a schematic diagram of another topological structure of a phoneme. Among them, the first topological structure belongs to a three-state structure, and the second topological structure belongs to a six-state structure or a nine-state structure. Figure 3 Taking the nine-state structure as an example. The basic duration corresponding to each state is fixed. A larger number of states indicates that the pronunciation duration of the phoneme it represents is longer. Therefore, a topological structure with a larger number of states can be used to represent a phoneme with a drawn-out feature.

[0057] Based on the above embodiments, the drawn-out feature can be further specifically divided into a short drawn-out feature and a long drawn-out feature. Among them, the short drawn-out feature can be a drawn-out duration with a pronunciation duration less than a preset duration threshold, and the long drawn-out feature can be a drawn-out duration with a pronunciation duration not less than the preset duration threshold.

[0058] Based on this, in the training stage and the application stage (also called the inference stage) of the phoneme sequence annotation model, the first phoneme can be represented by the first topological structure, the second phoneme can be represented by the second topological structure, and the third phoneme can be represented by the third topological structure. Specifically, the first phoneme is a phoneme without a drawn-out feature, that is, a phoneme without a drawn-out pronunciation. The second phoneme is a phoneme with a short drawn-out feature, that is, a phoneme with a short drawn-out duration. The third phoneme is a phoneme with a long drawn-out feature, that is, a phoneme with a long drawn-out duration.

[0059] Among them, the first topological structure includes a preset first number of states with a chronological relationship, the second topological structure includes a preset second number of states with a chronological relationship, and the third topological structure includes a preset third number of states with a chronological relationship. Specifically, the preset third number is greater than the preset second number, and the preset second number is greater than the preset first number.

[0060] Referring to Figure 4 , an embodiment of the present disclosure provides another schematic diagram of the topological structure of a phoneme. Among them, the first topological structure belongs to a three-state structure, the second topological structure belongs to a six-state structure, and the third topological structure belongs to a nine-state structure. The basic duration corresponding to each state is fixed. A larger number of states indicates that the pronunciation duration of the phoneme it represents is longer. Therefore, a topological structure with a larger number of states can be used to represent a phoneme with a longer drawn-out duration.

[0061] In the phoneme sequence annotation method provided by the embodiments of the present disclosure, in the training stage and the application stage of the phoneme sequence annotation model, different topological structures can be used to represent phonemes with and without a drawl feature respectively, and different topological structures can also be used to represent phonemes with a short drawl feature and phonemes with a long drawl feature respectively, which can express the drawl feature of phonemes more precisely. It can be seen that when using the trained phoneme sequence annotation model for phoneme sequence annotation, on the basis of improving the efficiency of phoneme sequence annotation, the drawl feature in phonemes can also be annotated.

[0062] On the basis of the above embodiments, the embodiments of the present disclosure further provide a training method for a phoneme sequence annotation model. In the model training stage, first, audio segment samples and phoneme sequence samples with corresponding relationships are obtained, where the phoneme sequence samples can be obtained by performing phoneme sequence annotation on the audio segment samples, and usually, the phoneme sequence samples are obtained by manually performing phoneme sequence annotation.

[0063] Before using the sample pairs of phoneme segment samples and phoneme sequence samples for model training, they are first preprocessed. Specifically, at least one of the above-mentioned first topological structure, second topological structure, and third topological structure can be used to model the phonemes in the phoneme sequence samples to obtain the modeling results corresponding to the phoneme sequence samples. Then, the model is trained using the audio segment samples and the modeling results corresponding to the phoneme sequence samples to obtain a trained phoneme sequence annotation model for intelligently annotating phoneme sequences.

[0064] In addition, the embodiments of the present disclosure further provide a phoneme sequence annotation device. Refer to Figure 5 , which is a schematic structural diagram of a phoneme sequence annotation device provided by the embodiments of the present disclosure. Specifically, the device includes:

[0065] A first acquisition module 501, configured to acquire a target audio text pair; wherein, the target audio text pair includes a target audio segment and a target text segment with a corresponding relationship;

[0066] An annotation processing module 502, configured to input the target audio text pair into a phoneme sequence annotation model, and after being processed by the phoneme sequence annotation model, obtain a target phoneme sequence of the target audio text pair; wherein, the target phoneme sequence includes phonemes with a time sequence relationship, the first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model, the first topological structure includes a preset first number of states with a time sequence relationship, and the state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence annotation model is trained using audio segment samples and phoneme sequence samples with a corresponding relationship.

[0067] In an alternative embodiment, the second phoneme in the target phoneme sequence is represented by a second topological structure in the phoneme sequence annotation model. The first phoneme is a phoneme without a trailing feature, and the second phoneme is a phoneme with a first trailing feature. The second topological structure includes a preset second number of states with a temporal order relationship, and the preset second number is greater than the preset first number.

[0068] In an alternative embodiment, the third phoneme in the target phoneme sequence is represented by a third topological structure in the phoneme sequence annotation model. The third phoneme is a phoneme with a second trailing feature, the first trailing feature is a short trailing feature, and the second trailing feature is a long trailing feature. The third topological structure includes a preset third number of states with a temporal order relationship, and the preset third number is greater than the preset second number.

[0069] In an alternative embodiment, the apparatus further comprises:

[0070] A second acquisition module, configured to acquire audio segment samples and phoneme sequence samples with corresponding relationships;

[0071] A modeling module, configured to model the phonemes in the phoneme sequence sample by using at least one of the first topological structure, the second topological structure, and the third topological structure to obtain a modeling result corresponding to the phoneme sequence sample;

[0072] A model training module, configured to train a model by using the audio segment sample and the modeling result corresponding to the phoneme sequence sample to obtain the phoneme sequence annotation model.

[0073] In an alternative embodiment, the first topological structure belongs to a three-state structure, and the second topological structure belongs to a six-state structure or a nine-state structure.

[0074] In an alternative embodiment, the first topological structure belongs to a three-state structure, the second topological structure belongs to a six-state structure, and the third topological structure belongs to a nine-state structure.

[0075] In an alternative embodiment, the state is further configured with fields for storing phoneme identifiers and state identifiers.

[0076] In the phoneme sequence annotation device provided by the embodiments of the present disclosure, first, a target audio text pair is obtained. Then, the target audio segment and the target text segment in the target audio text are input into a trained phoneme sequence annotation model. After being processed by the phoneme sequence annotation model, the target phoneme sequence of the target audio text pair is obtained. Among them, the target phoneme sequence includes phonemes with a chronological relationship. The first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model. The first topological structure includes a preset first number of states with a chronological relationship, and each state is configured with a time parameter for characterizing the pronunciation duration. It can be seen that the embodiments of the present disclosure use a trained phoneme sequence annotation model to implement the annotation of the phoneme sequence, which can improve the efficiency of phoneme sequence annotation.

[0077] In addition, the embodiments of the present disclosure can also improve the accuracy of phoneme sequence annotation by combining the audio segment and the text segment to implement the annotation of the phoneme sequence.

[0078] In addition to the above methods and devices, the embodiments of the present disclosure also provide a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device implements the phoneme sequence annotation method described in the embodiments of the present disclosure.

[0079] The embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the phoneme sequence annotation method described in the embodiments of the present disclosure is implemented.

[0080] In addition, the embodiments of the present disclosure also provide a phoneme sequence annotation device. As shown in Figure 6 it may include:

[0081] A processor 601, a memory 602, an input device 603, and an output device 604. The number of processors 601 in the phoneme sequence annotation device may be one or more. Figure 6 Taking one processor as an example. In some embodiments of the present disclosure, the processor 601, the memory 602, the input device 603, and the output device 604 may be connected by a bus or other means. Among them, Figure 6 taking the connection by bus as an example.

[0082] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing of the phoneme sequence annotation device by running the software programs and modules stored in the memory 602. The memory 602 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 603 can be used to receive input digital or character information, and generate signal inputs related to the user settings and function controls of the phoneme sequence annotation device.

[0083] Specifically in this embodiment, the processor 601 will, according to the following instructions, load the executable files corresponding to the processes of one or more application programs into the memory 602, and the processor 601 will run the application programs stored in the memory 602, so as to implement various functions of the above-mentioned phoneme sequence annotation device.

[0084] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0085] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for phoneme sequence annotation, characterized in that, the method includes: obtaining a target audio-text pair; wherein, the target audio-text pair includes a target audio segment and a target text segment with a corresponding relationship; inputting the target audio-text pair into a phoneme sequence annotation model, and after being processed by the phoneme sequence annotation model, obtaining a target phoneme sequence of the target audio-text pair; wherein, the target phoneme sequence includes phonemes with a time sequence relationship, the first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model, the first topological structure includes a preset first number of states with a time sequence relationship, and the state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence annotation model is trained using audio segment samples and phoneme sequence samples with a corresponding relationship.

2. The method according to claim 1, characterized in that, the second phoneme in the target phoneme sequence is represented by a second topological structure in the phoneme sequence annotation model, the first phoneme is a phoneme without a trailing sound feature, the second phoneme is a phoneme with a first trailing sound feature, the second topological structure includes a preset second number of states with a time sequence relationship, and the preset second number is greater than the preset first number.

3. The method according to claim 2, characterized in that, the third phoneme in the target phoneme sequence is represented by a third topological structure in the phoneme sequence annotation model, the third phoneme is a phoneme with a second trailing sound feature, the first trailing sound feature is a short trailing sound feature, the second trailing sound feature is a long trailing sound feature, the third topological structure includes a preset third number of states with a time sequence relationship, and the preset third number is greater than the preset second number.

4. The method according to any one of claims 1-3, characterized in that, before inputting the target audio-text pair into the phoneme sequence annotation model and obtaining the target phoneme sequence corresponding to the target audio-text pair after being processed by the phoneme sequence annotation model, it further includes: obtaining audio segment samples and phoneme sequence samples with a corresponding relationship; modeling the phonemes in the phoneme sequence samples using at least one of the first topological structure, the second topological structure, and the third topological structure to obtain a modeling result corresponding to the phoneme sequence samples; training a model using the audio segment samples and the modeling results corresponding to the phoneme sequence samples to obtain the phoneme sequence annotation model.

5. The method according to claim 2, characterized in that, the first topological structure belongs to a three-state structure, and the second topological structure belongs to a six-state structure or a nine-state structure.

6. The method according to claim 3, characterized in that, the first topological structure belongs to a three-state structure, the second topological structure belongs to a six-state structure, and the third topological structure belongs to a nine-state structure.

7. The method according to claim 1, characterized in that, the state is further configured with fields for storing phoneme identifiers and state identifiers.

8. A phoneme sequence annotation device, characterized in that, the device comprises: a first acquisition module, configured to acquire a target audio-text pair; wherein, the target audio-text pair includes a target audio segment and a target text segment with a corresponding relationship; an annotation processing module, configured to input the target audio-text pair into a phoneme sequence annotation model, and after being processed by the phoneme sequence annotation model, obtain a target phoneme sequence of the target audio-text pair; wherein, the target phoneme sequence includes phonemes with a time sequence relationship, a first phoneme in the target phoneme sequence is represented by a first topological structure in the phoneme sequence annotation model, the first topological structure includes a preset first number of states with a time sequence relationship, and the state is configured with a time parameter for characterizing the pronunciation duration; the phoneme sequence annotation model is trained by using audio segment samples and phoneme sequence samples with a corresponding relationship.

9. A computer-readable storage medium, characterized in that, instructions are stored in the computer-readable storage medium, and when the instructions run on a terminal device, the terminal device is enabled to implement the method according to any one of claims 1-7.

10. A phoneme sequence annotation device, characterized in that, it includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method according to any one of claims 1-7 is implemented.