Timestamp labeling method and device for phonemes, equipment and storage medium

By entering the audio clips and phoneme sequences into the phoneme timestamp labeling model for annotation, and combining the boundary correction model, the problem of low phoneme timestamp labeling in the prior art is solved, and more efficient phoneme timestamp labeling is achieved.

CN120148481APending Publication Date: 2025-06-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311695523.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, phoneme timestamp labeling mainly relies on manual methods, resulting in higher time and labor costs and lower efficiency.

Method used

By determining the target audio clip and its corresponding phoneme sequence, it is input into the phoneme timestamp labeling model, the phoneme timestamp labeling model obtained by the trained phoneme timestamp labeling model, and it can be further corrected through the boundary correction model.

Benefits of technology

It improves the efficiency of phoneme timestamp labeling, reduces the dependence of manual labeling, and reduces time and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148481A_ABST
    Figure CN120148481A_ABST
Patent Text Reader

Abstract

The invention provides a phoneme timestamp labeling method and device, equipment and a storage medium, and the method comprises the steps: firstly, determining a target audio clip, obtaining a phoneme sequence corresponding to the target audio clip, taking the phoneme sequence as a target phoneme sequence, inputting the target audio clip and the target phoneme sequence into a phoneme timestamp labeling model, and obtaining a phoneme timestamp labeling model; after phonemes in the target phoneme sequence are subjected to timestamp labeling processing through a phoneme timestamp labeling model, a phoneme sequence labeling result corresponding to the target audio clip is obtained, and the phoneme timestamp labeling model is obtained by training audio clip samples and labeled phoneme sequence samples which have the corresponding relation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to a method, device, equipment and storage medium for timestamp labeling of phonemes. Background Art

[0002] The timestamp annotation of phonemes refers to the annotation of the pronunciation start and end time of phonemes. Phonemes are the smallest speech units divided according to the natural properties of language. They are analyzed based on the pronunciation actions in syllables. One pronunciation action constitutes one phoneme. For example, the phonemes corresponding to the pronunciation of Mandarin can be represented by the initial consonants and finals in Pinyin. For example, the phonemes corresponding to the pronunciation of the three words "我和你" can be represented by "w, o, h, e, n, i" respectively. The phonemes corresponding to the pronunciation of English can be represented by phonetic symbols, such as / I / , / e / , etc.

[0003] Currently, timestamp annotation of phonemes is mainly achieved manually, which has high time and labor costs, and the efficiency of annotating phoneme timestamps is low. Summary of the invention

[0004] In order to solve the above technical problem, an embodiment of the present disclosure provides a method for timestamp marking of phonemes.

[0005] In a first aspect, the present disclosure provides a method for timestamp labeling of phonemes, the method comprising:

[0006] Determine a target audio segment, and obtain a phoneme sequence corresponding to the target audio segment as a target phoneme sequence;

[0007] The target audio segment and the target phoneme sequence are input into a phoneme timestamp labeling model, and after the phoneme timestamp labeling model performs timestamp labeling processing on the phonemes in the target phoneme sequence, a phoneme sequence labeling result corresponding to the target audio segment is obtained; wherein the phoneme sequence labeling result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp labeling model is trained using audio segment samples and labeled phoneme sequence samples having a corresponding relationship.

[0008] In an optional implementation, after the target audio segment and the target phoneme sequence are input into a phoneme timestamp labeling model, and the phoneme timestamp labeling model performs timestamp labeling processing on the phonemes in the target phoneme sequence to obtain a phoneme sequence labeling result corresponding to the target audio segment, the method further includes:

[0009] Input the phoneme sequence annotation result into a boundary correction model. After the boundary correction model corrects the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, a corrected phoneme sequence annotation result is obtained. Among them, the boundary timestamps include the end timestamp of the previous phoneme in the adjacent phonemes and the start timestamp of the next phoneme in the adjacent phonemes. The boundary correction model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

[0010] In an optional implementation, before inputting the target audio segment and the target phoneme sequence into a phoneme timestamp annotation model and obtaining the phoneme sequence annotation result corresponding to the target audio segment after the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, it further includes:

[0011] Perform multi-task training using pre-acquired sample pairs to obtain a trained phoneme timestamp annotation model. Among them, the sample pairs include audio segment samples and annotated phoneme sequence samples, and the annotated phoneme sequence samples include model-annotated phoneme sequence samples and manually-annotated phoneme sequence samples.

[0012] In an optional implementation, before performing multi-task training using pre-acquired sample pairs to obtain a trained phoneme timestamp annotation model, it further includes:

[0013] Input the first audio segment sample and the first phoneme sequence sample with a corresponding relationship into a preset first model. After the preset first model performs timestamp annotation on the phonemes in the first phoneme sequence sample, obtain the primary timestamp annotation result corresponding to the first audio segment sample.

[0014] Input the primary timestamp annotation result and the first audio segment sample into a preset second model. After the preset second model performs timestamp annotation calibration on the phonemes in the primary timestamp annotation result, obtain the secondary timestamp annotation result corresponding to the first audio segment sample. Among them, the timestamp annotation accuracy of the preset second model is higher than that of the preset first model.

[0015] Determine the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample.

[0016] In an optional implementation, after determining the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample, it further includes:

[0017] Obtain the manually-annotated phoneme sequence sample corresponding to the first audio segment sample.

[0018] Construct sample pairs based on the first audio clip sample, the model-annotated phoneme sequence sample, and the manually-annotated phoneme sequence sample.

[0019] In an alternative implementation, before inputting the phoneme sequence annotation result into the boundary correction model and obtaining the corrected phoneme sequence annotation result after the boundary correction model corrects the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, it further includes:

[0020] Train the boundary correction model using the second audio clip sample and the manually-annotated phoneme sequence sample with a corresponding relationship to obtain a trained boundary correction model.

[0021] In an alternative implementation, obtaining the phoneme sequence corresponding to the target audio clip as the target phoneme sequence includes:

[0022] Input the target audio clip into a phoneme annotation model, and after the processing of the phoneme annotation model, obtain the phoneme sequence corresponding to the target audio clip as the target phoneme sequence.

[0023] In a second aspect, the present disclosure provides a device for timestamp annotation of phonemes, the device includes:

[0024] A first acquisition module, configured to determine a target audio clip and acquire the phoneme sequence corresponding to the target audio clip as the target phoneme sequence;

[0025] A first input module, configured to input the target audio clip and the target phoneme sequence into a phoneme timestamp annotation model, and after the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, obtain the phoneme sequence annotation result corresponding to the target audio clip; wherein, the phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using the audio clip sample and the annotated phoneme sequence sample with a corresponding relationship.

[0026] In a third aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on a terminal device, the terminal device implements the above method.

[0027] In a fourth aspect, the present disclosure provides a device for timestamp annotation of phonemes, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the above method is implemented.

[0028] Fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the above-mentioned method is implemented.

[0029] The technical solutions provided by the embodiments of the present disclosure have at least the following advantages compared with the prior art:

[0030] The embodiments of the present disclosure provide a method for timestamp annotation of phonemes. First, a target audio segment is determined, and a phoneme sequence corresponding to the target audio segment is obtained as a target phoneme sequence. The target audio segment and the target phoneme sequence are input into a phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, a phoneme sequence annotation result corresponding to the target audio segment is obtained. The phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence. The phoneme timestamp annotation model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship. The embodiments of the present disclosure perform timestamp annotation on the phonemes in the target phoneme sequence through the phoneme timestamp annotation model, improving the efficiency of phoneme timestamp annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic structural diagram of a phoneme timestamp annotation system provided by an embodiment of the present disclosure;

[0034] Figure 2 It is a schematic process diagram of a training model provided by an embodiment of the present disclosure;

[0035] Figure 3 It is a schematic diagram of a phoneme timestamp annotation method provided by an embodiment of the present disclosure;

[0036] Figure 4 It is a schematic structural diagram of a phoneme timestamp annotation device provided by an embodiment of the present disclosure;

[0037] Figure 5 It is a schematic structural diagram of a phoneme timestamp annotation device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0039] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0040] A phoneme sequence, including a plurality of phonemes having a chronological relationship. A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing according to the pronunciation actions in a syllable, one pronunciation action constitutes one phoneme, and phonemes are divided into two major categories: vowels and consonants. In Chinese, it generally includes initials and finals in the pronunciation of characters.

[0041] Currently, the timestamp annotation of phonemes is mainly achieved through manual methods, with both high time cost and labor cost, and the efficiency of annotating the timestamps of phonemes is relatively low.

[0042] To this end, the present disclosure provides a method for timestamp annotation of phonemes. Specifically, a target audio segment is determined, and the phoneme sequence corresponding to the target audio segment is obtained as the target phoneme sequence. The target audio segment and the target phoneme sequence are input into a phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, a phoneme sequence annotation result corresponding to the target audio segment is obtained. The phoneme sequence annotation result includes the start timestamp and end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using audio segment samples and annotated phoneme sequence samples with corresponding relationships. Through the phoneme timestamp annotation model of the embodiments of the present disclosure to perform timestamp annotation on the phonemes in the target phoneme sequence, the efficiency of timestamp annotation of phonemes is improved.

[0043] Based on this, the present disclosure provides a method for timestamp annotation of phonemes, referring to Figure 1 , which is a flowchart of a method for timestamp annotation of phonemes provided by the embodiments of the present disclosure, specifically including:

[0044] S101: Determine a target audio segment, and obtain the phoneme sequence corresponding to the target audio segment as the target phoneme sequence.

[0045] In the embodiments of the present disclosure, the target audio segment may be any audio segment. For example, a segment of audio data in a certain song, a segment of audio data in a certain movie, etc.

[0046] In the disclosed embodiment, a phoneme sequence corresponding to a target audio segment is obtained through a target audio segment, wherein a phoneme is the smallest speech unit divided according to the natural properties of a language, and one pronunciation action forms a phoneme. For example, the phonemes corresponding to the pronunciation of Mandarin can be represented by pinyin, and the phonemes corresponding to the pronunciation of the three words "you and I" can be represented by "w, o, h, e, n, i" respectively. The phonemes corresponding to the pronunciation of English can be represented by phonetic symbols, such as / I / , / e / , etc. The phoneme sequence can be a sequence including multiple continuous phonemes. For example, the phoneme sequence can be A = {n, i, sh, i, w, o, x, in, zh, ong, d, e, m, i, an, h, u, a, t, ang}.

[0047] In the embodiment of the present disclosure, the phoneme sequence corresponding to the target audio segment may include a sequence having a time order relationship composed of pronunciation units, ie, phonemes, in the target audio segment. In the embodiment of the present disclosure, the phoneme sequence corresponding to the target audio segment is used as the target phoneme sequence.

[0048] In an optional implementation, a phoneme sequence corresponding to a target audio segment is obtained as a target phoneme sequence. Specifically, this may include identifying text content in the target audio segment by manually marking timestamps, then spelling out the identified text content in pinyin, and phonetically marking the spelling results to obtain a phoneme sequence corresponding to the target audio segment as the target phoneme sequence.

[0049] In practical applications, the model can also be used to obtain the phoneme sequence corresponding to the target audio segment. In an optional implementation, obtaining the phoneme sequence corresponding to the target audio segment as the target phoneme sequence can specifically include inputting the target audio segment into a trained phoneme annotation model, and obtaining the phoneme sequence corresponding to the target audio segment as the target phoneme sequence after processing by the phoneme annotation model.

[0050] Among them, the phoneme annotation model is used to annotate the phonemes in the target audio segment.

[0051] S102: Input the target audio segment and the target phoneme sequence into a phoneme timestamp labeling model, and after the phoneme timestamp labeling model performs timestamp labeling processing on the phonemes in the target phoneme sequence, obtain a phoneme sequence labeling result corresponding to the target audio segment.

[0052] The phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using audio clip samples and annotated phoneme sequence samples having a corresponding relationship.

[0053] In the embodiments of the present disclosure, a phoneme timestamp annotation model is used to perform timestamp annotation on phonemes in a target phoneme sequence. Among them, the phoneme timestamp annotation model can be trained using audio clip samples and annotated phoneme sequence samples with a corresponding relationship.

[0054] Specifically, multi-task training is performed using pre-acquired sample pairs to obtain a trained phoneme timestamp annotation model. Among them, the sample pairs include audio clip samples and annotated phoneme sequence samples, and the annotated phoneme sequence samples include model-annotated phoneme sequence samples and manually annotated phoneme sequence samples.

[0055] Among them, multi-task training means simultaneously learning multiple related tasks in one model.

[0056] In the embodiments of the present disclosure, before multi-task training is performed using pre-acquired sample pairs to obtain a trained phoneme timestamp annotation model, it is necessary to pre-acquire the sample pairs.

[0057] Specifically, a first audio clip sample and a first phoneme sequence sample with a corresponding relationship are input into a preset first model. After the preset first model performs timestamp annotation on the phonemes in the first phoneme sequence sample, a primary timestamp annotation result corresponding to the first audio clip sample is obtained.

[0058] Among them, the first audio clip sample can be any audio clip sample, and the first phoneme sequence sample is the phoneme sequence corresponding to the first audio clip sample. The preset first model is used to perform timestamp annotation on the phonemes in the first phoneme sequence sample. The preset first model can be set based on requirements, and the embodiments of the present disclosure do not impose any restrictions on this. The first phoneme sequence sample can be obtained by the method of manually annotating phonemes or by the method of model-annotating phonemes.

[0059] The primary timestamp annotation result can be understood as frame-level alignment, that is, each frame of audio in the first audio clip sample is aligned with each frame in the first phoneme sequence, and timestamp annotation is performed on the phonemes in the first phoneme sequence sample based on the first audio clip sample. The timestamps annotated for the phonemes can include the start timestamp and end timestamp of the phoneme, that is, the start time and end time of the phoneme pronunciation.

[0060] Then, the primary timestamp annotation result and the first audio clip sample are input into a preset second model. After the preset second model calibrates the timestamp annotation of the phonemes in the primary timestamp annotation result, a secondary timestamp annotation result corresponding to the first audio clip sample is obtained. Among them, the timestamp annotation accuracy of the preset second model is higher than that of the preset first model. The preset second model can be set based on requirements, and the embodiments of the present disclosure do not impose any restrictions on this.

[0061] Among them, the secondary timestamp annotation result can be understood as secondary frame-level alignment, that is, on the basis of the primary timestamp annotation result, timestamp annotation calibration is performed to obtain a secondary timestamp annotation result with higher timestamp annotation accuracy.

[0062] Determine the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample.

[0063] After obtaining the model-annotated phoneme sequence sample of the first audio segment sample, it is also necessary to obtain the manually annotated phoneme sequence sample corresponding to the first audio segment sample. Then, a sample pair is constructed based on the first audio segment sample, the model-annotated phoneme sequence sample, and the manually annotated phoneme sequence sample.

[0064] Among them, to obtain the manually annotated phoneme sequence sample corresponding to the first audio segment sample, professionals can annotate the phonemes in the phoneme sequence of the first audio segment sample to obtain the manually annotated phoneme sequence sample.

[0065] The phoneme timestamp annotation method provided by the embodiments of the present disclosure. Specifically, determine the target audio segment, and obtain the phoneme sequence corresponding to the target audio segment as the target phoneme sequence. Input the target audio segment and the target phoneme sequence into the phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, the phoneme sequence annotation result corresponding to the target audio segment is obtained. Among them, the phoneme sequence annotation result includes the start timestamp and end timestamp of the phonemes in the target phoneme sequence. The phoneme timestamp annotation model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship. The embodiments of the present disclosure perform timestamp annotation on the phonemes in the target phoneme sequence through the phoneme timestamp annotation model, improving the efficiency of phoneme timestamp annotation.

[0066] In practical applications, after obtaining the phoneme sequence annotation result corresponding to the target audio segment, in order to improve the accuracy of the phoneme sequence annotation result, the phoneme sequence annotation result can also be corrected.

[0067] Specifically, input the phoneme sequence annotation result into the boundary correction model. After the boundary correction model performs correction processing on the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, the corrected phoneme sequence annotation result is obtained.

[0068] Among them, the boundary timestamps include the end timestamp of the previous phoneme in adjacent phonemes and the start timestamp of the next phoneme in adjacent phonemes. The boundary correction model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

[0069] For example, assume that the end timestamp of the previous phoneme in adjacent phonemes in the phoneme sequence annotation result is 130 ms, and the start timestamp of the subsequent phoneme is also 130 ms. After inputting the phoneme sequence annotation result into the boundary correction model, the boundary timestamps of the adjacent phonemes are corrected to obtain the corrected phoneme sequence annotation result. The end timestamp of the previous phoneme in the adjacent phonemes is 130 ms, and the start timestamp of the subsequent phoneme is 140 ms.

[0070] The boundary correction model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

[0071] Specifically, the boundary modification model can be trained using the second audio segment sample and the manually annotated phoneme sequence sample with a corresponding relationship to obtain the trained boundary correction model. Among them, the second audio segment sample can be any audio segment sample, and the manually annotated phoneme sequence sample is obtained by a professional annotating the timestamps of the phonemes in the phoneme sequence of the second audio segment sample.

[0072] In the embodiments of the present disclosure, the first audio segment sample and the second audio segment sample can be the same audio segment sample or different audio segment samples.

[0073] By correcting the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result to obtain the corrected phoneme sequence annotation result, the accuracy of the phoneme annotation result is further improved.

[0074] To facilitate the understanding of the above embodiments, the embodiments of the present disclosure provide a schematic diagram of a process for training a model, as Figure 2 shown. Among them, the first audio segment sample and the second audio segment sample are the same audio segment sample.

[0075] First, the first audio segment sample and the first phoneme sequence sample with a corresponding relationship are input into the preset first model 201. The preset first model annotates the timestamps of the phonemes in the first phoneme sequence sample to obtain the primary timestamp annotation result corresponding to the first audio segment sample. The primary timestamp annotation result and the first audio segment sample are input into the preset second model 202. The preset second model calibrates the timestamps of the phonemes in the primary timestamp annotation result to obtain the secondary timestamp annotation result corresponding to the first audio sample. The secondary timestamp annotation result is determined as the model-annotated phoneme sequence sample of the first audio segment sample.

[0076] Then, the manually annotated phoneme sequence sample corresponding to the first audio segment sample is obtained, and a sample pair is constructed based on the first audio segment sample, the model-annotated phoneme sequence sample, and the manually annotated phoneme sequence sample.

[0077] Perform multi-task training on the sample pairs to obtain a trained phoneme timestamp annotation model 203.

[0078] Then, use the first audio clip sample (second audio clip) with a corresponding relationship and the manually annotated phoneme sequence sample to train the boundary correction model, and obtain a trained boundary correction model 204.

[0079] Based on the above embodiments, the embodiments of the present disclosure also provide a schematic diagram of a method for timestamp annotation of phonemes, as Figure 3 shown.

[0080] First, the target audio clip and the target phoneme sequence are input into the phoneme timestamp annotation model 301. The phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence to obtain a phoneme sequence annotation result corresponding to the target audio clip. Then, the phoneme sequence annotation result is used as an input parameter and input into the boundary correction model 302. The boundary correction model corrects the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result to obtain a corrected phoneme sequence annotation result, which is used as the final phoneme sequence annotation result.

[0081] Based on this, the present disclosure also provides a device for timestamp annotation of phonemes. Refer to Figure 4 , which is a schematic structural diagram of a device for timestamp annotation of phonemes provided by the embodiments of the present disclosure. Specifically, the device includes:

[0082] A first acquisition module 401, configured to determine a target audio clip and acquire a phoneme sequence corresponding to the target audio clip as a target phoneme sequence;

[0083] A first input module 402, configured to input the target audio clip and the target phoneme sequence into the phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, a phoneme sequence annotation result corresponding to the target audio clip is obtained; wherein, the phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using audio clip samples and annotated phoneme sequence samples with a corresponding relationship.

[0084] In an optional implementation manner, the device further includes:

[0085] A second input module, configured to input the phoneme sequence annotation result into a boundary correction model. After the boundary correction model corrects the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, a corrected phoneme sequence annotation result is obtained. Wherein, the boundary timestamps include the end timestamp of the previous phoneme in the adjacent phonemes and the start timestamp of the next phoneme in the adjacent phonemes, and the boundary correction model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

[0086] In an alternative embodiment, the apparatus further includes:

[0087] A first training module, configured to perform multi-task training using pre-acquired sample pairs to obtain a trained phoneme timestamp annotation model. Wherein, the sample pairs include audio segment samples and annotated phoneme sequence samples, and the annotated phoneme sequence samples include model-annotated phoneme sequence samples and manually-annotated phoneme sequence samples.

[0088] In an alternative embodiment, the apparatus further includes:

[0089] A third input module, configured to input a first audio segment sample and a first phoneme sequence sample with a corresponding relationship into a preset first model. After the preset first model performs timestamp annotation on the phonemes in the first phoneme sequence sample, a primary timestamp annotation result corresponding to the first audio segment sample is obtained;

[0090] A fourth input module, configured to input the primary timestamp annotation result and the first audio segment sample into a preset second model. After the preset second model calibrates the timestamp annotation of the phonemes in the primary timestamp annotation result, a secondary timestamp annotation result corresponding to the first audio segment sample is obtained. Wherein, the timestamp annotation accuracy of the preset second model is higher than that of the preset first model;

[0091] A determination module, configured to determine the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample.

[0092] In an alternative embodiment, the apparatus further includes:

[0093] A second acquisition module, configured to acquire a manually-annotated phoneme sequence sample corresponding to the first audio segment sample;

[0094] A construction module, configured to construct sample pairs based on the first audio segment sample, the model-annotated phoneme sequence sample, and the manually-annotated phoneme sequence sample.

[0095] In an alternative embodiment, the apparatus further includes:

[0096] A second training module, configured to train a boundary correction model by using a second audio clip sample and an artificially annotated phoneme sequence sample with a corresponding relationship, so as to obtain a trained boundary correction model.

[0097] In an alternative implementation, the first acquisition module includes:

[0098] Input the target audio clip into a phoneme annotation model, and after being processed by the phoneme annotation model, obtain the phoneme sequence corresponding to the target audio clip as the target phoneme sequence.

[0099] In the phoneme timestamp annotation device provided by the embodiments of the present disclosure, specifically, a target audio clip is determined, and the phoneme sequence corresponding to the target audio clip is obtained as the target phoneme sequence. The target audio clip and the target phoneme sequence are input into a phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, a phoneme sequence annotation result corresponding to the target audio clip is obtained. The phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence. The phoneme timestamp annotation model is trained by using an audio clip sample and an annotated phoneme sequence sample with a corresponding relationship. The embodiments of the present disclosure perform timestamp annotation on the phonemes in the target phoneme sequence through the phoneme timestamp annotation model, improving the efficiency of phoneme timestamp annotation.

[0100] In addition to the above methods and devices, the embodiments of the present disclosure further provide a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device implements the phoneme timestamp annotation method described in the embodiments of the present disclosure.

[0101] The embodiments of the present disclosure further provide a computer program product. The computer program product includes computer programs / instructions. When the computer programs / instructions are executed by a processor, the phoneme timestamp annotation method described in the embodiments of the present disclosure is implemented.

[0102] In addition, the embodiments of the present disclosure further provide a phoneme timestamp annotation device. As shown in Figure 5 it may include:

[0103] A processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 in the phoneme timestamp annotation device may be one or more. Figure 5 Taking one processor as an example. In some embodiments of the present disclosure, the processor 501, the memory 502, the input device 503, and the output device 504 may be connected through a bus or other means. Among them, Figure 5 taking the connection through a bus as an example.

[0104] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing of the phoneme timestamp annotation device by running the software programs and modules stored in the memory 502. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory 502 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 503 can be used to receive input digital or character information and generate signal inputs related to the user settings and function controls of the phoneme timestamp annotation device.

[0105] Specifically, in this embodiment, the processor 501 loads the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 runs the application programs stored in the memory 502 to implement various functions of the above-mentioned phoneme timestamp annotation device.

[0106] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0107] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for timestamp annotation of phonemes, characterized in that, the method includes: determining a target audio segment, and obtaining a phoneme sequence corresponding to the target audio segment as a target phoneme sequence; inputting the target audio segment and the target phoneme sequence into a phoneme timestamp annotation model, and after the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, obtaining a phoneme sequence annotation result corresponding to the target audio segment; wherein, the phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

2. The method according to claim 1, characterized in that, after inputting the target audio segment and the target phoneme sequence into the phoneme timestamp annotation model, and after the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence to obtain a phoneme sequence annotation result corresponding to the target audio segment, it further includes: inputting the phoneme sequence annotation result into a boundary correction model, and after the boundary correction model performs correction processing on the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, obtaining a corrected phoneme sequence annotation result; wherein, the boundary timestamps include the end timestamp of the previous phoneme in the adjacent phonemes and the start timestamp of the next phoneme in the adjacent phonemes, and the boundary correction model is trained using audio segment samples and annotated phoneme sequence samples with a corresponding relationship.

3. The method according to claim 1, characterized in that, before inputting the target audio segment and the target phoneme sequence into the phoneme timestamp annotation model, and after the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence to obtain a phoneme sequence annotation result corresponding to the target audio segment, it further includes: performing multi-task training using pre-obtained sample pairs to obtain a trained phoneme timestamp annotation model; wherein, the sample pairs include audio segment samples and annotated phoneme sequence samples, and the annotated phoneme sequence samples include model-annotated phoneme sequence samples and manually-annotated phoneme sequence samples.

4. The method according to claim 3, characterized in that, before performing multi-task training using pre-obtained sample pairs to obtain a trained phoneme timestamp annotation model, it further includes: inputting a first audio segment sample and a first phoneme sequence sample with a corresponding relationship into a preset first model, and after the preset first model performs timestamp annotation on the phonemes in the first phoneme sequence sample, obtaining a primary timestamp annotation result corresponding to the first audio segment sample; Input the primary timestamp annotation result and the first audio segment sample into a preset second model. After the preset second model calibrates the timestamps of the phonemes in the primary timestamp annotation result, obtain the secondary timestamp annotation result corresponding to the first audio segment sample; wherein, the timestamp annotation accuracy of the preset second model is higher than that of the preset first model. Determine the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample.

5. The method according to claim 4, wherein, after determining the secondary timestamp annotation result as the model-annotated phoneme sequence sample of the first audio segment sample, further include: Obtain the manually annotated phoneme sequence sample corresponding to the first audio segment sample; Construct a sample pair based on the first audio segment sample, the model-annotated phoneme sequence sample, and the manually annotated phoneme sequence sample.

6. The method according to claim 2, wherein, before inputting the phoneme sequence annotation result into the boundary correction model and obtaining the corrected phoneme sequence annotation result after the boundary correction model corrects the boundary timestamps of adjacent phonemes in the phoneme sequence annotation result, further include: Use the second audio segment sample and the manually annotated phoneme sequence sample with a corresponding relationship to train the boundary correction model to obtain a trained boundary correction model.

7. The method according to claim 1, wherein, the obtaining the phoneme sequence corresponding to the target audio segment as the target phoneme sequence includes: Input the target audio segment into a phoneme annotation model. After being processed by the phoneme annotation model, obtain the phoneme sequence corresponding to the target audio segment as the target phoneme sequence.

8. A device for timestamp annotation of phonemes, wherein, the device includes: A first acquisition module, configured to determine a target audio segment and acquire the phoneme sequence corresponding to the target audio segment as the target phoneme sequence; A first input module, configured to input the target audio segment and the target phoneme sequence into a phoneme timestamp annotation model. After the phoneme timestamp annotation model performs timestamp annotation processing on the phonemes in the target phoneme sequence, obtain the phoneme sequence annotation result corresponding to the target audio segment; wherein, the phoneme sequence annotation result includes the start timestamp and the end timestamp of the phonemes in the target phoneme sequence, and the phoneme timestamp annotation model is trained using an audio segment sample and an annotated phoneme sequence sample with a corresponding relationship.

9. A computer-readable storage medium, wherein, instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device implements the method according to any one of claims 1-7.

10. A device for timestamp annotation of phonemes, wherein, includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1-7 is implemented.