Two-level speech alignment method, electronic device, and storage medium

Through the two-level speech alignment method, combined with word-level and phoneme-level alignment models, the problems of misalignment and tailing of speech alignment results in the prior art are solved, and the high accuracy correspondence between speech and transcripts is achieved.

CN114882904BActive Publication Date: 2025-05-06BEIJING INTENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210274508.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-05-06
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

In the prior art, due to the high model sensitivity of the phoneme-level alignment model, the speech alignment results may be misaligned at the beginning position or the end position will be tailed, resulting in incomplete correspondence between the speech and the transcript.

Method used

The two-level speech alignment method is adopted, and the initial word-level alignment result is first obtained through the pre-constructed word-level alignment model, and the final word-level alignment transcript is generated by combining adjacent frame data and placeholder processing. Then, the word-level alignment results are input into the phoneme-level alignment model, and the final phoneme-level alignment transcript is generated by combining phoneme-level frame data and placeholders.

Benefits of technology

Through the two-level alignment method, the misalignment and tailing of speech alignment results are effectively avoided, the correspondence between speech and transcript is ensured, and errors are minimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882904B_ABST
    Figure CN114882904B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention relates to a two-level speech alignment method, electronic device, and storage medium, including: obtaining speech data and a word-level transcript; inputting the two into a word-level alignment model to obtain an initial word-level alignment result; traversing each frame of word-level frame data, merging adjacent frames belonging to the same word, obtaining first merged frame data, and recording its start time and end time; merging the word-level frame data corresponding to the first type of placeholder and the word adjacent to the first type of placeholder to obtain second merged frame data; updating its start time and end time to obtain a word-level alignment transcript; obtaining a first spectral feature sequence and an information vector; inputting the two and the word-level alignment transcript into a phoneme-level alignment model to obtain an initial phoneme-level alignment result; traversing each frame of phoneme-level frame data, merging the second type of placeholder with the phoneme-level frame data corresponding to the phoneme unit, and obtaining a phoneme-level alignment transcript. In the above, no error will occur at the beginning or end of the transcript.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a two-level speech alignment method, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, related technologies in the field of speech are becoming more and more mature, and speech technology applications are becoming more and more extensive. As an important part of speech processing, speech alignment technology can provide important support for speech recognition, speech synthesis, speech evaluation and other technologies. The accuracy of speech alignment boundaries determines the accuracy of speech recognition and speech evaluation results, as well as the quality of speech synthesis audio.

[0003] In current technology, the phoneme-level alignment model has a high sensitivity, which may cause misalignment at the beginning or tailing at the end of the speech alignment result. As a result, the final transcript corresponding to the speech may not completely correspond to the speech itself and may contain certain errors. Summary of the invention

[0004] The present application provides a two-level speech alignment method, an electronic device, and a storage medium to solve some or all of the above-mentioned technical problems in the prior art.

[0005] In a first aspect, the present application provides a two-level speech alignment method, the method comprising:

[0006] obtaining speech data, and word-level transcripts corresponding to the speech data;

[0007] Inputting the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model to obtain an initial word-level alignment result corresponding to the speech data, wherein the initial word-level alignment result includes a plurality of frames of word-level frame data, and each frame of word-level frame data corresponds to a word in the word-level transcript, or corresponds to a first-category placeholder, and a word in the word-level transcript corresponds to at least one frame of word-level frame data;

[0008] Traverse each frame of word-level frame data, merge adjacent frames belonging to the same word, obtain first merged frame data, and record the start time and end time of the first merged frame data;

[0009] According to a first preset rule, the first type of placeholder is merged with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder to obtain second merged frame data;

[0010] Updating the start time and end time of the second merged frame data, and finally obtaining a word-level aligned transcript corresponding to the speech data;

[0011] Acquire a first frequency spectrum feature sequence corresponding to the speech data and an information vector corresponding to the speech data;

[0012] Inputting the word-level aligned transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model to obtain an initial phoneme-level alignment result corresponding to the speech data, wherein the initial phoneme-level alignment result includes multiple frames of phoneme-level frame data, each frame of the phoneme-level frame data corresponds to a phoneme unit, or corresponds to a second-category placeholder;

[0013] Traversing each frame of phoneme-level frame data, merging at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data;

[0014] The start time and the end time of the third merged frame data are recorded as the start time and the end time corresponding to the phoneme unit, and finally a phoneme-level aligned transcript corresponding to the speech data is obtained.

[0015] Optionally, according to a first preset rule, merging the first type of placeholder with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder to obtain second merged frame data specifically includes:

[0016] When the left side of the first type of placeholder is word-level frame data corresponding to the word, the first type of placeholder is merged with the word-level frame data corresponding to the word on the left side, and the end time of the first type of placeholder is used as the end time of the second merged frame, and the start time of the word-level frame data with the word on the left side is used as the start time of the second merged frame;

[0017] Alternatively, when the first type of placeholder is word-level frame data corresponding to the word on the right side, the first type of placeholder is merged with the word-level frame data corresponding to the word on the right side, and the start time of the first type of placeholder is used as the start time of the second merged frame, and the end time of the word-level frame data corresponding to the word on the right side is used as the end time of the second merged frame.

[0018] In a second aspect, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;

[0019] Memory, used to store computer programs;

[0020] The processor is used to implement the steps of the two-level speech alignment method of any embodiment of the first aspect when executing the program stored in the memory.

[0021] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the two-level speech alignment method as in any embodiment of the first aspect are implemented.

[0022] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0023] The method provided in the embodiment of the present application is for speech data, and the word-level transcript corresponding to the speech data. Then the speech data and the word-level transcript corresponding to the speech data are input into a pre-built word-level alignment model, and then the initial word-level alignment result is obtained. Each frame of word-level frame data in the initial word-level alignment result is traversed, and the frame data corresponding to the word in the word-level frame data and the first type of placeholder are merged and processed respectively. The placeholder is used to represent non-speech data. The time of the placeholder in the speech data is allocated to the adjacent words to ensure that each word of speech in the speech data can correspond to the word in the transcript, so as to achieve word-level transcript alignment.

[0024] On this basis, each character in the character-level transcript is used as a phoneme transcript, and the first spectral feature sequence corresponding to the speech data and the information vector corresponding to the speech data are input into the pre-built phoneme-level alignment model to obtain the initial phoneme-level alignment result. Then, each frame of phoneme-level frame data is traversed, and a merging process of the phoneme-level frame data is performed to merge the second type of placeholder with the frame data corresponding to the adjacent phoneme to obtain a merged frame. Among them, the second type of placeholder is also a phoneme-level placeholder corresponding to non-speech data. The start time and end time of the merged frame are used as the start time and end time of the phoneme. In this way, phoneme-level transcript alignment is achieved. As mentioned above, because the time corresponding to all non-speech data has been reasonably allocated to the phoneme-level frame data corresponding to the adjacent phoneme-level transcripts, that is, the time corresponding to the non-speech data is allocated to the adjacent phonemes, the obtained phoneme-level transcripts will not be misplaced at the beginning position of the speech alignment result or have tailing at the end position. This can ensure that the final phoneme-level transcript can basically correspond to the speech itself, thus eliminating errors as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of a two-stage speech alignment method provided by an embodiment of the present invention;

[0026] Figure 2 A schematic diagram of the effect of word-level alignment transcript provided by an embodiment of the present invention;

[0027] Figure 3 A schematic flow chart of a third method for merging frame data provided by an embodiment of the present invention;

[0028] Figure 4 A schematic diagram of the effect of obtaining a phoneme-level aligned transcript provided by an embodiment of the present invention;

[0029] Figure 5A schematic diagram of the overall process of the method for constructing a character-level alignment model provided by the present invention;

[0030] Figure 6 A schematic diagram of the overall process of the method for constructing a phoneme-level alignment model provided by the present invention;

[0031] Figure 7 A schematic diagram of the structure of a two-stage speech alignment device provided by an embodiment of the present invention;

[0032] Figure 8 A schematic diagram of the structure of an electronic device is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] To facilitate understanding of the embodiments of the present invention, specific embodiments will be further explained below in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the embodiments of the present invention.

[0035] In view of the technical problems mentioned in the background technology, the present application embodiment provides a two-level speech alignment method, see Figure 1 As shown, Figure 1 A schematic flow chart of a two-level speech alignment method provided by an embodiment of the present invention, the method comprising the steps of:

[0036] Step 110, obtaining speech data and a word-level transcript corresponding to the speech data.

[0037] Specifically, the voice data is audio data collected by a sound pickup device. The word-level transcript can be a pre-set or post-annotated transcript. The pre-set may be, for example, a book or a line of dialogue, or any transcript. Then, find a corresponding person to record the transcript of the book or line of dialogue to obtain the voice data.

[0038] Step 120 , input the speech data and the character-level transcript corresponding to the speech data into a pre-built character-level alignment model to obtain an initial character-level alignment result corresponding to the speech data.

[0039] Specifically, the word-level alignment model is a pre-built model, for example, it can be a pre-built end-to-end neural network model. The construction process of the word-level alignment model will be described in detail below, and no further explanation is given here.

[0040] After passing through the word-level alignment model, the initial word-level alignment result obtained may include multiple frames of word-level frame data, and each frame of word-level frame data corresponds to a word in the word-level transcript, or corresponds to a first-class placeholder, and a word in the word-level transcript corresponds to at least one frame of word-level frame data.

[0041] For example, the speech data is "non-speech data + hello". Non-speech data, such as noise, corresponds to the character-level transcript with the 1st to 5th frame data as noise, and the noise data is reflected in the form of the first type of placeholder. The word "you" corresponds to the 6th to 7th frame data, and the word "good" corresponds to the 8th to 9th frame data. In order to distinguish it from the phoneme-level frame data below, the frame data output by the character-level alignment model is defined as character-level frame data. In essence, it is just one frame of data. In addition, in the initial character-level alignment result, a character-level placeholder can be inserted at the beginning, end, and between characters of the transcript. The reason for adding character-level placeholders at these positions is that it is considered that in the initial character-level alignment results that may be obtained, not every frame of data can correspond to the corresponding character, and some places may not correspond to the pronunciation position, so it needs to be reflected in the form of a character-level placeholder.

[0042] Step 130 , traverse each frame of word-level frame data, merge adjacent frames belonging to the same word, obtain first merged frame data, and record the start time and end time of the first merged frame data.

[0043] Step 140: According to a first preset rule, the first type of placeholder is merged with the character-level frame data corresponding to the character immediately adjacent to the first type of placeholder to obtain second merged frame data.

[0044] Step 150, updating the start time and the end time of the second merged frame data, and finally obtaining a word-level aligned transcript corresponding to the speech data.

[0045] Specifically, as described above, the initial word-level alignment result includes multiple frames of word-level frame data. Each frame of word-level frame data will be mapped to a word or a word-level placeholder. According to the decoding rules of CTC (Connectionist Temporal Classification), multiple frames of input correspond to one output. For example, the word "you" corresponds to the 6th frame data and the 7th frame data. Therefore, it is necessary to traverse each frame of word-level frame data, merge adjacent frames belonging to the same word, and obtain the first merged frame data.

[0046] At the same time, it is also necessary to merge the first type of placeholders with the word-level frame data corresponding to the words immediately adjacent to the first type of placeholders according to the first preset rule to obtain second merged frame data.

[0047] Specifically, according to the first preset rule, the first type of placeholder is merged with the character-level frame data corresponding to the character adjacent to the first type of placeholder. For the implementation of obtaining the second merged frame data, see the following:

[0048] When the left side of the first type of placeholder is the character-level frame data corresponding to a character, the first type of placeholder is merged with the character-level frame data corresponding to the character on the left side, and the end time of the first type of placeholder is used as the end time of the second merged frame, and the start time of the character-level frame data of the character on the left side is the start time of the second merged frame;

[0049] Or, when the first type of placeholder is the character-level frame data corresponding to a character on the right side, the first type of placeholder is merged with the character-level frame data corresponding to the character on the right side, and the start time of the first type of placeholder is used as the start time of the second merged frame, and the end time of the character-level frame data corresponding to the character on the right side is used as the end time of the second merged frame.

[0050] For example, as mentioned above, frames 1 to 5 are all the first type of placeholders, and frame 6 is the character "你" (assuming that the two-frame character-level frame data corresponding to the current character "你" has been merged to form the first merged frame, and the first merged frame is currently frame 6 data).

[0051] In an optional embodiment, the first type of placeholder in frame 5 can be first merged with the character-level frame data corresponding to the character "你" in the first merged frame. Then, the first type of placeholder in frame 4 is merged with the current character-level frame data corresponding to the character "你", and so on, until all the first type of placeholders are merged with the character "你".

[0052] In the second optional embodiment, the character-level placeholders corresponding to the first 5 frames can be merged into one frame of character-level frame data, and then the merged character-level frame data is merged with the character-level frame data corresponding to the character "你" in the first merged frame.

[0053] In the third optional example, assume that frames 6 and 7 in the initial character-level alignment result have not been merged, that is, the character "你" still corresponds to two frames of character-level frame data. The first 5 frames are all the first type of placeholders.

[0054] Then, the first 5 frames of the first type of placeholders can be merged using the merging method of the first type of placeholders introduced in either the first or the second method introduced above, and then merged with the character-level frame data corresponding to the character "你" in frame 6, and then the merged frame is merged with the character-level frame data in frame 7 to obtain the final character-level frame data corresponding to the character "你".

[0055] Considering that there is still the first type of placeholder between the character-level frame data corresponding to the character "你" and the character-level frame data corresponding to the character "好".

[0056] Then, this first type of placeholder can be set to be merged with the character-level frame data corresponding to the "character" on the left, or it can be set to be merged with the character-level frame data corresponding to the "character" on the right. The specific merging method can be set according to the actual situation. Similarly, assume that the character "好" also has a first type of placeholder, and there are no other characters after the character "好", so the first type of placeholder after the character "好" is merged with the character-level frame data corresponding to the character "好".

[0057] In addition, when merging the character-level frame data, it is also necessary to determine the start time and end time of the merged character-level frame data. For example, the start time corresponding to the character "你" mentioned above is the start time corresponding to the first frame of character-level frame data. And the end time is the end time of the last frame of character-level frame data corresponding to the character "你" (assuming here that the first type of placeholder after the character-level frame data corresponding to the character "你" is merged with the character-level frame data corresponding to the character "好").

[0058] It should be noted that the execution order of step 140 and step 150 is not fixed. Instead, it can be freely switched according to the actual situation. For how to switch specifically, refer to the several situations introduced above.

[0059] In short, through this method, after merging all the character-level frame data in the above manner, a character-level alignment transcript corresponding to the speech data can be obtained.

[0060] Figure 2 Shows the character-level alignment transcript obtained according to the above method. For details, refer to Figure 2 as shown. Figure 2 Take "noise + refrigeration mode + noise" as an example for illustration. Therefore, the first type of character-level placeholder corresponding to the first half of the noise is merged with the character-level frame data corresponding to the character "制" to obtain the start time and end time of the character "制". For the rest, refer to the above introduction and will not be elaborated here.

[0061] Step 160, obtain the first spectral feature sequence corresponding to the speech data and the information vector corresponding to the speech data.

[0062] Specifically, after obtaining the speech data, the speech data can be subjected to feature extraction to obtain the first spectral feature sequence. The specific extraction process is prior art. Briefly speaking, it is to first extract the time-frequency signal corresponding to the speech signal and convert it into a frequency-domain signal, and then perform feature extraction on the frequency-domain signal to obtain the first spectral feature. This will not be elaborated here in detail. And the information vector corresponding to the speech data is also obtained through prior art and will not be elaborated here too much.

[0063] Step 170, input the character-level alignment transcript, the first spectral feature sequence, and the information vector into the pre-constructed phoneme-level alignment model to obtain the initial phoneme-level alignment result corresponding to the speech data.

[0064] Specifically, similar to the initial word-level alignment result, the initial phoneme-level alignment result includes multiple frames of phoneme-level frame data, each frame of phoneme-level frame data corresponds to a phoneme unit, or corresponds to a second-class placeholder. The specific construction process of the pre-built phoneme-level alignment model is also described in detail below, and will not be described in detail here. At present, only its function is introduced. The phoneme-level alignment model can be, for example, a Gaussian mixture model.

[0065] Step 180, traverse each frame of phoneme-level frame data, merge at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit, and obtain third merged frame data.

[0066] Step 190 , recording the start time and end time of the third merged frame data as the start time and end time corresponding to the phoneme unit, and finally obtaining a phoneme-level aligned transcript corresponding to the speech data.

[0067] Specifically, in an optional example, each frame of phoneme-level frame data is traversed, and at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result is merged with the phoneme-level frame data corresponding to the phoneme unit. The specific implementation process of obtaining the third merged frame data can be referred to as Figure 3 As shown, the method steps include:

[0068] Step 310: When the phoneme-level frame data is adjacent to at least one second-category placeholder, merge the at least one second-category placeholder to obtain fourth merged frame data, and determine the start time and end time of the fourth merged frame data.

[0069] Step 320: merge the phoneme-level frame data with the fourth merged frame data to obtain third merged frame data.

[0070] Specifically, the second type of placeholder is actually a common placeholder similar to the first type of placeholder, and its function is to distinguish the phoneme-level (or word-level) frame data corresponding to the voice data and the non-voice data.

[0071] In an optional example, the second type of placeholders may include phoneme-level placeholders and / or silence markers.

[0072] Then, when the phoneme-level frame data is adjacent to at least one second-category placeholder, at least one second-category placeholder is merged to obtain fourth merged frame data, which may specifically include the following situations.

[0073] First, when the second type of placeholders only includes phoneme-level placeholders, each phoneme-level placeholder is converted into a silence mark, and then the silence marks are merged to obtain fourth merged frame data.

[0074] Second, when the second type of placeholders includes phoneme-level placeholders and silence placeholders in sequence (two placeholders are adjacent), the phoneme-level placeholders are merged with the silence placeholders adjacent thereto, and then all silence placeholders are merged to obtain fourth merged frame data.

[0075] Further optionally, recording the start time and the end time of the third merged frame data as the start time and the end time corresponding to the phoneme unit specifically includes the following method steps:

[0076] When the fourth merged frame data is on the left side of the phoneme-level frame data, the start time of the fourth merged frame data is used as the start time of the third merged frame data, and the end time of the phoneme-level frame data is used as the end time of the third merged frame data;

[0077] Alternatively, when the fourth merged frame data is on the right side of the phoneme level frame data, the start time of the phoneme level frame data is used as the start time of the third merged frame data, and the end time of the fourth merged frame data is used as the end time of the third merged frame data.

[0078] In the following examples, the second type of placeholders including phoneme-level placeholders and silence marks are taken as an example for explanation.

[0079] In a specific example, assuming that when the second category of placeholders includes phoneme-level placeholders and silence marks, and the combination order is phoneme-level placeholders and silence marks, it is necessary to first convert the phoneme-level placeholder into a silence mark, and then merge it with the right adjacent silence into a silence mark, and merge the boundary time at the same time.

[0080] In another specific example, assuming that when the second category of placeholders includes phoneme-level placeholders and silence marks, and the combination order is silence marks and phoneme-level placeholders, it is necessary to first convert the phoneme-level placeholder into a silence mark, and then merge it with the left adjacent silence into a silence mark, and merge the boundary time at the same time.

[0081] In another specific example, assuming that the second category of placeholders includes phoneme-level placeholders and silence marks, and the combination order is silence mark, phoneme-level placeholder and silence mark, it is necessary to first convert the phoneme-level placeholder into a silence mark, and then merge it with the left adjacent silence mark and the right adjacent silence mark into one silence mark, and merge the boundary time at the same time.

[0082] In the above examples, although the steps involve converting phoneme-level placeholders into silence marks. However, in the actual process, this operation is not actually necessary. For example, if the second type of placeholders only includes phoneme-level placeholders, then the phoneme-level placeholders can be directly merged, and then the phoneme-level frame data corresponding to the adjacent phoneme units can be merged (including time merging) without converting to silence marks. However, when there are both phoneme-level placeholders and silence marks, it is best to unify them into the same type, and then merge them with the phoneme-level frame data corresponding to the phoneme unit. From the perspective of the calculation process, it will be more efficient.

[0083] Of course, it is not necessary to convert the phoneme-level placeholders into silence marks, and vice versa. In short, the purpose of the above operation is to improve the efficiency of merging frame data. Therefore, whether the phoneme-level placeholders are converted into silence marks first, and then all silence marks are merged, or the phoneme-level placeholders are directly merged with the phoneme-level frame data, and then the phoneme-level frame data is merged with the silence marks, or merged in sequence, for example, the phoneme-level aligned transcript includes: phoneme-level placeholders, phoneme-level placeholders, phoneme-level frame data, silence marks, and phoneme-level placeholders, then the first two frames of phoneme-level placeholders are merged, then the phoneme-level frame data is merged, and then merged with the silence marks, and finally merged with the phoneme-level placeholders.

[0084] The above and other aspects are all within the scope of implementation of this embodiment, and specific implementation methods will not be described one by one by examples here.

[0085] Figure 4 The effect diagram of obtaining the phoneme-level aligned transcript by using the above method steps of the embodiment of the present application is shown. Figure 4 As shown, Figure 4 In the same example, "non-speech data + cooling mode + non-speech data" is used as speech data. After following the above steps, the effect diagram is obtained. It can be seen from the whole process that phoneme-level alignment can be fully achieved.

[0086] Further optionally, after obtaining the phoneme-level aligned transcript corresponding to the speech data, the method may further include:

[0087] The word-level aligned transcripts corresponding to the speech data are optimized using the phoneme-level aligned transcripts corresponding to the speech data.

[0088] Specifically, after obtaining the phoneme-level aligned transcript, in fact, for the voice data, the frame data corresponding to the pinyin of each word in the voice data is composed of frame data corresponding to the phoneme. It is equivalent to completing the alignment operation at a finer granularity. Then, using the phoneme combination to complete the alignment operation of the pinyin will be more accurate, which is to complete the optimization operation of the word-level aligned transcript. The start time of the first phoneme corresponding to each pinyin is used as the start time of the pinyin, and the end time of the last phoneme of the pinyin is used as the end time of the pinyin, so as to complete the optimization of the word-level alignment operation.

[0089] Further, as described above, when obtaining a word-level aligned transcript, it is necessary to input the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model. When obtaining a phoneme-level aligned transcript, it is necessary to input the word-level aligned transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model.

[0090] In the following, we will explain in detail how to build a word-level alignment model and a phoneme-level alignment model. Figure 5 and Figure 6 shown. Figure 5 shows the process of building a word-level alignment model, Figure 6 The process of building a phoneme-level alignment model is shown.

[0091] First, see Figure 5 As shown, the method steps include:

[0092] Step 510: Acquire multiple voice sample data and acquire a character-level transcript corresponding to each voice sample data.

[0093] Specifically, the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data. The specific acquisition process is referred to the process of acquiring voice data and the word-level transcript corresponding to the voice data above, which will not be repeated here.

[0094] Step 520 , extract feature vectors from the multiple voice sample data of the i-th person respectively to obtain the i-th information vector.

[0095] Step 530: extract frequency domain features from the multiple voice sample data of the i-th person respectively to obtain a second frequency spectrum feature sequence corresponding to each voice sample data.

[0096] Specifically, the process of obtaining the feature vector and the process of extracting the frequency domain features can be obtained through existing technologies, which will not be described in detail here.

[0097] Step 540, iteratively train the word-level alignment model according to each second spectral feature sequence, the word-level transcript corresponding to each speech sample data, and the i-th information vector, until the word-level alignment model reaches a first preset standard, and obtain the final word-level alignment model.

[0098] Specifically, the model can be trained according to the conventional neural network model iterative training process. However, in order to reduce the influence of channel and additive noise, before iteratively training the word-level alignment model, each second spectrum feature sequence can be normalized to obtain a third spectrum feature sequence that corresponds to each second spectrum feature sequence and conforms to the normal distribution;

[0099] Then, the third spectral feature sequence is sorted according to the sequence length; finally, according to the sorting order, during the jth training, the third spectral feature sequence sorted in the jth position, the word-level transcript corresponding to each voice data, and the i-th information vector are input into the word-level alignment model to complete the j-th training of the word-level alignment model.

[0100] In a specific example, the third spectral feature sequence needs to be sorted from long to short. The reason why the third spectral feature sequence is sorted according to the sequence length and then trained in the jth order is to provide a step-by-step process for the model and improve the robustness of the model. The reason why the speaker's information vector is added is to improve the generalization of the model.

[0101] In an optional example, whether the word-level alignment model meets the first preset standard can be determined according to a preconfigured loss function. As time and the number of iterations increase, the loss becomes smaller. When the loss is less than a preset threshold, the standard is met. Alternatively, after the number of iterations reaches a preset number, the training is stopped. By default, the model obtained after reaching the preset number of training times is the word-level alignment model that is ultimately desired, where i and j are both positive integers.

[0102] Figure 6 The construction process of the phoneme-level alignment model is shown below:

[0103] Step 610: Acquire multiple voice sample data and acquire a character-level transcript corresponding to each voice sample data.

[0104] Among them, the word-level transcript is based on Figure 6 The corresponding method steps are obtained. The multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data.

[0105] Step 620: extract frequency domain features from each piece of speech sample data to obtain a fourth frequency spectrum feature sequence corresponding to each piece of speech sample data.

[0106] Specifically, the feature extraction process is a prior art and will not be described in detail here.

[0107] Step 630: Normalize each fourth frequency spectrum feature sequence to obtain a fifth frequency spectrum feature sequence that corresponds to each fourth frequency spectrum feature sequence and conforms to normal distribution.

[0108] Specifically, the purpose of normalizing the fourth frequency spectrum feature sequence is also to reduce the influence of channel and additive noise.

[0109] Step 640: construct a data set by combining each fifth spectral feature sequence, the character-level transcript and the phoneme-level transcript corresponding to each fifth spectral feature sequence.

[0110] For example, these three types of data are bound together and unique ID information is generated.

[0111] Step 650: randomly split all data sets into a plurality of subsets, each subset including a data set corresponding to each of a plurality of persons.

[0112] Specifically, each subset includes a data set corresponding to each of the multiple persons in order to ensure that the model is more robust.

[0113] Step 660, using multiple subsets, iteratively train the phoneme-level alignment model in sequence, and end when the phoneme-level alignment model reaches a second preset standard to obtain the final phoneme-level alignment model, that is, the pre-built phoneme-level alignment model.

[0114] Specifically, the model generated by the previous subset training is used as the initialization model for the next subset training, but the model in the first step is randomly initialized. The last step is to iteratively train using all the training data. The subset data is sequentially entered into the phoneme-level alignment model from small to large to perform alignment training, so that the model has a step-by-step process and improves its complexity.

[0115] In addition to the training process, the testing process of the conventional neural network model is also included, which will not be elaborated here.

[0116] Further optionally, in addition to the above operations, the method may also include: constructing a character-level pronunciation dictionary and a phoneme-level pronunciation dictionary. When constructing a character-level pronunciation dictionary, the transcript in the speech sample data may be split into units of characters to generate character-level transcripts, and the character-level pronunciation dictionary may be generated after deduplication of characters in all transcripts, and this process does not require word segmentation processing.

[0117] When constructing a phoneme pronunciation dictionary, since each Chinese character may have multiple pronunciations, each Chinese character may map to multiple phoneme combinations. The character-level transcript corresponding to the speech template data can generate a phoneme-level transcript according to semantics and pronunciation rules. In a specific scenario, each Chinese character basically has its fixed pronunciation and corresponding meaning, so each Chinese character corresponds to a unique pair of phoneme combinations, with spaces between phonemes for convenient program processing.

[0118] In the above text, when generating the initial character-level alignment result using the character-level alignment model, it can be assisted by the character-level pronunciation dictionary. Similarly, when generating the initial phoneme-level alignment result using the phoneme-level alignment model, it can also be assisted by the phoneme-level pronunciation dictionary.

[0119] Moreover, in the above initial character-level alignment result and phoneme-level alignment result, in the character-level frame data or phoneme-level frame data involving tones, there will also be corresponding tone annotations. For example, in the phoneme-level alignment result, "你好" (nǐ hǎo) finally generates "n i3 h ao3", and the 3 inside represents the tone.

[0120] The two-level speech alignment method provided by the embodiments of the present invention is applicable to speech data and the character-level transcript corresponding to the speech data. Then, the speech data and the character-level transcript corresponding to the speech data are input into a pre-constructed character-level alignment model, and then the initial character-level alignment result is obtained. Each frame of character-level frame data in the initial character-level alignment result is traversed, and the frame data corresponding to the character and the first type of placeholder in the character-level frame data are respectively merged. The placeholder is used to represent non-speech data. The time of the placeholder in the speech data is allocated to the adjacent character to ensure that the speech of each character in the speech data can correspond to the character in the transcript, realizing the alignment of the character-level transcript.

[0121] On this basis, each character in the character-level transcript is used as a phoneme transcript, and the first spectral feature sequence corresponding to the speech data and the information vector corresponding to the speech data are input into the pre-built phoneme-level alignment model to obtain the initial phoneme-level alignment result. Then, each frame of phoneme-level frame data is traversed, and a merging process of the phoneme-level frame data is performed to merge the second type of placeholder with the frame data corresponding to the adjacent phoneme to obtain a merged frame. Among them, the second type of placeholder is also a phoneme-level placeholder corresponding to non-speech data. The start time and end time of the merged frame are used as the start time and end time of the phoneme. In this way, phoneme-level transcript alignment is achieved. As mentioned above, because the time corresponding to all non-speech data has been reasonably allocated to the phoneme-level frame data corresponding to the adjacent phoneme-level transcripts, that is, the time corresponding to the non-speech data is allocated to the adjacent phonemes, the obtained phoneme-level transcripts will not be misplaced at the beginning position of the speech alignment result or have tailing at the end position. This can ensure that the final phoneme-level transcript can basically correspond to the speech itself, thus eliminating errors as much as possible.

[0122] The above are several method embodiments of two-level speech alignment provided by the present application. The following describes other embodiments of two-level speech alignment provided by the present application. Please refer to the following for details.

[0123] Figure 7 A two-level speech alignment device is provided in an embodiment of the present invention. The device includes: an acquisition module 701, a processing module 702, a traversal module 703, a recording module 704, and an updating module 705.

[0124] The acquisition module 701 is used to acquire speech data and word-level transcripts corresponding to the speech data.

[0125] The processing module 702 is used to input the speech data and the character-level transcript corresponding to the speech data into a pre-built character-level alignment model to obtain an initial character-level alignment result corresponding to the speech data, wherein the initial character-level alignment result includes multiple frames of character-level frame data, and each frame of character-level frame data corresponds to a character in the character-level transcript, or corresponds to a first-class placeholder, and a character in the character-level transcript corresponds to at least one frame of character-level frame data;

[0126] A traversal module 703, used for traversing each frame of word-level frame data;

[0127] The processing module 702 is further used to merge adjacent frames belonging to the same word to obtain first merged frame data;

[0128] Recording module 704, used to record the start time and end time of the first merged frame data;

[0129] The processing module 702 is further configured to merge the first type of placeholder with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder according to the first preset rule to obtain second merged frame data;

[0130] An updating module 705 is used to update the start time and the end time of the second merged frame data, and finally obtain a word-level aligned transcript corresponding to the speech data;

[0131] The acquisition module 701 is further used to acquire a first frequency spectrum feature sequence corresponding to the speech data and an information vector corresponding to the speech data;

[0132] The processing module 702 is further used to input the word-level alignment transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model to obtain an initial phoneme-level alignment result corresponding to the speech data, wherein the initial phoneme-level alignment result includes multiple frames of phoneme-level frame data, each frame of phoneme-level frame data corresponds to a phoneme unit, or corresponds to a second type of placeholder;

[0133] The traversal module 703 is also used to traverse each frame of phoneme-level frame data;

[0134] The processing module 702 is further configured to merge at least one second type of placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data;

[0135] The recording module 704 is further used to record the start time and the end time of the third merged frame data as the start time and the end time corresponding to the phoneme unit, and finally obtain the phoneme-level aligned transcript corresponding to the speech data.

[0136] Optionally, the apparatus further includes: an optimization module 706, configured to optimize the word-level aligned transcript corresponding to the speech data using the phoneme-level aligned transcript corresponding to the speech data.

[0137] Optionally, the processing module 702 is specifically configured to: when the phoneme-level frame data is adjacent to at least one second-type placeholder, merge the at least one second-type placeholder to obtain fourth merged frame data, and determine the start time and end time of the fourth merged frame data;

[0138] The phoneme-level frame data is merged with the fourth merged frame data to obtain third merged frame data.

[0139] Optionally, the recording module 704 is specifically configured to, when the fourth merged frame data is on the left side of the phoneme-level frame data, use the start time of the fourth merged frame data as the start time of the third merged frame data, and use the end time of the phoneme-level frame data as the end time of the third merged frame data;

[0140] Alternatively, when the fourth merged frame data is on the right side of the phoneme level frame data, the start time of the phoneme level frame data is used as the start time of the third merged frame data, and the interpretation time of the fourth merged frame data is used as the end time of the third merged frame data.

[0141] Optionally, the second type of placeholders includes phoneme-level placeholders and / or silence markers; the processing module 702 is specifically configured to:

[0142] When the second type of placeholders all include phoneme-level placeholders, each phoneme-level placeholder is converted into a silence mark, and then the silence marks are merged to obtain fourth merged frame data;

[0143] Alternatively, when the second type of placeholders includes phoneme-level placeholders and silence placeholders, after merging the phoneme-level placeholders with the silence placeholders adjacent thereto, all silence placeholders are merged to obtain fourth merged frame data.

[0144] Optionally, the acquisition module 701 is further used to acquire multiple voice sample data and acquire a word-level transcript corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data;

[0145] The processing module 702 is further used to extract feature vectors from the multiple voice sample data of the i-th person respectively to obtain the i-th information vector;

[0146] Perform frequency domain feature extraction on multiple voice sample data of the i-th person respectively to obtain a second frequency spectrum feature sequence corresponding to each voice sample data;

[0147] According to each second spectral feature sequence, the word-level transcript corresponding to each speech sample data, and the i-th information vector, the word-level alignment model is iteratively trained until the word-level alignment model reaches a first preset standard, and a final word-level alignment model is obtained, that is, a pre-constructed word-level alignment model wherein i is a positive integer.

[0148] Optionally, the processing module 702 is further configured to perform normalization processing on each second spectrum feature sequence respectively to obtain a third spectrum feature sequence that corresponds to each second spectrum feature sequence and conforms to normal distribution;

[0149] sorting the third spectrum feature sequence according to sequence length;

[0150] In the jth training, the third spectral feature sequence ranked in the jth position, the word-level transcript corresponding to each voice data, and the i-th information vector are input into the word-level alignment model in order to complete the j-th training of the word-level alignment model, where j is a positive integer.

[0151] Optionally, the acquisition module 701 is further used to acquire multiple voice sample data, and acquire a character-level transcript and a phoneme-level transcript corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data;

[0152] The processing module 702 is further used to extract frequency domain features for each piece of speech sample data to obtain a fourth frequency spectrum feature sequence corresponding to each piece of speech sample data;

[0153] Normalizing each fourth frequency spectrum feature sequence respectively to obtain a fifth frequency spectrum feature sequence that corresponds to each fourth frequency spectrum feature sequence and conforms to normal distribution;

[0154] Constructing a data set by using each fifth spectral feature sequence, a word-level transcript and a phoneme-level transcript corresponding to each fifth spectral feature sequence;

[0155] randomly splitting all data sets into a plurality of subsets, each subset including a data set corresponding to each of a plurality of persons;

[0156] The phoneme-level alignment model is iteratively trained in sequence using multiple subsets, and the training ends when the phoneme-level alignment model reaches a second preset standard, thereby obtaining a final phoneme-level alignment model, that is, a pre-built phoneme-level alignment model.

[0157] The functions performed by the various components in the two-stage speech alignment device provided in the embodiment of the present invention have been described in detail in any of the above method embodiments, and therefore will not be repeated here.

[0158] A two-level speech alignment device provided by an embodiment of the present invention is for speech data and a character-level transcript corresponding to the speech data. Then the speech data and the character-level transcript corresponding to the speech data are input into a pre-built character-level alignment model, and then an initial character-level alignment result is obtained. Each frame of character-level frame data in the initial character-level alignment result is traversed, and the frame data corresponding to the character in the character-level frame data and the first type of placeholder are merged and processed respectively. The placeholder is used to represent non-speech data. The time of the placeholder in the speech data is allocated to the adjacent characters to ensure that each character speech in the speech data can correspond to the character in the transcript, thereby realizing character-level transcript alignment.

[0159] On this basis, each character in the character-level transcript is used as a phoneme transcript, and the first spectral feature sequence corresponding to the speech data and the information vector corresponding to the speech data are input into the pre-built phoneme-level alignment model to obtain the initial phoneme-level alignment result. Then, each frame of phoneme-level frame data is traversed, and a merging process of the phoneme-level frame data is performed to merge the second type of placeholder with the frame data corresponding to the adjacent phoneme to obtain a merged frame. Among them, the second type of placeholder is also a phoneme-level placeholder corresponding to non-speech data. The start time and end time of the merged frame are used as the start time and end time of the phoneme. In this way, phoneme-level transcript alignment is achieved. As mentioned above, because the time corresponding to all non-speech data has been reasonably allocated to the phoneme-level frame data corresponding to the adjacent phoneme-level transcripts, that is, the time corresponding to the non-speech data is allocated to the adjacent phonemes, the obtained phoneme-level transcripts will not be misplaced at the beginning position of the speech alignment result or have tailing at the end position. This can ensure that the final phoneme-level transcript can basically correspond to the speech itself, thus eliminating errors as much as possible.

[0160] like Figure 8 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0161] Memory 113, used for storing computer programs;

[0162] In one embodiment of the present application, the processor 111 is used to execute the program stored in the memory 113 to implement the two-level speech alignment method provided by any of the above method embodiments, including:

[0163] obtaining speech data, and word-level transcripts corresponding to the speech data;

[0164] Inputting the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model to obtain an initial word-level alignment result corresponding to the speech data, wherein the initial word-level alignment result includes a plurality of frames of word-level frame data, and each frame of word-level frame data corresponds to a word in the word-level transcript, or corresponds to a first-category placeholder, and a word in the word-level transcript corresponds to at least one frame of word-level frame data;

[0165] Traverse each frame of word-level frame data, merge adjacent frames belonging to the same word, obtain first merged frame data, and record the start time and end time of the first merged frame data;

[0166] According to a first preset rule, the first type of placeholder is merged with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder to obtain second merged frame data;

[0167] Updating the start time and end time of the second merged frame data, and finally obtaining a word-level aligned transcript corresponding to the speech data;

[0168] Acquire a first frequency spectrum feature sequence corresponding to the speech data and an information vector corresponding to the speech data;

[0169] Inputting the word-level aligned transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model to obtain an initial phoneme-level alignment result corresponding to the speech data, wherein the initial phoneme-level alignment result includes multiple frames of phoneme-level frame data, each frame of the phoneme-level frame data corresponds to a phoneme unit, or corresponds to a second-category placeholder;

[0170] Traversing each frame of phoneme-level frame data, merging at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data;

[0171] The start time and the end time of the third merged frame data are recorded as the start time and the end time corresponding to the phoneme unit, and finally a phoneme-level aligned transcript corresponding to the speech data is obtained.

[0172] Optionally, after obtaining the phoneme-level aligned transcript corresponding to the speech data, the method further includes:

[0173] The word-level aligned transcripts corresponding to the speech data are optimized using the phoneme-level aligned transcripts corresponding to the speech data.

[0174] Optionally, according to a first preset rule, merging the first type of placeholder with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder to obtain second merged frame data specifically includes:

[0175] When the left side of the first type of placeholder is word-level frame data corresponding to the word, the first type of placeholder is merged with the word-level frame data corresponding to the word on the left side, and the end time of the first type of placeholder is used as the end time of the second merged frame, and the start time of the word-level frame data with the word on the left side is used as the start time of the second merged frame;

[0176] Alternatively, when the first type of placeholder is word-level frame data corresponding to the word on the right side, the first type of placeholder is merged with the word-level frame data corresponding to the word on the right side, and the start time of the first type of placeholder is used as the start time of the second merged frame, and the end time of the word-level frame data corresponding to the word on the right side is used as the end time of the second merged frame.

[0177] Optionally, traversing each frame of phoneme-level frame data, merging at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data, specifically including:

[0178] When the phoneme-level frame data is adjacent to at least one second-category placeholder, merge the at least one second-category placeholder to obtain fourth merged frame data, and determine the start time and end time of the fourth merged frame data;

[0179] The phoneme-level frame data is merged with the fourth merged frame data to obtain third merged frame data.

[0180] Optionally, recording the start time and the end time of the third merged frame data as the start time and the end time corresponding to the phoneme unit specifically includes:

[0181] When the fourth merged frame data is on the left side of the phoneme-level frame data, the start time of the fourth merged frame data is used as the start time of the third merged frame data, and the end time of the phoneme-level frame data is used as the end time of the third merged frame data;

[0182] Alternatively, when the fourth merged frame data is on the right side of the phoneme level frame data, the start time of the phoneme level frame data is used as the start time of the third merged frame data, and the interpretation time of the fourth merged frame data is used as the end time of the third merged frame data.

[0183] Optionally, the second type of placeholder includes a phoneme-level placeholder and / or a silence mark; when the phoneme-level frame data is adjacent to at least one second type of placeholder, merging at least one second type of placeholder to obtain fourth merged frame data, specifically comprising: when the second type of placeholders all include phoneme-level placeholders, converting each phoneme-level placeholder into a silence mark, merging the silence marks, and obtaining fourth merged frame data;

[0184] Alternatively, when the second type of placeholders includes phoneme-level placeholders and silence placeholders, after merging the phoneme-level placeholders with the silence placeholders adjacent thereto, all silence placeholders are merged to obtain fourth merged frame data.

[0185] Optionally, before inputting the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model, the method further comprises:

[0186] Acquire multiple voice sample data, and acquire word-level transcripts corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data;

[0187] Extract feature vectors from multiple voice sample data of the i-th person respectively to obtain the i-th information vector;

[0188] Perform frequency domain feature extraction on multiple voice sample data of the i-th person respectively to obtain a second frequency spectrum feature sequence corresponding to each voice sample data;

[0189] According to each second spectral feature sequence, the word-level transcript corresponding to each speech sample data, and the i-th information vector, the word-level alignment model is iteratively trained until the word-level alignment model reaches a first preset standard, and a final word-level alignment model is obtained, that is, a pre-constructed word-level alignment model wherein i is a positive integer.

[0190] Optionally, iteratively training the word-level alignment model according to each second spectral feature sequence, the word-level transcript corresponding to each speech sample data, and the i-th information vector specifically includes:

[0191] Normalizing each second frequency spectrum feature sequence respectively to obtain a third frequency spectrum feature sequence that corresponds to each second frequency spectrum feature sequence and conforms to normal distribution;

[0192] sorting the third spectrum feature sequence according to sequence length;

[0193] In the jth training, the third spectral feature sequence ranked in the jth position, the word-level transcript corresponding to each voice data, and the i-th information vector are input into the word-level alignment model in order to complete the j-th training of the word-level alignment model, where j is a positive integer.

[0194] Optionally, before inputting the word-level aligned transcript, the first spectral feature sequence, and the information vector into a pre-built phoneme-level alignment model, the method further comprises:

[0195] Acquire multiple voice sample data, and acquire word-level transcripts and phoneme-level transcripts corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data;

[0196] Perform frequency domain feature extraction on each piece of speech sample data to obtain a fourth frequency spectrum feature sequence corresponding to each piece of speech sample data;

[0197] Normalizing each fourth frequency spectrum feature sequence respectively to obtain a fifth frequency spectrum feature sequence that corresponds to each fourth frequency spectrum feature sequence and conforms to normal distribution;

[0198] Constructing a data set by using each fifth spectral feature sequence, a word-level transcript and a phoneme-level transcript corresponding to each fifth spectral feature sequence;

[0199] randomly splitting all data sets into a plurality of subsets, each subset including a data set corresponding to each of a plurality of persons;

[0200] The phoneme-level alignment model is iteratively trained in sequence using multiple subsets, and the training ends when the phoneme-level alignment model reaches a second preset standard, thereby obtaining a final phoneme-level alignment model, that is, a pre-built phoneme-level alignment model.

[0201] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the two-level speech alignment method provided in any of the aforementioned method embodiments are implemented.

[0202] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0203] The above is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features applied herein.

Claims

1. A two-level speech alignment method, characterized in that: The method comprises: Acquiring speech data and a word-level transcript corresponding to the speech data; Inputting the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model to obtain an initial word-level alignment result corresponding to the speech data, wherein the initial word-level alignment result includes a plurality of frames of word-level frame data, and each frame of the word-level frame data corresponds to a word in the word-level transcript, or corresponds to a first-class placeholder, and a word in the word-level transcript corresponds to at least one frame of the word-level frame data; Traversing each frame of the word-level frame data, merging adjacent frames belonging to the same word, obtaining first merged frame data, and recording the start time and end time of the first merged frame data; According to a first preset rule, the first type of placeholder is merged with the word-level frame data corresponding to the word immediately adjacent to the first type of placeholder to obtain second merged frame data; Updating the start time and end time of the second merged frame data, and finally obtaining a word-level aligned transcript corresponding to the speech data; Acquire a first frequency spectrum feature sequence corresponding to the speech data and an information vector corresponding to the speech data; Inputting the word-level aligned transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model to obtain an initial phoneme-level alignment result corresponding to the speech data, wherein the initial phoneme-level alignment result includes multiple frames of phoneme-level frame data, each frame of the phoneme-level frame data corresponds to a phoneme unit, or corresponds to a second-category placeholder; Traversing each frame of the phoneme-level frame data, merging at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data; The start time and the end time of the third merged frame data are recorded as the start time and the end time corresponding to the phoneme unit, and finally a phoneme-level aligned transcript corresponding to the speech data is obtained.

2. The method according to claim 1, characterized in that After obtaining the phoneme-level aligned transcript corresponding to the speech data, the method further includes: The word-level aligned transcript corresponding to the speech data is optimized using the phoneme-level aligned transcript corresponding to the speech data.

3. The method according to claim 1, characterized in that Traversing each frame of the phoneme-level frame data, merging at least one second-type placeholder adjacent to the phoneme-level frame data corresponding to the phoneme unit in the initial phoneme-level alignment result with the phoneme-level frame data corresponding to the phoneme unit to obtain third merged frame data, specifically comprising: When the phoneme-level frame data is adjacent to at least one second-category placeholder, merge at least one second-category placeholder to obtain fourth merged frame data, and determine the start time and end time of the fourth merged frame data; The phoneme-level frame data is merged with the fourth merged frame data to obtain the third merged frame data.

4. The method according to claim 3, characterized in that Recording the start time and end time of the third merged frame data as the start time and end time corresponding to the phoneme unit specifically includes: When the fourth merged frame data is on the left side of the phoneme-level frame data, the start time of the fourth merged frame data is used as the start time of the third merged frame data, and the end time of the phoneme-level frame data is used as the end time of the third merged frame data; Alternatively, when the fourth merged frame data is on the right side of the phoneme level frame data, the start time of the phoneme level frame data is used as the start time of the third merged frame data, and the interpretation time of the fourth merged frame data is used as the end time of the third merged frame data.

5. The method according to claim 3 or 4, characterized in that: The second type of placeholders include phoneme-level placeholders and / or silence marks; when the phoneme-level frame data is adjacent to at least one second type of placeholder, merging at least one second type of placeholder to obtain fourth merged frame data specifically includes: When the second type of placeholders all include the phoneme-level placeholders, converting each phoneme-level placeholder into a silence mark, merging the silence marks, and acquiring the fourth merged frame data; Alternatively, when the second type of placeholders includes phoneme-level placeholders and silence placeholders, the phoneme-level placeholders are merged with the silence placeholders adjacent thereto, and then all silence placeholders are merged to obtain the fourth merged frame data.

6. The method according to any one of claims 1 to 4, characterized in that: Before inputting the speech data and the word-level transcript corresponding to the speech data into a pre-built word-level alignment model, the method further comprises: Acquire multiple voice sample data, and acquire word-level transcripts corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data; Extract feature vectors from multiple voice sample data of the i-th person respectively to obtain the i-th information vector; Perform frequency domain feature extraction on multiple voice sample data of the i-th person respectively to obtain a second frequency spectrum feature sequence corresponding to each voice sample data; According to each of the second spectral feature sequences, the word-level transcripts corresponding to each speech sample data, and the i-th information vector, the word-level alignment model is iteratively trained until the word-level alignment model reaches a first preset standard, and a final word-level alignment model is obtained, that is, the pre-constructed word-level alignment model, wherein i is a positive integer.

7. The method according to claim 6, characterized in that Iteratively training the word-level alignment model according to each of the second spectral feature sequences, the word-level transcripts corresponding to each of the speech sample data, and the i-th information vector specifically includes: Normalizing each of the second frequency spectrum feature sequences respectively to obtain a third frequency spectrum feature sequence that corresponds to each of the second frequency spectrum feature sequences and conforms to normal distribution; Sorting the third frequency spectrum feature sequence according to sequence length; In the jth training, in sequence, the third spectral feature sequence ranked at the jth position, the word-level transcript corresponding to each piece of speech data, and the i-th information vector are input into the word-level alignment model to complete the j-th training of the word-level alignment model, where j is a positive integer.

8. The method according to any one of claims 1 to 4, characterized in that: Before inputting the word-level aligned transcript, the first spectral feature sequence and the information vector into a pre-built phoneme-level alignment model, the method further comprises: Acquire multiple voice sample data, and acquire word-level transcripts and phoneme-level transcripts corresponding to each voice sample data; wherein the multiple voice sample data are voice sample data of multiple persons, and each person corresponds to at least two voice sample data; Perform frequency domain feature extraction on each piece of speech sample data to obtain a fourth frequency spectrum feature sequence corresponding to each piece of speech sample data; Normalizing each of the fourth frequency spectrum characteristic sequences respectively to obtain a fifth frequency spectrum characteristic sequence that corresponds to each of the fourth frequency spectrum characteristic sequences and conforms to normal distribution; Constructing a data set by using each of the fifth spectral feature sequences, the word-level transcripts and the phoneme-level transcripts corresponding to each of the fifth spectral feature sequences; randomly splitting all data sets into a plurality of subsets, each subset including a data set corresponding to each of a plurality of persons; The phoneme-level alignment model is iteratively trained in sequence using multiple subsets, and the training ends when the phoneme-level alignment model reaches a second preset standard, thereby obtaining a final phoneme-level alignment model, namely the pre-constructed phoneme-level alignment model.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor is used to implement the steps of the two-level speech alignment method described in any one of claims 1-8 when executing a program stored in a memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the two-level speech alignment method as described in any one of claims 1-8 are implemented.

Citation Information

Patent Citations

  • Voice alignment method and device

    CN108682436A

  • Method and device for aligning synthesized voice with text, and computer storage medium

    CN112420016A