Miao language ternary database construction method, model training method and speech recognition method

By constructing the Miao language ternary database, the data loss problem in the construction of the Miao language speech recognition model is solved, the high-accurate speech recognition effect is achieved, and the threshold for data construction is lowered.

CN120011473APending Publication Date: 2025-05-16SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411838698.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The lack of text records in Miao language leads to difficulties in building speech recognition models and database construction, especially in obtaining complete speech-text-phoneme correspondence.

Method used

A Miao language ternary database construction method is proposed. Through text processing, audio processing and data mapping steps, a complete database including text, audio and phoneme parts is obtained, and a speech recognition model is trained based on the database.

Benefits of technology

It has realized the construction of a Miao language ternary database suitable for all types of speech recognition processes, which has improved the accuracy and applicability of the speech recognition model, lowered the threshold for data construction, and promoted the development of Miao language speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011473A_ABST
    Figure CN120011473A_ABST
Patent Text Reader

Abstract

The invention relates to a germchit language ternary database construction method, a model training method and a speech recognition method. The construction method comprises the following steps: acquiring Miao language text data as a text part of a Miao language ternary database according to a preset text data rule; according to a first mapping rule and a preset audio data rule, Miao language audio data is obtained to serve as an audio part of a Miao language ternary database, and the first mapping rule is used for determining a mapping relation between the Miao language text data and the Miao language audio data; and according to a second mapping rule, obtaining Miao speech element data as a phoneme part of the Miao language ternary database, the second mapping rule being used for determining a mapping relationship between the Miao language text data and the Miao language element data, so that a complete Miao language ternary database comprising a text part, an audio part and the phoneme part is obtained. Compared with the prior art, the Miao language recognition device has the advantages of being higher in applicability, more reasonable in structure, capable of obtaining an accurate Miao language recognition result and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of media asset management, and in particular to a Miao language triple database construction method, a model training method and a speech recognition method. Background Art

[0002] Miao is a language but does not have its own written language. Nowadays, most Miao language is recorded directly by pronunciation and using English letters. The lack of written language brings difficulties to speech recognition workflow and database construction.

[0003] For the traditional speech recognition process, the original audio needs to be processed for features to obtain a phoneme sequence, and then the language model decodes the phoneme sequence to obtain the final speech recognition result. Taking Chinese as an example, there is a strict correspondence between Chinese characters and pinyin, and pinyin can be mapped to phonemes. When annotating, Chinese characters can be used directly for annotation. Since Miao language has no characters, Chinese characters are used as a substitute to express semantics, but Chinese characters and Miao pinyin are not strictly one-to-one corresponding, which makes it difficult to build a pronunciation dictionary and corpus that adapts to the speech recognition model, and it also adds more processes to the decoding of the language model.

[0004] There are also difficulties in collecting Miao speech corpus. The annotated data pairs required for the speech recognition process include data pairs of original audio and phonemes and data pairs of phonemes and Chinese characters. The original audio, phonemes, and Chinese characters must correspond one to one. However, in fact, the existing data cannot meet this requirement. First of all, many of the existing Miao voice programs that can be used as corpora do not have Chinese subtitles. Secondly, there is no corresponding Miao pinyin. Even if they are annotated, there are only data pairs from original audio to Chinese characters; and the existing Miao dictionaries are data pairs from Miao pinyin to Chinese characters, lacking original audio. Therefore, how to obtain an effective Miao speech recognition model and then obtain accurate Miao recognition results has become a problem that needs to be solved in this field. Summary of the invention

[0005] The purpose of the present invention is to provide a Miao language triple database construction method, a model training method and a speech recognition method in order to overcome the defect of the Miao language speech recognition database in the above-mentioned prior art.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] According to a first aspect of the present invention, there is provided a method for constructing a Qiandongnan Miao language triple database for speech recognition, comprising the following steps: a text processing step: according to a preset text data rule, Miao language text data is obtained as the text part of the Miao language triple database; an audio processing step: according to a first mapping rule and a preset audio data rule, Miao language audio data is obtained as the audio part of the Miao language triple database, wherein the first mapping rule is used to determine the mapping relationship between the Miao language text data and the Miao language audio data; a Miao language triple database acquisition step: according to a second mapping rule, Miao language phoneme data is obtained as the phoneme part of the Miao language triple database, wherein the second mapping rule is used to determine the mapping relationship between the Miao language text data and the Miao language phoneme data, thereby obtaining a complete Miao language triple database including the text part, the audio part and the phoneme part.

[0008] As a preferred technical solution, the text data rules include text database construction rules and text data format rules, and the text processing steps specifically include: according to the text database construction rules, collating and screening the pre-acquired Miao language physical books to determine the basic text content; scanning the basic text content and performing text recognition to obtain editable text; correcting the erroneous content of the editable text; and according to the text data format rules, performing format correction on the corrected text to obtain the Miao language text data.

[0009] As a preferred technical solution, the text database construction rules include usage requirements and pronunciation requirements. The usage requirements are used to ensure that the basic text content covers the daily life expressions and uncommon words of the Miao nationality, and the proportion of the daily life expressions in the basic text content is higher than that of the uncommon words. The pronunciation requirements are used to ensure that the pronunciation of the basic text content covers all initials and finals of the Miao language.

[0010] As an optimal technical solution, the text data format rules specifically include: taking lines as text data units, each line of text data includes a Chinese word and the corresponding Miao pinyin, wherein the Chinese word is in front and the Miao pinyin is in the back, and the two are separated by a space; when a Chinese word corresponds to multiple Miao pinyins, different Miao pinyins are separated by spaces; punctuation marks are not entered, and punctuation marks are used as line break marks.

[0011] As a preferred technical solution, the audio data rules include audio data collection rules and audio recording rules, and the audio processing steps specifically include: recording the original audio corresponding to the Miao language text data according to the first mapping rules and the audio data collection rules, and complying with the audio recording rules during recording; and editing, volume balancing, audio noise reduction and classification of the original audio in sequence to obtain the Miao language audio data.

[0012] As a preferred technical solution, the audio data collection rules specifically include: selecting the persons to be collected and recording representative Miao pronunciations; the residence of the persons to be collected is a typical village where ethnic minorities gather; the representative Miao pronunciation is that the collected vocabulary contains all initials, finals and tones and the pronunciation phonemes are evenly distributed; the audio recording rules specifically include: recording by person, using ZOOM H6 as a recorder and a directional microphone as a recording device, with a sampling rate of 16 bits, a frequency of 48kHz, and a recorder level of -3dB.

[0013] As a preferred technical solution, the second mapping rule is determined based on a mapping table between Miao pinyin and international phonetic symbols.

[0014] As a preferred technical solution, in the Miao language triple database, the data of the text part is stored in txt format, the data of the audio part is stored in wav format, and the data of the phoneme part is stored in txt format, wherein each data of the audio part is configured with two txt text annotation files, which respectively represent the Chinese annotation and the Miao language annotation.

[0015] According to a second aspect of the present invention, a method for training a speech recognition model is provided, wherein the speech recognition model includes a language model, an acoustic model and a dictionary, and the model training method includes the following steps: training the language model using the text part of a Miao language triple database; training the acoustic model using the audio part of a Miao language triple database; constructing a dictionary using the text part and the corresponding factor part of the Miao language triple database; wherein the Miao language triple database is obtained based on the construction method.

[0016] According to a third aspect of the present invention, there is provided a speech recognition method, comprising the following steps: obtaining a Miao language text to be processed; inputting the Miao language text to be processed into a trained speech recognition model to obtain a Miao language speech recognition result; wherein the speech recognition model is trained using a Miao language triple database, and the Miao language triple database is obtained based on the construction method.

[0017] Compared with the prior art, the present invention has the following beneficial effects:

[0018] 1. The present invention provides a method for constructing a Miao language ternary database for speech recognition. Based on the construction method, a Miao language ternary database with a complete speech-text-phoneme correspondence relationship can be obtained. The data in the database can be used for all types of existing speech recognition processes, including end-to-end speech recognition, unsupervised speech recognition, and traditional speech recognition. The database has stronger applicability and more reasonable structure. The Miao language speech recognition model is trained by using the ternary database, and the trained model is used for speech recognition, so that accurate Miao language recognition results can be obtained.

[0019] 2. The present invention provides a process and construction standard for constructing a three-yuan Miao language database, which is normative;

[0020] 3. In the present invention, ternary data is derived from unary data, which has low requirements on source data and a lower construction threshold. The source data can be used to derive the most information with the least process, which can improve data utilization efficiency and save training resources.

[0021] 4. The method of the present invention links the Miao language to the International Phonetic Alphabet to perform phoneme-sharing speech recognition, so that the Miao language speech recognition can be developed under low resource conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flowchart of the method for constructing a Miao language triple database in an embodiment of the present invention;

[0023] Figure 2 Schematic diagram of the flow of the speech recognition model training method in an embodiment of the present invention;

[0024] Figure 3 The figure is a flow chart of a speech recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0026] Example

[0027] In order to protect the traditional minority languages, it is necessary to collect the relevant voice and text materials of the Miao language, build a complete three-element Miao language database, digitize the traditional culture, study the speech recognition method of the Miao language, and help the Miao language to be passed down for a long time. With a complete database, learners can easily compare the Miao language and the Chinese language, and unlike paper dictionaries, learners can also learn related pronunciations. Based on this, the present invention targets the pain points in the development of Miao language speech recognition, uses the method of data expansion to take the existing most easily accessible data as the starting point, expands the unary data and finally obtains the ternary data, which has the characteristics of wide applicability and rich information, and the constructed Miao language ternary database can fully promote the speech recognition research of the Miao language.

[0028] like Figure 1As shown, the present embodiment provides a method for constructing a Qiandongnan Miao language ternary database for speech recognition, the method specifically comprising: the original unigram Miao language text data is subjected to a text processing step and an audio processing step to obtain binary Miao language data, and then data mapping is performed according to a ternary data mapping rule to obtain a final Miao language ternary database. Wherein, the ternary data mapping rule includes a first mapping rule and a second mapping rule. The first mapping rule is used to determine the mapping relationship between the Miao language text data and the Miao language audio data. Optionally, the reader is made to read the Miao language pronunciation according to the Miao language text data in a one-to-one correspondence, and the first mapping rule can be determined according to the pronunciation order. The second mapping rule is used to determine the mapping relationship between the Miao language text data and the Miao language phoneme data. Optionally, the first mapping rule is determined based on the Miao language pinyin and the international phonetic symbol mapping table. The ternary database can be applied to the module required for speech recognition, can solve the difficulties existing in the Miao language speech recognition, and is conducive to promoting the development of the Miao language speech recognition.

[0029] Specifically, the Miao language triple database includes:

[0030] A1) The trigram of Miao language trigram database refers to the audio part, text part and phoneme part;

[0031] A2) In the Miao language triple database, the text data is Miao language pinyin, the phoneme data is composed of the corresponding phonemes of the text data, and the audio data is the pronunciation corresponding to the text data. The three correspond one to one with the text data as the core.

[0032] On this basis, the storage formats and requirements of the three parts of data are as follows:

[0033] A11) The audio data is stored in wav format and named "RE001P44098.wav", where "RE" is a fixed character, "001" indicates the speaker number, and "P44098" indicates the text number. Each audio is equipped with two txt text annotation files, which are named "RE001P44098H.txt" and "RE001P44098M.txt" respectively. The first part has the same meaning as the audio, and the last "H" indicates Chinese annotation, and "M" indicates Miao annotation;

[0034] A12) The text data is stored in a file in txt format, with each line storing a phrase or sentence consisting of Miao pinyin, and pronunciations are separated by spaces, for example, "ax bub nenx cenb dlenl mongljox fangb deis mongl yangx";

[0035] A13) The phoneme part is stored in a file in txt format, with lines as units. Each line contains text and corresponding international phonetic symbols. The first line is the Miao pinyin, followed by the international phonetic symbols of the initial consonants, and finally the international phonetic symbols of the finals. The words following the finals represent the tone, for example, "hfaid ff et3".

[0036] Specifically, the text processing steps include:

[0037] B1) sorting and screening existing Miao language physical books (such as Miao language dictionaries) according to preset text database construction rules to determine the basic text content to be entered into the database;

[0038] B2) Scanning the determined physical book to obtain an image, and then using existing text recognition software to perform text recognition on the image to obtain editable text;

[0039] B3) Since the editable text obtained by scanning has a lot of errors in format and content, it is necessary to compare with the original physical book to correct the errors in the editable text;

[0040] B4) performing format correction on the corrected text according to preset text data format rules;

[0041] B5) Thus, a Miao language text database with standardized content and format is obtained. Miao language does not have its own writing system, and all pronunciations are recorded in pinyin. The text here refers to the text in Miao language pinyin.

[0042] After obtaining the text data, the text data enters the ternary database and becomes the text part thereof. The text data entering the ternary database is mapped to phoneme data according to the existing Miao language pinyin and international phonetic phoneme mapping table (second mapping rule), forming the phoneme part in the ternary database. Among them, the Miao language pinyin is mapped to the international phonetic symbol so that the Miao language can use the same phoneme unit as other languages, that is, phoneme sharing, so that the Miao language speech recognition can develop under low resource conditions. On the other hand, the text data will also enter the audio processing step as the recording and reading reference data of the audio processing part (naturally forming the first mapping rule).

[0043] In step B1), the text database construction rules include usage requirements and pronunciation requirements, specifically:

[0044] Text data collection should be considered from two aspects: usage rate and pronunciation distribution. The usage rate means that the collected text data should cover as much as possible the daily language used by the Miao people in their daily lives. The daily language should account for the majority of the text data, while also covering a certain degree of rare words, which account for a small part of the text data. The pronunciation distribution means that the text pronunciations that make up the text data should cover as many initials and finals of the Miao language as possible, and the proportion of each initial and final should be evenly distributed.

[0045] In step B4, the text data format rules specifically include:

[0046] The text data is grouped by line, with each line as a group. Each line is written in the format of "Chinese Miao language pinyin", such as "cowpox def". The Chinese word and the Miao language pinyin are separated by a space. If there are multiple pronunciations of the Miao language corresponding to a Chinese word, the subsequent Miao language pronunciations are also separated by spaces, such as "stubbornness ves ninx ves liod". For the phenomenon of multiple translations of a word, the Miao language uses " / " as a separator, such as "calf ghab daib liod / ghab daib ninx". When sorting and entering the text data, punctuation marks are not entered. If there is a sentence separated by a comma, it is regarded as two lines and entered separately.

[0047] Specifically, the audio processing steps include:

[0048] C1) The text data is used as the reference data for reading aloud the recording. According to the preset audio data collection rules, a recorder is selected to read aloud the text data and record it at the same time, and the preset audio recording rules are observed during recording;

[0049] C2) In the original audio, for the convenience of marking, there are Chinese explanations recorded by the recorder, as well as recording phenomena such as dental sounds, plosive sounds, and saliva sounds caused by the recorder's own pronunciation problems. Therefore, the recorded original audio is clipped. This part will cut the Chinese part and unfavorable recording phenomena in the original audio of the recording, and finally leave only the audio containing the Miao language;

[0050] C3) Since the recorders are generally not professional recording workers, there will be a situation where the recording level fluctuates greatly during recording, which is not conducive to feature extraction and model training. Therefore, it is necessary to balance the volume of the clipped audio to ensure that the level of the entire audio is maintained at an average level, without the situation of excessive or too small level;

[0051] C4) There may be noise in the original audio recording, and these noises may be amplified after the volume balance, and new noises may also appear. Therefore, it is necessary to perform audio noise reduction on the audio after the volume balance;

[0052] C5) After the above steps, audio data is obtained, and after classification and sorting, the audio data enters the ternary database as the audio part.

[0053] In step C1), the audio data collection rules specifically include:

[0054] Record representative pronunciations, where representative means covering basic daily use, the collected vocabulary contains all initials, finals and tones, and the phonemes of each pronunciation are evenly distributed. The database collects voices from at least 20 people, including 10 Miao people aged 15-25 and 10 Miao people aged 35-45. The residence of the collected people should be a typical minority village, the Miao language they speak should be authentic and have the highest fidelity, retain the most original Miao pronunciation characteristics, and be representative.

[0055] The audio recording rules include:

[0056] The audio recording was conducted by two people, using ZOOM H6 as the recorder and a directional microphone as the recording device. The sampling rate was 16 bits and the frequency was 48kHz. The speaker faced the microphone during the recording to ensure that the recorder level was kept at around -3dB. During the recording, the speaker first read Chinese and then Miao language. If an error occurred during the recording process, the sentence was repeated, but the recorder did not stop recording until the end of a recording section, which consisted of 50 sentences.

[0057] Based on the above steps, the data entering the final database are audio data and text data, which are binary data. Since the audio data is recorded from the text data, there is a mapping relationship between them. Since the Miao language has no written language, additional expansion is required to form ternary data. The text data is mapped according to the Miao language pinyin and international phonetic symbol mapping table to obtain international phonetic symbol data. As a result, the final database becomes a ternary database with two-to-two mappings.

[0058] Furthermore, if Figure 2 As shown, this embodiment also provides a method for training a speech recognition model. Specifically:

[0059] D1) Generally, a speech recognition model includes three modules: a language model, an acoustic model, and a dictionary. The Miao language triple database obtained by the construction method proposed in this embodiment can directly input data and train these three modules. The model training method includes the following steps:

[0060] D2) Using the text part of the Miao language triple database to train a language model. Optionally, the text part exists in the form of phrases and sentences, which can be directly used for training language models, including mainstream language models such as n-gram and BERT.

[0061] D3) Using the audio part of the Miao language triple database to train the acoustic model. Optionally, the audio part as the original audio signal can be directly transformed to obtain MFCC or other acoustic features, or it can be converted into multi-dimensional features through a neural network for analysis in subsequent processes.

[0062] D4) Constructing a dictionary using the text part and the corresponding factor part of the Miao language trigram database. Specifically, the text corresponds to the phoneme part one by one, and the two can be used together to construct a dictionary.

[0063] Furthermore, if Figure 3 As shown, this embodiment also provides a speech recognition method. After obtaining the Miao language text to be processed, the method inputs the Miao language text to be processed into a trained speech recognition model to obtain a Miao language speech recognition result. Among them, the speech recognition model is trained using the aforementioned training method of the Miao language triple database, and the Miao language triple database is obtained based on the construction method proposed in this embodiment, and the specific process is not repeated here.

[0064] In summary, the Miao language ternary database construction method, model training method and speech recognition method provided by the present invention expand the existing unigram Miao language dictionary data into a ternary database for each link of the speech recognition process, and at the same time optimize the disadvantage that the Miao language has no written text. It can make up for the gap in the non-end-to-end speech recognition method of the Miao language and promote the development of speech recognition of the Miao language.

[0065] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A method for constructing a Qiandongnan Miao language triple database for speech recognition, characterized in that: The following steps are involved: Text processing step: according to the preset text data rules, obtain Miao language text data as the text part of the Miao language triple database; An audio processing step, according to a first mapping rule and a preset audio data rule, obtaining Miao language audio data as the audio part of the Miao language ternary database, wherein the first mapping rule is used to determine a mapping relationship between the Miao language text data and the Miao language audio data; Steps for acquiring the Miao language tri-gram database: According to the second mapping rule, Miao phoneme data is acquired as the phoneme part of the Miao language tri-gram database. The second mapping rule is used to determine the mapping relationship between the Miao language text data and the Miao phoneme data, thereby obtaining a complete Miao language tri-gram database including the text part, the audio part and the phoneme part.

2. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 1, characterized in that: The text data rules include text database construction rules and text data format rules, and the text processing steps specifically include: According to the text database construction rules, the pre-acquired Miao language physical books are sorted and screened to determine the basic text content; Scan the basic text content and perform text recognition to obtain editable text; Correcting erroneous content of the editable text; According to the text data format rules, the corrected text is formatted to obtain the Miao language text data.

3. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 2, characterized in that: The text database construction rules include usage requirements and pronunciation requirements. The usage requirements are used to ensure that the basic text content covers the daily life expressions and uncommon words of the Miao nationality, and the proportion of the daily life expressions in the basic text content is higher than that of the uncommon words. The pronunciation requirements are used to ensure that the pronunciation of the basic text content covers all initials and finals of the Miao language.

4. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 2, characterized in that: The text data format rules specifically include: The text data unit is a line, and each line of text data includes a Chinese word and the corresponding Miao pinyin, wherein the Chinese word is in front and the Miao pinyin is in the back, and the two are separated by a space; When a Chinese word corresponds to multiple Miao pinyins, different Miao pinyins are separated by spaces; Do not enter punctuation marks; use them as line break marks.

5. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 1, characterized in that: The audio data rules include audio data collection rules and audio recording rules, and the audio processing steps specifically include: According to the first mapping rule and the audio data collection rule, recording the original audio corresponding to the Miao language text data, and complying with the audio recording rule during recording; The Miao language audio data is obtained by editing, volume balancing, audio noise reduction and classification of the original audio.

6. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 5, characterized in that: The audio data collection rules specifically include: selecting the persons to be collected and recording representative Miao language pronunciations; the residence of the persons to be collected is a typical village where ethnic minorities gather; the representative Miao language pronunciations are that the collected words contain all initials, finals and tones and the pronunciation phonemes are evenly distributed; The audio recording rules specifically include: recording by different people, using ZOOM H6 as a recorder, a directional microphone as a recording device, a sampling rate of 16 bits, a frequency of 48kHz, and a recorder level of -3dB.

7. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to claim 1, characterized in that: The second mapping rule is determined based on a mapping table of Miao pinyin and international phonetic symbols.

8. The method for constructing a Qiandongnan Miao language triple database for speech recognition according to any one of claims 1 to 7, characterized in that: In the Miao language triple database, the data of the text part is stored in txt format, the data of the audio part is stored in wav format, and the data of the phoneme part is stored in txt format, wherein each data of the audio part is configured with two txt text annotation files, which respectively represent the Chinese annotation and the Miao language annotation.

9. A method for training a speech recognition model, wherein the speech recognition model comprises a language model, an acoustic model and a dictionary, characterized in that: The model training method comprises the following steps: Training the language model using a text portion of a Miao language triple database; Training the acoustic model using the audio portion of the Miao language tri-gram database; The dictionary is constructed using the text part and the corresponding factor part of the Miao language trigram database; Wherein, the Miao language triple database is obtained based on the construction method described in any one of claims 1-7.

10. A speech recognition method, characterized in that: The following steps are involved: Get the Miao language text to be processed; Inputting the Miao language text to be processed into a trained speech recognition model to obtain a Miao language speech recognition result; The speech recognition model is trained using a Miao language triple database, and the Miao language triple database is obtained based on the construction method described in any one of claims 1 to 7.