Vowel recovery method and device, electronic equipment and storage medium
By using a vowel recovery model to synchronously predict multi-note labels, the problem of rule conflicts in Hebrew vowel recovery is solved, achieving high-accuracy vowel recovery even without a dictionary. This method is applicable to multilingual mixed texts and texts with special symbols.
Patent Information
- Application Number
- CN202511925907.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-12-19
AI Technical Summary
Existing technologies, when recovering Hebrew vowels, often employ a serial prediction method, which can easily lead to rule conflicts when a single character has multiple diacritics, resulting in inaccurate recovered vowels.
A vowel recovery model is used for synchronous prediction of altered note labels. The model learns rules autonomously through training data and combines character samples and three types of altered note labels to avoid rule conflicts and improve accuracy.
It enables accurate recovery of Hebrew vowels even without a dictionary, improving the accuracy and applicability of vowel recovery. It is suitable for multilingual mixed texts and texts with special symbols, covering application scenarios such as news, technical documents, and social media.
Smart Images

Figure CN121393418A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and in particular to a vowel restoration method and device, electronic equipment and storage medium. BACKGROUND
[0002] At present, Hebrew adopts a full letter writing form, however, this form omits the diacritics appearing in full diacritics or dotted variants, which can cause pronunciation errors and semantic ambiguities of words.
[0003] In order to avoid pronunciation errors and semantic ambiguities of words, it is necessary to annotate vowels of Hebrew to ensure accurate pronunciation of words and avoid semantic ambiguities.
[0004] From the perspective of vowel restoration, the restoration of a single word is relatively difficult, and usually needs to combine the context to add diacritics in the word without vowels, and then restore the actual complete spelling, and the related technical solution adopts a serial prediction mode for prediction, and for the case that a single character has multiple diacritics, rule conflicts are prone to occur, so that the restored vowels are inaccurate. SUMMARY
[0005] The present application provides a vowel restoration method, device, electronic equipment and storage medium to solve the problem that the related technical solution adopts a serial prediction mode for prediction, and for the case that a single character has multiple diacritics, rule conflicts are prone to occur, so that the restored vowels are inaccurate.
[0006] The present application provides a vowel restoration method, comprising the following steps: Obtaining a first to-be-processed text, the first to-be-processed text comprising a first text requiring vowel restoration; Inputting the first to-be-processed text into a vowel restoration model for multi-diacritic label synchronous prediction to obtain a prediction label output by the vowel restoration model, the prediction label being three types of diacritic labels corresponding to each character in the first text, the three types of diacritic labels comprising a first dot symbol label for consonant fricative, a second dot symbol label for correcting consonant sound value, and a symbol label for annotating vowels, the vowel restoration model being trained based on a character sample and three types of diacritic labels associated with the character sample; Determining a target text after vowel restoration based on the first to-be-processed text and the prediction label output by the vowel restoration model.
[0007] The present application provides a vowel restoration method, wherein the first to-be-processed text further comprises a second text not requiring vowel restoration. The first to-be-processed text is input into a vowel restoration model for multi-vowel symbol label synchronous prediction to obtain a predicted label output by the vowel restoration model, and the method comprises the following steps: The first text in the first to-be-processed text is kept unchanged, the second text in the first to-be-processed text is replaced with a third text, and a first to-be-processed text after replacement is obtained, wherein the third text is a text that does not affect vowel restoration; The first to-be-processed text after replacement is input into a vowel restoration model for multi-vowel symbol label synchronous prediction to obtain a predicted label output by the vowel restoration model; The target text after vowel restoration is determined based on the first to-be-processed text and the predicted label output by the vowel restoration model, and the method comprises the following steps: The third text in the first to-be-processed text after replacement is kept unchanged, the first text in the first to-be-processed text after replacement is replaced with the first text after vowel annotation, and a first to-be-processed text after annotation is obtained, wherein the first text after vowel annotation is obtained by annotating the first text in the first to-be-processed text after replacement with the predicted label output by the vowel restoration model; The first text after vowel annotation in the first to-be-processed text after annotation is kept unchanged, the third text in the first to-be-processed text after annotation is replaced with the second text to obtain the target text after vowel restoration.
[0008] In the vowel restoration method provided by the application, the second text comprises one or more of the following: Text in a language other than the first language, space, number, and punctuation mark; The first language is a language corresponding to the first text.
[0009] In the vowel restoration method provided by the application, the first to-be-processed text is input into a vowel restoration model for multi-vowel symbol label synchronous prediction to obtain a predicted label output by the vowel restoration model, and the method comprises the following steps: Each character in the first to-be-processed text is represented by a basic symbol to obtain a basic symbol string, wherein the basic symbol is a symbol constituting each character; The basic symbol string is split with a delimiter between different characters in the first to-be-processed text as a boundary to obtain a first sequence composed of a plurality of basic symbol subsequences; The first sequence is input into a vowel restoration model for multi-vowel symbol label synchronous prediction to obtain a predicted label output by the vowel restoration model.
[0010] In the vowel restoration method provided by the application, the vowel restoration model comprises an embedding layer, a first extraction layer, a second extraction layer, and an output layer. The embedding layer is configured to convert each of the basic symbol substrings into a corresponding text distribution feature, and the text distribution feature and a position code corresponding to each of the text distribution features are fused to obtain a first feature; The first extraction layer is configured to perform feature extraction on the first feature to obtain a first extraction result, wherein the first extraction result is subjected to an element repetition enhancement process to obtain an enhanced first extraction result, the padding bits in the enhanced first extraction result are subjected to a masking operation to obtain an updated first extraction result, and the updated first extraction result and an updated position code are fused to obtain a second feature; The second extraction layer is configured to perform feature extraction on the second feature to obtain a second extraction result; The output layer is configured to output a predicted label based on the second extraction result.
[0011] The element repetition enhancement process includes: Based on the first extraction result, each target basic symbol in each basic symbol substring at a character level is replaced by an enhanced target basic symbol to obtain an enhanced first extraction result; The target basic symbol is a basic symbol in the basic symbol substring other than the first basic symbol and the last basic symbol, and the enhanced target basic symbol is obtained by repeating embedding each basic symbol in the target basic symbol N times and splicing.
[0012] The vowel restoration method provided by the application includes the following steps: A training sample set is constructed, and the training sample set includes a plurality of character samples, and each character sample is associated with three types of diacritic labels; The character sample is input into the vowel restoration model to obtain a probability value of the vowel restoration model outputting the three types of diacritic labels associated with the character sample; The probability value and the three types of diacritic labels associated with the character sample are used to determine a total training loss; The trainable parameters of the vowel restoration model are updated based on the total training loss.
[0013] The vowel restoration device provided by the application includes the following modules: The obtaining module is configured to obtain a first to-be-processed text, and the first to-be-processed text includes a first text that needs to restore a vowel; a processing module configured to input the first text to be processed into a vowel restoration model to perform multi-allophone label synchronous prediction, and obtain a predicted label output by the vowel restoration model, the predicted label being three types of allophone labels corresponding to each character in the first text, the three types of allophone labels including a first diacritic label for a consonant fricative, a second diacritic label for a consonant phonetic value, and a diacritic label for a vowel, the vowel restoration model being trained based on a character sample and three types of allophone labels associated with the character sample; a restoration module configured to determine a target text with restored vowels based on the first text to be processed and the predicted label output by the vowel restoration model.
[0014] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the vowel restoration method according to any one of the above when executing the program.
[0015] The present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the vowel restoration method according to any one of the above.
[0016] The present application also provides a computer program product, which includes a computer program, and the computer program is executable on a processor to implement the vowel restoration method according to any one of the above.
[0017] The present application provides a vowel restoration method, device, electronic device and storage medium, for a first text to be processed including a first text needing to restore vowels, inputting the first text to be processed into a vowel restoration model to perform multi-allophone label synchronous prediction, and obtaining a predicted label output by the vowel restoration model, in the process, three types of allophone labels corresponding to each character in the first text can be synchronously predicted, avoiding conflicts between different allophone labels, so as to accurately predict the allophone label for each character in the first text to be processed, thereby improving the accuracy of vowel restoration. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0019] Figure 1 is one of the flowcharts of the vowel restoration method provided by the present application; Figure 2 is another flowchart of the vowel restoration method provided by the present application; Figure 3 is a structural diagram of the vowel restoration model provided by the application; Figure 4 is a flowchart of a first text processing process provided by the application; Figure 5 is a schematic block diagram of the vowel restoration device provided by the application; Figure 6 is a structural diagram of the electronic device provided by the application.
[0020] Reference signs: 501, acquisition module; 502, processing module; 503, restoration module; 610, processor; 620, communication interface; 630, memory; 640, communication bus. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0022] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of another identical element in the process, method, article or device comprising the element. The terms "upper", "lower" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. Unless otherwise specified and limited, the terms "mount", "connect", "connect" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0023] The terms "first", "second", and the like in the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in a "or" relationship.
[0024] The following will be described in conjunction with Figures 1-6 The vowel restoration method, device, electronic equipment and storage medium provided by the present application are described, aiming to solve the problem that the related technical solutions use a serial prediction method for prediction, and for the case that a single character has multiple diacritics, rule conflicts are prone to occur, making the restored vowel inaccurate.
[0025] Figure 1 is one of the flowcharts of the vowel restoration method provided by the present application, as Figure 1 shown, including but not limited to the following steps: Step 101, obtaining a first to-be-processed text, the first to-be-processed text including a first text requiring vowel restoration.
[0026] Among them, the first text can be a Hebrew text.
[0027] In some embodiments, the first text is a segment or all of the first to-be-processed text.
[0028] Step 102, inputting the first to-be-processed text to a vowel restoration model for multi-diacritic label synchronous prediction, obtaining a prediction label output by the vowel restoration model, the prediction label being a three-class diacritic label corresponding to each character in the first text, the three-class diacritic label including a first dot symbol label for consonant fricative, a second dot symbol label for correcting consonant sound value, and a symbol label for labeling vowel, the vowel restoration model being trained based on a character sample and three-class diacritic labels associated with the character sample.
[0029] Step 103, determining a target text after restoring the vowel based on the first to-be-processed text and the prediction label output by the vowel restoration model.
[0030] In this embodiment, for a first to-be-processed text including a first text requiring restored vowels, the first to-be-processed text is input into the vowel restoration model for multi-diacritic synchronous prediction, and a predicted label output by the vowel restoration model can be obtained. In this process, three types of diacritic labels corresponding to each character in the first text can be synchronously predicted, avoiding conflicts between different diacritic labels, thereby accurately predicting the diacritic label of each character in the first to-be-processed text, and improving the accuracy of vowel restoration.
[0031] The predicted label is three types of diacritic labels corresponding to each character in the first text, and the three types of diacritic labels include a first diacritic label for a consonant fricative, a second diacritic label for correcting a consonant phonetic value, and a diacritic label for labeling a vowel. The three types of diacritic labels one-to-one correspond to three types of diacritics used in Hebrew orthography.
[0032] Specifically, in Hebrew orthography, Hebrew uses three types of diacritics, namely, Sin / Shin dots (4 types of representation, used to distinguish consonant fricatives, dot symbols located at the upper right corner or upper left corner of the letter), Dagesh dots (3 types of representation, used to correct the phonetic value of a consonant, dot symbols located at the center of a consonant letter), and Niqqud symbols (16 types of representation, used to label vowels, located above or below the letter).
[0033] The first diacritic label for a consonant fricative is a Sin / Shin dot, the second diacritic label for correcting a consonant phonetic value is a Dagesh dot, and the diacritic label for labeling a vowel is a Niqqud symbol.
[0034] In the above embodiment, after obtaining the predicted label output by the vowel restoration model, the first text in the first to-be-processed text can be labeled according to the labeling manner of the three types of diacritics on the character, thereby obtaining a target text after the vowels are restored.
[0035] In the above embodiment, the vowel restoration model is trained based on character samples and three types of diacritic labels associated with the character samples, so that the process of obtaining the predicted label does not need to rely on the diacritic combination rules built in the dictionary, but learns the rules autonomously through the training data. Therefore, vowel restoration can be realized without a dictionary.
[0036] In some embodiments, the first to-be-processed text further includes a second text that does not require restored vowels; the first to-be-processed text is input into the vowel restoration model for multi-diacritic synchronous prediction, and a predicted label output by the vowel restoration model is obtained, including: keeping the first text in the first to-be-processed text unchanged, replacing the second text in the first to-be-processed text with the third text, obtaining the first to-be-processed text after replacement, the third text being text that does not affect the restoration of vowels; inputting the first to-be-processed text after replacement into the vowel restoration model for multi-vowel symbol label synchronous prediction to obtain predicted labels output by the vowel restoration model; determining the target text after restoration of vowels based on the first to-be-processed text and the predicted labels output by the vowel restoration model, comprising: keeping the third text in the first to-be-processed text after replacement unchanged, replacing the first text in the first to-be-processed text after replacement with the first text after annotation of vowels, obtaining the first to-be-processed text after annotation, the first text after annotation of vowels being obtained by annotating the first text in the first to-be-processed text after replacement with the predicted labels output by the vowel restoration model, keeping the first text after annotation of vowels in the first to-be-processed text after annotation unchanged, replacing the third text in the first to-be-processed text after annotation with the second text to obtain the target text after restoration of vowels.
[0037] In this embodiment, by replacing the second text that does not need to restore vowels with the third text that does not affect the restoration of vowels, the range of text in the first to-be-processed text that needs to restore vowels can be clearly determined. In this process, the vowel restoration model can focus on the text that needs to restore vowels, thereby reducing the redundancy of the prediction task.
[0038] For example, the first to-be-processed text is: (Hebrew "Genesis" + number "1" + English phrase).
[0039] In this embodiment, the "1" and "In the beginning" belong to the second text (no need to restore Hebrew vowels), and can be replaced with the third text "5" (uniform symbol of number) and "O" (illegal character MASK symbol).
[0040] In this embodiment, compared with shielding or deleting the second text, the embodiment of the present application can avoid information loss caused by shielding or deleting the second text, and ensure that the target text after restoration of vowels contains all the contents of the first to-be-processed text.
[0041] In the above embodiment, if the second text is directly input into the model, the model may learn incorrect character-vowel symbol associations, such as incorrect binding of English characters and Hebrew vowel symbols, which reduces the prediction accuracy. By replacing the second text with the third text, the non-target text is converted into a non-interference text known to the model and does not affect the training rules, thereby avoiding noise from damaging the model's learning of the target character-vowel symbol association.
[0042] Specifically, after the illegal character MASK in the second text is replaced by O, the word accuracy (WOR) of the vowel restoration model is improved by 20%. Especially when the test set contains a large number of mixed texts, the WOR can be improved from 87.41% to about 89%, close to the performance of the dictionary scheme.
[0043] In addition, for the multi-lingual mixed special symbol containing scene, by replacing the second text with the third text, the applicable scene of the present application can be expanded from pure Hebrew text to multi-lingual mixed text and special symbol containing text, covering more than 80% of real application scenarios such as news, technical documents, social media, and breaking through the scene limitation of traditional solutions.
[0044] For example, Hebrew news (containing the number 555 and English Breaking News), the text input into the model after replacement is , and after prediction, it is restored to the original number and English, and only the vowels of the Hebrew part are supplemented.
[0045] For example, Hebrew technical documents (containing English AI), after replacing AI with O O, the prediction is restored to retain AI and supplement Hebrew vowels.
[0046] Table 1 shows the replacement relationship of replacing the second text with the third text, as shown in Table 1: Table 1
[0047] In some embodiments, the second text includes one or more of the following: text in a language other than the first language, space, number, punctuation mark; wherein the first language is the language corresponding to the first text.
[0048] In some embodiments, the first to-be-processed text is input into the vowel restoration model for multi-vowel symbol label synchronous prediction to obtain the prediction label output by the vowel restoration model, including: representing each character in the first to-be-processed text with a basic symbol to obtain a basic symbol string, the basic symbol being a symbol constituting each character; splitting the basic symbol string with the separator between different characters in the first to-be-processed text as a boundary to obtain a first sequence composed of multiple basic symbol subsequences; inputting the first sequence into the vowel restoration model for multi-vowel symbol label synchronous prediction to obtain the prediction label output by the vowel restoration model.
[0049] In this embodiment, the basic symbols are the elements that make up the first text, which are equivalent to the radicals in Chinese characters. By representing each character in the first text to be processed with basic symbols, the vowel restoration model can learn the underlying rule of basic symbol combination → diacritic. In this process, the vowel restoration model accurately predicts the labels based on the basic symbols, avoiding misjudgment caused by the overall input at the character level.
[0050] In addition, by representing each character in the first text to be processed with basic symbols, a first sequence is finally obtained, which can unify the dimensions for the vowel restoration model to learn and reduce the cognitive load of the vowel restoration model.
[0051] Exemplarily, different forms of (with Dagesh / without Dagesh), after decomposition, the core basic symbols are the same, and the vowel restoration model can reuse the same set of radical → diacritic rules without separately learning the character rules of different forms.
[0052] In this process, the learning efficiency of the vowel restoration model for character features is increased by 35%, the number of training convergence steps is reduced from 140k steps to 100k steps, and the overfitting risk in low-resource scenarios is significantly reduced.
[0053] The delimiters (spaces, punctuation marks) in Hebrew text are like the word boundaries in Chinese text, naturally dividing semantic units (words, phrases), enabling the vowel restoration model to accurately capture the collaborative rules of basic symbols (radicals) within the same semantic unit (such as the association between the combination and pronunciation of "艹", "平", "木", "果" in the word "apple"), avoiding interference of basic symbols across semantic units (such as not wrongly associating the "木" in "apple" with the "艹" in "banana"), making the diacritic prediction more accurate, and solving the pain point of fuzzy semantic boundaries in long sequences in traditional solutions.
[0054] The length of the decomposed basic symbol substrings (radical groups) is more compact (corresponding to the basic symbol combination of the original word, just like an idiom is still grouped by words after being decomposed into radicals, with a controllable length), avoiding the redundant attention calculation caused by long character sequences + no clear grouping in traditional solutions and avoiding the problem of gradient decay caused by long sequences.
[0055] Among them, the basic symbols are the effective inputs of the vowel restoration model, which are represented by Letter-Symbols, and the effective outputs of the vowel restoration model include the first dot symbol labels (Sin-Symbols) for consonant fricatives, the second dot symbol labels (Dagesh-Symbols) for correcting consonant values, and the symbol labels (Niqqud-Symbols) for annotating vowels.
[0056] Among them, the first dot symbol labels for consonant fricatives can distinguish and For the second point symbol label used to correct the consonant phonetic value, the center point of some consonant pronunciation can be affected, and for the symbol label used to mark the vowel, it is used for all other diacritics. The combination of the above three diacritics determines the pronunciation.
[0057] Table 2 shows the specific relationship between the input data type, the output data type and the corresponding label of the vowel restoration model, as shown in Table 2: Table 2
[0058] In some embodiments, the vowel restoration model includes an embedding layer, a first extraction layer, a second extraction layer, and an output layer; The embedding layer is used to convert each basic symbol substring into a corresponding text distribution feature, and the text distribution feature and the position encoding corresponding to each text distribution feature are fused to obtain a first feature; The first extraction layer is used for feature extraction of the first feature to obtain a first extraction result, wherein the first extraction result is subjected to an element repetition enhancement processing to obtain an enhanced first extraction result, the padding bits in the enhanced first extraction result are subjected to a shielding operation to obtain an updated first extraction result, and the updated first extraction result and the updated position encoding are fused to obtain a second feature; The second extraction layer is used for feature extraction of the second feature to obtain a second extraction result; The output layer is used to output a predicted label based on the second extraction result.
[0059] In this embodiment, the double extraction layer scheme of the first extraction layer and the second extraction layer can focus on the local features of the basic symbol substring (radical group) by the first extraction layer, such as understanding the collocation rule of the radical in the same semantic unit (such as the local association of day + month to form Ming), and outputting the first extraction result. The second extraction layer captures global context association based on the enhanced second feature (such as the semantic association of Ming in tomorrow), solving the cross-substring feature dependency problem of long sentences and compound words.
[0060] In this process, the diacritic prediction accuracy of long words and compound words is further improved by 8%-10%, especially for Hebrew compound words This cross-substring diacritic collaborative prediction is more accurate, and the WOR is further broken through to about 90%.
[0061] In the above embodiment, by performing element repetition enhancement processing, the model can be forced to focus on the semantic integrity of the basic symbol combination, avoid semantic fragmentation caused by radical splitting, and significantly improve the long word segmentation accuracy.
[0062] In addition, the execution of the shielding operation can remove invalid placeholders, reduce noise interference, reduce prediction errors caused by padding bits, and improve the prediction stability of the vowel restoration model in long sequences and irregular texts to 99%. The vowel restoration model can adapt to scenarios where the lengths of basic symbol subsequences are inconsistent, and the position encoding after reintegration is updated to ensure the accuracy of the enhanced sequence position information.
[0063] In the above embodiments, the simplified modular design of the embedding layer, the first extraction layer, the second extraction layer, and the output layer can reduce the computing power consumption by 25% during training, increase the training speed on low-configuration devices by 30%, and improve the model debugging and iteration efficiency by 40% due to the structured process.
[0064] In some embodiments, the structure of the vowel restoration model is a Transformer Encoder structure, which can also be any one of a fully connected layer network, a convolutional neural network, a recurrent neural network, a long short-term memory neural network, a residual neural network, and an attention mechanism deep learning model.
[0065] In some embodiments, the structure of the first extraction layer and the second extraction layer is a Transformer Encoder structure, which uses a 6-head attention 4-layer Encoder structure, the optimizer Adam is set to 0.005, dropout=0.1 is performed on the downstream task, the batch-size is set to 256, and 140k steps of training are used as the current optimal model (State-of-the-Art Model, SOTA).
[0066] In some embodiments, the first extraction layer and the second extraction layer can be absolute position encoding, relative position encoding, or rotational position encoding.
[0067] In some embodiments, the element repetition enhancement processing includes: Based on the first extraction result, replacing each target basic symbol in each basic symbol subsequence at the character level with an enhanced target basic symbol to obtain an enhanced first extraction result; wherein the target basic symbol is a basic symbol in the basic symbol subsequence except the first and last basic symbols, and the enhanced target basic symbol is obtained by repeating embedding each basic symbol in the target basic symbol N times and concatenating.
[0068] For example, the basic symbol subsequence is “APPLE”, and the target basic symbol is “PPL” when N is 3. The enhanced target basic symbol is “PPPPPPLLL”.
[0069] In some embodiments, the vowel restoration model is trained based on the following manner: constructing a training sample set, the training sample set including a plurality of character samples, each character sample being associated with three types of diacritic labels; inputting the character samples into the vowel restoration model to obtain probability values of the three types of diacritic labels associated with the character samples output by the vowel restoration model; determining a total training loss based on the probability values and the three types of diacritic labels associated with the character samples; updating the trainable parameters of the vowel restoration model based on the total training loss.
[0070] In this embodiment, the unified target of the character samples is bound with three types of labels, the model is forced to learn the symbol co-occurrence constraint, which can improve the collaborative prediction accuracy of the three types of diacritic symbols by 15%-20% and reduce the symbol combination error rate by 60%, and completely solve the rule conflict problem caused by independent training.
[0071] In addition, the character is taken as the minimum unit instead of the word or the sentence, so that the limited low-resource labeled data can be decomposed into more independent training samples, and the waste of training data in the low-resource scene is avoided.
[0072] As shown in Figure 2 In the case where the first text to be processed is Hebrew, and the process of replacing the second text with the third text, the basic symbol transformation and the separator segmentation is taken as the preprocessing, the vowel restoration method mainly includes the following steps: Step 201, obtaining the original Hebrew text to be processed; Step 202, preprocessing the original Hebrew text in a character-level manner to distinguish between legal characters, illegal characters, numbers and punctuation, and performing index indexing operation thereon; Step 203, using the vowel restoration model to obtain the text distribution features of each character in the original Hebrew text to be processed and the feature distribution of the three types of diacritic labels; Step 204, determining the vowel restoration result corresponding to each character in the original Hebrew text to be processed based on the text distribution features of each character in the original Hebrew text to be processed and the learning rule of the feature distribution of the three types of diacritic labels.
[0073] The index indexing operation is a process of representing by basic symbols.
[0074] The input form of the preprocessed vowel-free text input into the vowel restoration model is shown in Table 3.
[0075] Table 3
[0076] In some embodiments, the accuracy is calculated with the real labeled data at the word level granularity, and two examples of the test set input and the predicted result are shown in Table 4. For example, the predicted result of "O OOOOOO" after the "I am Rich" in the table is masked; the real vowel recovery and the predicted result in the second sentence are different, and the word is recorded as a prediction error.
[0077] Table 4
[0078] In the present application, the statistical indicators of the prediction result of the vowel recovery model are word accuracy (WOR) and vocalization accuracy (VOC), the word accuracy refers to the part of the word without diacritic errors, and the vocalization accuracy refers to the part of the word that will not cause pronunciation errors regardless of any diacritic point errors, wherein the baseline effect is 67.08% (17076 / 25453 words).
[0079] In some embodiments, the training sample set is obtained by processing a 78,000 token-level diacritic corpus and a token-level fine pos corpus, wherein the pos corpus refers to a ready-made text collection in which each word in the text is labeled with a part-of-speech tag.
[0080] In some embodiments, as shown in Figure 3 , the initial model structure of the vowel recovery model includes an embedding layer, a first extraction layer, an output layer, and a full connection layer, wherein the first to-be-processed text (x1, x2, … xn) is input into the embedding layer, the output of the embedding layer is input into the first extraction layer after being embedded and positionally encoded, the output of the first extraction layer is input into the output layer after performing element repetition enhancement processing, the output of the output layer is input into the full connection layer, and the output predicted label is obtained.
[0081] In the case of using a model setting of 6 heads, 6 layers, and 768 hidden layer dimensions, the accuracy on the validation set is more than 90%, but the accuracy on the test set is only 67.08%, but from the loss performance, there is no obvious overfitting.
[0082] The best WOR of different configurations adjusting the hidden layer dimension to 192, 384, 512 and 1024 is 71.15%. At the same time, the number of heads and the number of layers are also modified, and after the model structure is adjusted to 6 heads 8 layers, 8 heads 6 layers and 8 heads 8 layers respectively using the 384 hidden layer dimension, the best WOR is 68.2%; the number of repeated embedding times, that is, the proportion, is adjusted, and the input data is expanded at the word level, and the WOR is obviously improved to 73% in the best 6 head 8 layer 384 node structure. After analyzing the model structure, considering the parameter size and performance of the model integrated deployment, the finally used structure is 6 head 4 layer 384 hidden layer dimension, and the repeated embedding is 3 times. After the input data passes through two layers of Encoder (that is, the extraction of the first extraction layer and the second extraction layer in the application), the dropout operation (that is, the operation performed by the output layer in the application) is performed to avoid overfitting caused by insufficient data.
[0083] The vowel restoration model provided in the application has a WOR of 87.41% (22250 / 25453 words) on a test set of 2124 sentences, which is 20% higher than the baseline effect of 67.08% (17076 / 25453 words) in accuracy. The WOR is close to the current optimal effect and the Morfix scheme using a word-level dictionary (89.43%) and the pre-training model (90.76%), but the cost and the number of model parameters are much lower than the above schemes.
[0084] In some embodiments, as shown in Figure 4 The first to-be-processed text is input into the vowel restoration model, and the target text after the vowels are restored is finally obtained. If the target text is a dictionary word, the phonemes corresponding to the target text are output according to the dictionary. If the target text is not a dictionary word, the C45 module is called to predict the speech corresponding to the target text, and the prediction result of the C45 module is input into the stress identification module. The stress identification module is used to mark the stress of the speech corresponding to the predicted target text to obtain the phonemes corresponding to the target text, wherein the C45 module is a speech prediction module.
[0085] It should be noted that the vowel restoration device provided by the application can execute the vowel restoration method of any of the above embodiments when actually running, and therefore the present embodiment will not be described in detail.
[0086] As shown in Figure 5 The vowel restoration device provided by the application comprises: The acquisition module 501 is configured to acquire a first to-be-processed text, and the first to-be-processed text comprises a first text in which vowels need to be restored. The processing module 502 is configured to input the first to-be-processed text into the vowel restoration model to perform multi-diacritic label synchronous prediction, to obtain a predicted label output by the vowel restoration model, the predicted label being three types of diacritic labels corresponding to each character in the first text, the three types of diacritic labels including a first dot symbol label for a consonant fricative, a second dot symbol label for correcting a consonant phonetic value, and a symbol label for labeling a vowel, and the vowel restoration model being trained based on a character sample and three types of diacritic labels associated with the character sample; The restoration module 503 is configured to determine a target text with restored vowels based on the first to-be-processed text and the predicted label output by the vowel restoration model.
[0087] In this embodiment, for the first to-be-processed text including the first text requiring vowel restoration, the first to-be-processed text is input into the vowel restoration model to perform multi-diacritic label synchronous prediction, and the predicted label output by the vowel restoration model can be obtained. In this process, three types of diacritic labels corresponding to each character in the first text can be synchronously predicted, conflicts between different diacritic labels are avoided, and diacritic labels for each character in the first to-be-processed text are accurately predicted, thereby improving the accuracy of vowel restoration.
[0088] Figure 6 is a structural schematic diagram of an electronic device provided by the present application, as Figure 6 shown, the electronic device can include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can invoke a logical instruction in the memory 630 to execute a vowel restoration method, the method including: obtaining a first to-be-processed text, the first to-be-processed text including a first text requiring vowel restoration; inputting the first to-be-processed text into a vowel restoration model to perform multi-diacritic label synchronous prediction, to obtain a predicted label output by the vowel restoration model, the predicted label being three types of diacritic labels corresponding to each character in the first text, the three types of diacritic labels including a first dot symbol label for a consonant fricative, a second dot symbol label for correcting a consonant phonetic value, and a symbol label for labeling a vowel, and the vowel restoration model being trained based on a character sample and three types of diacritic labels associated with the character sample; and determining a target text with restored vowels based on the first to-be-processed text and the predicted label output by the vowel restoration model.
[0089] Moreover, the logic instructions in the memory 630 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0090] In another aspect, the present application also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions that, when executed by a computer, cause the computer to perform the vowel restoration method provided by any of the above embodiments, the method comprising: obtaining a first to-be-processed text, the first to-be-processed text comprising a first text in which vowels need to be restored; inputting the first to-be-processed text into a vowel restoration model to perform multi-vowel symbol label synchronous prediction, to obtain a prediction label output by the vowel restoration model, the prediction label being three types of vowel symbol labels corresponding to each character in the first text, the three types of vowel symbol labels comprising a first dot symbol label for consonant fricative, a second dot symbol label for modifying consonant sound value, and a symbol label for labeling vowel, the vowel restoration model being trained based on character samples and the three types of vowel symbol labels associated with the character samples; and determining a target text with restored vowels based on the first to-be-processed text and the prediction label output by the vowel restoration model.
[0091] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the vowel restoration method provided by any of the above embodiments, the method comprising: obtaining a first to-be-processed text, the first to-be-processed text comprising a first text in which vowels need to be restored; inputting the first to-be-processed text into a vowel restoration model to perform multi-vowel symbol label synchronous prediction, to obtain a prediction label output by the vowel restoration model, the prediction label being three types of vowel symbol labels corresponding to each character in the first text, the three types of vowel symbol labels comprising a first dot symbol label for consonant fricative, a second dot symbol label for modifying consonant sound value, and a symbol label for labeling vowel, the vowel restoration model being trained based on character samples and the three types of vowel symbol labels associated with the character samples; and determining a target text with restored vowels based on the first to-be-processed text and the prediction label output by the vowel restoration model.
[0092] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0093] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of the embodiments or some parts of the embodiments.
[0094] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A vowel restoration method, characterized in that, include: Obtain the first text to be processed, which includes the first text in which vowels need to be recovered; The first text to be processed is input into the vowel recovery model for simultaneous prediction of multiple altered note labels, and the predicted labels output by the vowel recovery model are obtained. The predicted labels are three types of altered note labels corresponding to each character in the first text. The three types of altered note labels include a first dot symbol label for consonant fricatives, a second dot symbol label for correcting consonant pitch values, and a symbol label for marking vowels. The vowel recovery model is trained based on character samples and the three types of altered note labels associated with the character samples. The target text after vowel recovery is determined based on the first text to be processed and the predicted labels output by the vowel recovery model.
2. The vowel restoration method according to claim 1, characterized in that, The first text to be processed also includes a second text that does not require the recovery of vowels; The step of inputting the first text to be processed into the vowel recovery model for simultaneous prediction of variable note labels, and obtaining the predicted labels output by the vowel recovery model, includes: Keeping the first text in the first text to be processed unchanged, replacing the second text in the first text to be processed with the third text, and obtaining the replaced first text to be processed, wherein the third text is text that does not affect vowel recovery; The replaced first text to be processed is input into the vowel recovery model for synchronous prediction of variable note labels, and the predicted labels output by the vowel recovery model are obtained. The step of determining the target text after vowel recovery based on the first text to be processed and the predicted labels output by the vowel recovery model includes: Keeping the third text in the replaced first text to be processed unchanged, the first text in the replaced first text to be processed is replaced with the first text after the vowel is annotated, and the first text after the vowel is annotated is obtained by annotating the first text in the replaced first text to be processed with vowels using the predicted labels output by the vowel recovery model. Keeping the first text after the marked vowels in the first text to be processed unchanged, the third text in the first text to be processed after the marked vowels is replaced with the second text to obtain the target text after the vowels are restored.
3. The vowel restoration method according to claim 2, characterized in that, The second text includes one or more of the following: Non-native language text, spaces, numbers, and punctuation marks; Wherein, the first language is the language corresponding to the first text.
4. The vowel restoration method according to claim 1, characterized in that, The step of inputting the first text to be processed into the vowel recovery model for simultaneous prediction of variable note labels, and obtaining the predicted labels output by the vowel recovery model, includes: Each character in the first text to be processed is represented by a basic symbol to obtain a basic symbol string, wherein the basic symbols are the symbols that constitute each character; Using the delimiters between different characters in the first text to be processed as boundaries, the basic symbol string is split to obtain a first sequence composed of multiple basic symbol substrings; The first sequence is input into the vowel restoration model for synchronous prediction of variable note labels, and the predicted labels output by the vowel restoration model are obtained.
5. The vowel restoration method according to claim 4, characterized in that, The vowel restoration model includes an embedding layer, a first extraction layer, a second extraction layer, and an output layer; The embedding layer is used to convert each of the basic symbol substrings into a corresponding text distribution feature, and the first feature is obtained by fusing the text distribution feature and the positional encoding corresponding to each of the text features; The first extraction layer is used to extract features from the first feature to obtain a first extraction result. The first extraction result is enhanced after performing element repetition enhancement processing. The padding bits in the enhanced first extraction result are masked to obtain an updated first extraction result. The updated first extraction result and the updated position code are fused to obtain a second feature. The second extraction layer is used to extract features from the second feature to obtain the second extraction result; The output layer is used to output predicted labels based on the second extraction result.
6. The vowel restoration method according to claim 5, characterized in that, The element repetition enhancement process includes: Based on the first extraction result, the target basic symbol in each basic symbol substring at the character level is replaced with the enhanced target basic symbol to obtain the enhanced first extraction result; The target basic symbol is the basic symbol in the basic symbol substring except for the first and last basic symbols. The enhanced target basic symbol is obtained by repeatedly embedding each basic symbol in the target basic symbol N times and concatenating them.
7. The vowel restoration method according to any one of claims 1 to 6, characterized in that, The vowel restoration model was trained using the following method: Construct a training sample set, which includes multiple character samples, each of which is associated with three types of altered note labels; Input the character sample into the vowel recovery model and obtain the probability values of the three types of altered note labels associated with the character sample output by the vowel recovery model; The total training loss is determined based on the probability value and the three types of altered note labels associated with the character sample; The trainable parameters of the vowel recovery model are updated based on the total training loss.
8. A vowel restoration device, characterized in that, include: The acquisition module is used to acquire the first text to be processed, which includes the first text in which vowels need to be recovered; The processing module is used to input the first text to be processed into the vowel recovery model for synchronous prediction of multiple altered note labels, and obtain the predicted labels output by the vowel recovery model. The predicted labels are three types of altered note labels corresponding to each character in the first text. The three types of altered note labels include a first dot symbol label for consonant fricatives, a second dot symbol label for correcting consonant pitch values, and a symbol label for marking vowels. The vowel recovery model is trained based on character samples and the three types of altered note labels associated with the character samples. The recovery module is used to determine the target text after recovering vowels based on the first text to be processed and the predicted labels output by the vowel recovery model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the vowel restoration method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the vowel restoration method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Webpage content extraction method based on Markov random field
CN103309961A
Chinese address word segmentation and annotation method
CN104933024A
System and method for generating glyphs for unknown characters
CN1117160A
Arabic vowel recovery method and device, equipment and storage medium
CN113011135A
Punctuation point recovery method and device, computer equipment and storage medium
CN113822060A