Text processing method and device, electronic equipment and computer readable storage medium

By obtaining the target pronunciation text, multiple preset vocabulary and pre-constructed vocabulary, using pinyin similarity to filter the synonyms collection, establishing the association relationship between the preset vocabulary and the collection of pronunciation and close words, and using multiple association relationships to correct the errors of the target pronunciation text, solving the problem of inaccurate speech text recognition, and achieving improved accuracy and automation of the pronunciation text.

CN120180141APending Publication Date: 2025-06-20GUANGZHOU SHIRONG INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311768380.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the after-sales customer service staff recognizes voice calls as text, they are susceptible to the customer's accent and dialect, resulting in the identified voice text inaccurate.

Method used

By obtaining the target pronunciation text, multiple preset vocabulary and pre-constructed vocabulary, filtering the synonym set using pinyin similarity, establishing the association relationship between the preset vocabulary and the collection of pronunciation and pronunciation, and using multiple association relationships to correct the target pronunciation text.

Benefits of technology

It achieves improved accuracy of speech text, high degree of automation, no manual participation, and saves labor and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180141A_ABST
    Figure CN120180141A_ABST
Patent Text Reader

Abstract

The invention provides a text processing method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a plurality of preset vocabularies and a pre-constructed lexicon comprising a plurality of synonym sets; for each preset vocabulary in the plurality of preset vocabularies, determining a synonym set in which words which are mutually synonyms with the preset vocabularies in the word bank are located as a target synonym set; determining pinyin similarity between the preset vocabulary and each word included in the target synonym set; filtering the target synonym set according to the pinyin similarity to obtain a synonym set; establishing an association relationship between the preset vocabularies and the near-pronunciation word set to obtain a plurality of association relationships, each association relationship being associated with one preset vocabulary; and performing error correction on the target voice text by adopting the plurality of association relationships. According to the invention, error correction can be carried out on the voice text obtained by voice information recognition, and the accuracy of voice text description is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a text processing method, apparatus, electronic device, and computer-readable storage medium in the field of artificial intelligence technology. Background Art

[0002] Currently, in order to accurately understand the after-sales needs of customers during work, after-sales customer service staff usually record calls during the call. After the call ends, the recorded call voice is recognized as text, so as to record the after-sales needs of customers. Since the Mandarin of the customers answered by the after-sales customer service staff is not necessarily standard, when recognizing the call voice as text, it is easily affected by factors such as the customer's accent and dialect, resulting in inaccurate recognized voice text.

[0003] Therefore, how to recognize voice information as accurate voice text has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides a text processing method, apparatus, electronic device, and computer-readable storage medium. This method can correct the voice text obtained by recognizing voice information and improve the accuracy of the description of the voice text.

[0005] In a first aspect, a text processing method is provided. The text processing method includes: obtaining a target voice text, a plurality of preset words, and a pre-constructed thesaurus; wherein, the thesaurus includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other; for each preset word in the plurality of preset words, determining the synonym set where the word in the thesaurus that is a synonym of the preset word is located as the target synonym set; determining the phonetic similarity between the preset word and each word included in the target synonym set; filtering the target synonym set according to the phonetic similarity to obtain a set of words with similar sounds; establishing an association relationship between the preset word and the set of words with similar sounds to obtain a plurality of association relationships; wherein, each association relationship associates a preset word; and correcting the target voice text by using the plurality of association relationships.

[0006] In the above technical solution, in the embodiment of the present application, by obtaining a target speech text, a plurality of preset words, and a pre-constructed word library, the word library includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other. For each preset word in the plurality of preset words, the synonym set where the word in the word library that is a synonym of the preset word is located is determined as the target synonym set. The pinyin similarity between the preset word and each word included in the target synonym set is determined, and the target synonym set is filtered according to the pinyin similarity to obtain a set of words with similar sounds. An association relationship is established between the preset word and the set of words with similar sounds to obtain a plurality of association relationships. Among them, each association relationship associates a preset word. By using the plurality of association relationships to correct the target speech text, on the one hand, the automatic classification and mining of words with similar sounds for each preset word are realized, with a high degree of automation, no need for manual participation, and a large amount of labor cost and time cost are saved. On the other hand, it can correct the text recognized from the voice information. Since the words with similar sounds for each preset word are filtered based on the synonyms of each preset word, the words with similar sounds for each preset word are also its corresponding synonyms. When correcting the speech text through the association relationship between the preset word and its corresponding set of words with similar sounds, the preset word that can accurately replace the misrecognized word in the speech text can be determined, which is beneficial to improving the accuracy of the speech text description.

[0007] In combination with the first aspect, in some possible implementation manners, the text processing method further includes: splitting the sample text into a plurality of words; for each word in the plurality of words, inputting the word into a pre-constructed word vector model, and determining the first vector similarity between the first word vector of the word and the second word vectors of other words in the plurality of words by the word vector model to obtain a plurality of first vector similarities; dividing the word and the other words corresponding to the first vector similarities greater than the vector similarity threshold among the plurality of first vector similarities into a synonym set to obtain the word library.

[0008] In the above technical solution, the method of constructing the word library belongs to a training method for predicting context words based on core words. By training the sample text using the training method for predicting context words based on core words to construct the word library, the word vector similarity between each word and other words can be calculated, and it can be determined whether each word and other words are synonyms of each other according to the word vector similarity, which is beneficial to accurately realizing the classification of words.

[0009] Combined with the first aspect and the above implementation manners, in some possible implementation manners, the text processing method further includes: splitting the sample text into multiple words; determining a second vector similarity between a third word vector of the sample text and a fourth word vector of each word through a pre-constructed word vector model, to obtain a plurality of second vector similarities; dividing the words corresponding to the second vector similarities greater than a vector similarity threshold among the plurality of second vector similarities into a set of near-synonyms, so as to obtain a thesaurus.

[0010] In the above technical solution, the manner of constructing the thesaurus belongs to a training method for predicting a core word based on context sub-words. By training the sample text by using the training method for predicting a core word based on context sub-words, the construction of the thesaurus is realized, which is to calculate the word vector similarity between each word and the sample text, and determine the words that are near-synonyms of each other from multiple words according to the word vector similarity, and the classification of words can be quickly realized, which is beneficial to improving the construction efficiency of the thesaurus.

[0011] Combined with the first aspect and the above implementation manners, in some possible implementation manners, determining the pinyin similarity between a preset word and each word included in the target near-synonym set includes: determining the initial consonant similarity between the initial consonants in the pinyin of the preset word and the initial consonants in the pinyin of each word in the target near-synonym set; determining the final similarity between the finals in the pinyin of the preset word and the finals in the pinyin of each word in the target near-synonym set; determining the tone similarity between the tones in the pinyin of the preset word and the tones in the pinyin of each word in the target near-synonym set; and determining the pinyin similarity between the preset word and each word included in the target near-synonym set according to the initial consonant similarity, the final similarity, and the tone similarity.

[0012] In the above technical solution, by calculating the pinyin similarity through the initial consonant similarity, the final similarity, and the tone similarity, the words with similar pronunciations of the preset word can be accurately screened out from the near-synonyms corresponding to the preset word. When using the association relationship between the preset word and the set of words with similar pronunciations for speech text error correction, it is beneficial to accurately correct the errors existing in the speech text.

[0013] Combined with the first aspect and the above implementation manners, in some possible implementation manners, filtering the target near-synonym set according to the pinyin similarity to obtain a set of words with similar pronunciations includes: determining the target pinyin similarity as the pinyin similarities less than a pinyin similarity threshold among the pinyin similarities between the preset word and each word included in the target near-synonym set; and filtering the words corresponding to the target pinyin similarity in the target near-synonym set to obtain a set of words with similar pronunciations.

[0014] Combined with the first aspect and the above implementation manners, in some possible implementation manners, after establishing the association relationship between the preset vocabulary and the set of homophonic words to obtain multiple association relationships, the text processing method further includes: when the same word is included in the sets of homophonic words associated with at least two preset vocabulary respectively, determining the vector distance between the fifth word vector of the same word and the sixth word vectors of at least two preset vocabulary respectively, to obtain at least two vector distances; determining the minimum vector distance among the at least two vector distances as the first target distance, and determining the vector distances other than the first target distance among the at least two vector distances as the second target distance; retaining the same word in the set of homophonic words corresponding to the first target distance, and deleting the same word in the set of homophonic words corresponding to the second target distance.

[0015] In the above technical solution, it can be ensured that there is no intersection in each set of homophonic words, which is beneficial to improving the accuracy of text error correction.

[0016] Combined with the first aspect and the above implementation manners, in some possible implementation manners, after establishing the association relationship between the preset vocabulary and the set of homophonic words to obtain multiple association relationships, the text processing method further includes: for each preset vocabulary among the multiple preset vocabulary, when there is a word in the set of homophonic words associated with the preset vocabulary that is the same as the preset vocabulary, deleting the word that is the same as the preset vocabulary from the set of homophonic words associated with the preset vocabulary.

[0017] In the above technical solution, by deleting the word that is the same as the preset vocabulary from the set of homophonic words associated with the preset vocabulary, since the set of homophonic words associated with the multiple preset vocabulary does not include the preset vocabulary, when correcting the target speech text using the multiple association relationships, if some words in the target speech text are the preset vocabulary, they are correct and standard words, and there is no need to replace the correct and standard words in the target speech text, and only the misrecognized words need to be corrected, which is beneficial to improving the error correction efficiency of the speech text.

[0018] Combined with the first aspect and the above implementation manners, in some possible implementation manners, correcting the target speech text using multiple association relationships includes: performing word segmentation processing on the target speech text to obtain multiple target words to be modified; for each target word among the multiple target words to be modified, determining the set of homophonic words including the target word among the multiple association relationships as the target set of homophonic words; replacing the target word with the preset vocabulary associated with the target set of homophonic words.

[0019] In a second aspect, a text processing apparatus is provided. The text processing apparatus includes:

[0020] An acquisition module for acquiring a target speech text, a plurality of preset words, and a pre-constructed thesaurus; wherein, the thesaurus includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other;

[0021] A first screening module for, for each preset word among the plurality of preset words, determining the synonym set in which the words in the thesaurus that are synonyms of the preset word are located as the target synonym set;

[0022] A similarity calculation module for determining the pinyin similarity between the preset word and each word included in the target synonym set;

[0023] A second screening module for filtering the target synonym set according to the pinyin similarity to obtain a set of words with similar sounds;

[0024] A relationship construction module for establishing an association relationship between the preset word and the set of words with similar sounds to obtain a plurality of association relationships; wherein, each association relationship associates a preset word;

[0025] A text adjustment module for correcting the target speech text by using the plurality of association relationships.

[0026] Combined with the second aspect, in some possible implementation manners, the text processing device further includes:

[0027] A first construction unit for splitting the sample text into a plurality of words; for each word among the plurality of words, inputting the word into a pre-constructed word vector model, and determining the first vector similarity between the first word vector of the word and the second word vectors of other words among the plurality of words by the word vector model to obtain a plurality of first vector similarities; dividing the word and the other words corresponding to the first vector similarities greater than the vector similarity threshold among the plurality of first vector similarities into a synonym set to obtain the thesaurus.

[0028] Combined with the second aspect and the above implementation manner, in some possible implementation manners, the text processing device further includes:

[0029] A second construction unit for splitting the sample text into a plurality of words; determining the second vector similarity between the third word vector of the sample text and the fourth word vectors of each word by a pre-constructed word vector model to obtain a plurality of second vector similarities; dividing the words corresponding to the second vector similarities greater than the vector similarity threshold among the plurality of second vector similarities into a synonym set to obtain the thesaurus.

[0030] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the similarity calculation module is specifically configured to determine the initial consonant similarity between the initial consonants in the pinyin of the preset vocabulary and the initial consonants in the pinyin of each word in the target near-synonym set; determine the final vowel similarity between the final vowels in the pinyin of the preset vocabulary and the final vowels in the pinyin of each word in the target near-synonym set; determine the tone similarity between the tones in the pinyin of the preset vocabulary and the tones in the pinyin of each word in the target near-synonym set; and determine the pinyin similarity between the preset vocabulary and each word included in the target near-synonym set according to the initial consonant similarity, the final vowel similarity, and the tone similarity.

[0031] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the second screening module is specifically configured to determine, as the target pinyin similarity, the pinyin similarity that is less than the pinyin similarity threshold among the pinyin similarities between the preset vocabulary and each word included in the target near-synonym set; and filter the words corresponding to the target pinyin similarity in the target near-synonym set to obtain a set of phonetically similar words.

[0032] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the text processing device further includes:

[0033] The first adjustment unit is configured to, when the same word is included in the sets of phonetically similar words associated with at least two preset vocabularies, determine the vector distances between the fifth word vector of the same word and the sixth word vectors of at least two preset vocabularies respectively, to obtain at least two vector distances; determine the minimum vector distance among the at least two vector distances as the first target distance, and determine the vector distances other than the first target distance among the at least two vector distances as the second target distances; retain the same word in the set of phonetically similar words corresponding to the first target distance, and delete the same word in the set of phonetically similar words corresponding to the second target distances.

[0034] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the text processing device further includes:

[0035] The second adjustment unit is configured to, for each of the multiple preset vocabularies, when there is a word identical to the preset vocabulary among the words included in the set of phonetically similar words associated with the preset vocabulary, delete the word identical to the preset vocabulary from the set of phonetically similar words associated with the preset vocabulary.

[0036] Combined with the second aspect and the above implementation manners, in some possible implementation manners, the text adjustment module is specifically configured to perform word segmentation on the target speech text to obtain multiple target vocabulary words to be modified; for each of the multiple target vocabulary words to be modified, determine the set of phonetically similar words including the target vocabulary word among the multiple association relationships as the target set of phonetically similar words; and replace the target vocabulary word with the preset vocabulary associated with the target set of phonetically similar words.

[0037] In a third aspect, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, so that the electronic device executes the text processing method in the above first aspect or any possible implementation manner of the first aspect.

[0038] In a fourth aspect, a computer program product is provided, which includes: computer program code, when the computer program code runs on a computer, it causes the computer to execute the text processing method in the above first aspect or any possible implementation manner of the first aspect.

[0039] In a fifth aspect, a computer-readable storage medium is provided, which stores computer program code, and when the computer program code runs on a computer, it causes the computer to execute the text processing method in the above first aspect or any possible implementation manner of the first aspect. Description of the Drawings

[0040] Figure 1 A schematic flowchart of a text processing method provided by an embodiment of the present application is shown.

[0041] Figure 2 An exemplary flowchart of constructing a thesaurus using a word vector model is shown;

[0042] Figure 3 Another exemplary flowchart of constructing a thesaurus using a word vector model is shown;

[0043] Figure 4 A schematic structural diagram of a text processing device provided by an embodiment of the present application is shown;

[0044] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present application is shown. Detailed Embodiments

[0045] The technical solutions in the present application will be clearly and elaborately described below with reference to the drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B: "and / or" in the text is only a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.

[0046] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or suggesting relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0047] On the one hand, the currently existing data classification methods include: 1) Manually collecting receipts and then manually sorting and classifying them. This method is too traditional and consumes a lot of manpower and time; 2) Using a model of the Transformers type for supervised training. For the training of a supervised model, a large amount of supervised training data is required. These supervised training data need to be manually sorted and labeled in advance, which also consumes a lot of manpower and time.

[0048] On the other hand, currently, in order to accurately understand the after-sales needs of customers during work, after-sales customer service staff usually record calls during the call. After the call ends, the recorded call voice is recognized into text to record the after-sales needs of customers. Since the Mandarin of the customers answered by the after-sales customer service staff is not necessarily standard, when recognizing the call voice into text, it is easily affected by factors such as the customer's accent and dialect, resulting in inaccurate recognized voice text.

[0049] Based on the above problems, the embodiments of the present application provide a text processing method, device, electronic device, and computer-readable storage medium. Each word in the standard and correct voice text is classified through an unsupervised word vector model to obtain multiple sets of synonyms. Each set of synonyms includes words corresponding to different semantics, and a thesaurus is formed by multiple sets of synonyms; furthermore, multiple different preset words are preset in advance. Each preset word is a correct and standard word used in daily life. For each preset word, the set of synonyms to which each preset word belongs is obtained in the thesaurus, and the set of synonyms to which each preset word belongs is filtered to obtain the set of homophonic words for each preset word. In this way, automatic classification and mining of the homophonic words of each preset word are realized, with a high degree of automation, no need for manual participation, and a large amount of manpower and time costs are saved. After obtaining the set of homophonic words for each preset word, an association relationship between the preset word and its corresponding set of homophonic words is formed. Thus, during application, the text obtained by recognizing the voice information is corrected based on the association relationship, which is beneficial to improving the accuracy of the voice text description.

[0050] The following is an embodiment of a text processing method provided by the embodiments of the present application.

[0051] Figure 1 Shows a schematic flowchart of a text processing method provided by the embodiments of the present application. Exemplarily, as Figure 1As shown in the figure, the text processing method provided by the embodiments of the present application is applied to an electronic device with computing power, such as a smart phone, a computer, etc. The text processing method includes the following solutions:

[0052] S110: Obtain a target speech text, a plurality of preset words, and a pre-constructed word library.

[0053] In an exemplary embodiment, the target speech text is obtained by identifying target speech information. The preset words are manually set and belong to the correct and standard words used in daily life, such as "screen mirroring", "large screen", "blue screen", etc. The word library is pre-constructed and includes a plurality of synonym sets. The semantics corresponding to each synonym set are different. Each synonym set includes a plurality of words, and any two words are synonyms of each other.

[0054] In a possible implementation, the text processing method further includes: training the sample text using an unsupervised word vector model to classify the words that are synonyms of each other in the sample text, thereby realizing the construction of the word library. Among them, the unsupervised word vector model can select models such as Word2Vec, FastText, GloVe, etc. Training the sample text using the unsupervised word vector model can learn the co-occurrence information of the words from the sample text, map the words to a continuous vector space, so that words that are semantically similar are close in this space.

[0055] Since in terms of semantic information capture ability, the Word2Vec model is a widely used word vector learning method, the purpose is to find a dense continuous vector space and map the words to this space, so that words that are semantically similar are close in this space. In Word2Vec, words that are semantically similar have similar local contexts. Therefore, Word2Vec can better capture the semantic relationship between words. In terms of computational efficiency, the Word2Vec model is trained based on a neural network language model. It has higher computational efficiency than traditional neural network language models, which benefits from the negative sampling method adopted by the Word2Vec model during the training process to reduce the solution difficulty, thereby improving the optimization speed and computational effect of the model. In summary, based on the comprehensive advantages in terms of semantic information capture ability, computational efficiency, parameter setting flexibility, and mature implementation and community support, etc., the present application selects the Word2Vec model as the unsupervised word vector model. During the entire process of mining phonetically similar words of proper nouns, the Word2Vec model will provide strong support for mining phonetically similar words associated with the preset words.

[0056] The construction of the word library includes the following solutions:

[0057] Split the sample text into multiple words;

[0058] For each of multiple words, input the word into a pre-constructed word vector model. The word vector model determines the first vector similarity between the first word vector of the word and the second word vectors of other words among the multiple words, obtaining multiple first vector similarities.

[0059] Divide the word and the other words corresponding to the first vector similarities greater than the vector similarity threshold among the multiple first vector similarities into a set of near-synonyms to obtain a thesaurus.

[0060] Obtaining a sample text includes: one way is to obtain a text that has been written, and perform noise filtering and typo correction on the written text to obtain the sample text; another way is to perform noise filtering and typo correction on the text recognized from voice information to obtain the sample text. Among them, the voice information is authorized communication voice, such as telephone voice, chat voice, Q&A voice generated on an online Q&A system, etc. recorded by customer service staff. By performing noise filtering and typo correction on the text that has been written and the text recognized from voice information, the quality of the sample text can be ensured, which is beneficial to improving the accuracy of subsequent homophone mining.

[0061] After obtaining the sample text, split the sample into multiple sentences, and perform preprocessing on each sentence to obtain multiple preprocessed sentences. The preprocessing of a sentence includes: for each sentence, delete the modal particles and delete the sentences that are too short. A sentence that is too short refers to a sentence whose word count length is less than a preset length.

[0062] After obtaining multiple preprocessed sentences, use the N-Gram word segmentation method to perform word segmentation on all preprocessed sentences to obtain multiple words. For each of the multiple words, input each word into a pre-constructed word vector model. The word vector model divides the words that are near-synonyms among the multiple words into a set of near-synonyms to obtain a thesaurus.

[0063] Specifically, as Figure 2 shown, Figure 2 shows an exemplary flowchart of constructing a thesaurus using a word vector model. The word vector model includes a first embedding layer and a second embedding layer. Figure 2The shown training method belongs to the training method of predicting context words based on core words. For each word W(t) among multiple words, the word W(t) is input into the word vector model. The first embedding layer extracts the first word vector of the word W(t) and the second word vectors of other words (excluding the word W(t)), and calculates the first vector similarity between the first word vector and the second word vectors of other words, obtaining multiple first vector similarities. The second embedding layer determines the first vector similarities greater than the vector similarity threshold from the multiple first vector similarities. The other words corresponding to the first vector similarities greater than the vector similarity threshold among the multiple first vector similarities are the words that are synonyms of the word W(t). The second embedding layer outputs the other words corresponding to the first vector similarities greater than the vector similarity threshold. For example Figure 2 the words W(t-1), W(t-2), W(t+1), W(t+2), etc. output by the second embedding layer in Figure 2 are all synonyms of the word W(t) and are also the context words of the word W(t).

[0064] After obtaining the other words corresponding to the first vector similarities greater than the vector similarity threshold, the word W(t) and the other words corresponding to the first vector similarities greater than the vector similarity threshold are divided into a synonym set, thereby obtaining multiple synonym sets corresponding to different semantics, and then a thesaurus is generated according to the multiple synonym sets.

[0065] By training the sample text using the training method of predicting context words based on core words to construct the thesaurus, the word vector similarity between each word and other words can be calculated, and it can be judged whether each word and other words are synonyms according to the word vector similarity, which is beneficial to accurately classify the words.

[0066] In a possible implementation, an unsupervised word vector model is used to train the sample text to classify the words that are synonyms in the sample text, thereby constructing the thesaurus. There is also another solution:

[0067] The sample text is split into multiple words;

[0068] The second vector similarity between the third word vector of the sample text and the fourth word vector of each word is determined through a pre-constructed word vector model, obtaining multiple second vector similarities;

[0069] The words corresponding to the second vector similarities greater than the vector similarity threshold among the multiple second vector similarities are divided into a synonym set to obtain the thesaurus.

[0070] The method of tokenizing the sample text here is the same as that of the above-mentioned sample text tokenization, so it will not be elaborated here.

[0071] Specifically, as Figure 3 shown, Figure 3 Another exemplary flowchart of constructing a thesaurus using a word vector model is shown. The word vector model includes a first embedding layer and a second embedding layer. Figure 3 The shown training method belongs to the training method of predicting the core word through context sub-words. Multiple words are output to the first embedding layer. The first embedding layer extracts the word vectors of each of the multiple words, and sums the word vectors of each of the multiple words to obtain a vector sum, which is used to represent the third word vector of the sample text. Figure 3 Words such as W(t - 1), W(t - 2), W(t + 1), W(t + 2) in are all multiple words input to the first embedding layer, and words such as W(t - 1), W(t - 2), W(t + 1), W(t + 2) are context words. The first embedding layer extracts the word vectors of words such as W(t - 1), W(t - 2), W(t + 1), W(t + 2), and performs vector summation to obtain a fusion vector, which is the third word vector of the sample text. After obtaining the third word vector of the sample text, calculate the second vector similarity between the third word vector and the fourth word vector of each word to obtain multiple second vector similarities. The second embedding layer determines the second vector similarities greater than the vector similarity threshold from the multiple second vector similarities. The words corresponding to the second vector similarities greater than the vector similarity threshold among the multiple second vector similarities are words that are synonyms of each other. The second embedding layer outputs the words corresponding to the second vector similarities greater than the vector similarity threshold. For example Figure 3 the words output by the second embedding layer in that are corresponding to the second vector similarities greater than the vector similarity threshold are the word W(t).

[0072] After obtaining the words corresponding to the second vector similarities greater than the vector similarity threshold, divide the words corresponding to the second vector similarities greater than the vector similarity threshold into a set of synonyms, so as to obtain multiple sets of synonyms corresponding to different semantics, and then generate a thesaurus according to the multiple sets of synonyms.

[0073] By training the sample text using the training method of predicting the core word through context sub-words to realize the construction of the thesaurus, which belongs to calculating the word vector similarity between each word and the sample text, and determining the words that are synonyms of each other from multiple words according to the word vector similarity, the classification of words can be quickly realized, which is beneficial to improving the construction efficiency of the thesaurus.

[0074] S120: For each of the multiple preset words, determine the set of near-synonym words in the thesaurus that are near-synonyms of the preset word as the target near-synonym set.

[0075] After obtaining the multiple preset words, for each preset word, calculate the similarity between the word vector of each preset word and the word vectors of each word in the thesaurus to obtain multiple similarities. The words corresponding to the similarities greater than the similarity threshold among the multiple similarities are the near-synonyms of the preset word, that is, the words corresponding to the similarities greater than the similarity threshold among the multiple similarities are determined as the words that are near-synonyms of the preset word. Then, determine the set of near-synonym words where the words that are near-synonyms of the preset word are located as the target near-synonym set, that is, each word in the target near-synonym set is a near-synonym of the preset word.

[0076] S130: Determine the pinyin similarity between the preset word and each word included in the target near-synonym set.

[0077] After obtaining the target near-synonym set, calculate the pinyin similarity between the preset word and each word included in the target near-synonym set, that is, calculate the similarity between the pinyin of the preset word and the pinyin of each word included in the target near-synonym set.

[0078] In a possible implementation manner, the above determination of the pinyin similarity between the preset word and each word included in the target near-synonym set includes the following solutions:

[0079] Determine the initial consonant similarity between the initial consonants in the pinyin of the preset word and the initial consonants in the pinyin of each word in the target near-synonym set;

[0080] Determine the final similarity between the finals in the pinyin of the preset word and the finals in the pinyin of each word in the target near-synonym set;

[0081] Determine the tone similarity between the tones in the pinyin of the preset word and the tones in the pinyin of each word in the target near-synonym set;

[0082] According to the initial consonant similarity, final similarity, and tone similarity, determine the pinyin similarity between the preset word and each word included in the target near-synonym set.

[0083] Since the pinyin of a Chinese character includes an initial consonant, a final, and a tone, as shown in Table 1:

[0084] Table 1

[0085] Chinese character Pinyin Initial consonant Final Tone rare xi x i 1 meal fan4 f an 4

[0086] The calculation of the pinyin similarity specifically calculates the initial consonant similarity, final similarity, and tone similarity between pinyins, and then fuses the initial consonant similarity, final similarity, and tone similarity to obtain the pinyin similarity, thereby improving the accuracy of near-sound word mining.

[0087] After determining the set of target near-synonyms corresponding to each preset word, obtain the pinyin of the preset word and the pinyins of each word in the set of target near-synonyms, obtain the initial consonants, finals, and tones of the pinyin of the preset word, and obtain the initial consonants, finals, and tones of the pinyins of each word in the set of target near-synonyms.

[0088] For the initial consonants, calculate and determine the initial consonant similarity between the initial consonants in the pinyin of the preset word and the initial consonants in the pinyins of each word in the set of target near-synonyms, denoted as S1; for the finals, calculate the initial consonant similarity between the initial consonants in the pinyin of the preset word and the initial consonants in the pinyins of each word in the set of target near-synonyms, denoted as S2; for the tones, calculate the tone similarity between the tones in the pinyin of the preset word and the tones in the pinyins of each word in the set of target near-synonyms, denoted as S3. Among them, considering that some initial consonants have certain pronunciation similarities and some finals also have certain pronunciation similarities, for example, the pronunciations of g and k are similar, the pronunciations of b and p are similar, and the pronunciations of c and ch are similar. For these initial consonants and finals with similar pronunciations, the corresponding similarity values can be set manually in advance to generate a similarity value table for the initial consonants and finals with similar pronunciations. When the obtained pinyin includes these initial consonants and / or finals, the corresponding initial consonant similarity and / or final similarity can be obtained by querying the similarity value table.

[0089] After obtaining S1, S2, and S3, then fuse S1, S2, and S3 to obtain the pinyin similarity between the preset word and each word included in the set of target near-synonyms, denoted as S. For example, the calculation process of S is: S = S1 + S2 + S3, or S = w1 * S1 + w2 * S2 + w3 * S3, where w1 represents the weight of the initial consonants, w2 represents the weight of the finals, and w3 represents the weight of the tones. By calculating the pinyin similarity through the initial consonant similarity, final similarity, and tone similarity, the near-sound words of the preset word can be accurately screened out from the near-synonyms corresponding to the preset word. When using the association relationship between the preset word and the set of near-sound words for speech text error correction, it is beneficial to accurately modify the errors existing in the speech text.

[0090] In addition, in addition to calculating the pinyin similarity in the above manner, the present application can also calculate the pinyin similarity using the similarity algorithm DimSim.

[0091] S140: Filter the set of target near-synonyms according to the pinyin similarity to obtain a set of near-sound words.

[0092] After calculating the pinyin similarity between the preset vocabulary and each word included in the target set of near-synonyms, determine the words that are phonetically similar to the preset vocabulary from the target set of near-synonyms according to the pinyin similarity. Retain the words in the target set of near-synonyms that are phonetically similar to the preset vocabulary, and filter out the remaining words (the words that are not phonetically similar to the preset vocabulary), so as to filter the target set of near-synonyms and obtain a set of phonetically similar words. The words in the set of phonetically similar words are phonetically similar to the preset vocabulary. Phonetically similar words refer to words with similar pinyin. The phonetically similar words of the preset vocabulary in this application are filtered and obtained based on the near-synonyms of the preset vocabulary.

[0093] In a possible implementation manner, the above filtering the target set of near-synonyms according to the pinyin similarity to obtain a set of phonetically similar words includes the following solutions:

[0094] Determine the target pinyin similarity as the pinyin similarity less than the pinyin similarity threshold among the pinyin similarities between the preset vocabulary and each word included in the target set of near-synonyms;

[0095] Filter the words corresponding to the target pinyin similarity in the target set of near-synonyms to obtain a set of phonetically similar words.

[0096] After obtaining the pinyin similarities between the preset vocabulary and each word included in the target set of near-synonyms, obtain multiple pinyin similarities, compare the size relationship between each pinyin similarity and the pinyin similarity threshold, determine the pinyin similarity less than the pinyin similarity threshold among the multiple pinyin similarities as the target pinyin similarity, and filter the words corresponding to the target pinyin similarity from the target set of near-synonyms. The new target set of near-synonyms formed by the words retained in the target set of near-synonyms is the set of phonetically similar words.

[0097] S150: Establish an association relationship between the preset vocabulary and the set of phonetically similar words to obtain multiple association relationships; where each association relationship associates a preset vocabulary.

[0098] After obtaining the set of phonetically similar words for each preset vocabulary, associate each preset vocabulary with the corresponding set of phonetically similar words to obtain multiple association relationships. Each association relationship associates a preset vocabulary, as shown in Table 2:

[0099] Table 2

[0100] Preset vocabulary Set of words with similar pronunciations Association relationship C1 J1 C1-J1 C2 J2 C2-J2 ... ... ... Cn Jn Cn-Jn

[0101] In Table 2, C1-J1, C2-J2,..., Cn-Jn are the association relationships corresponding to C1, C2,..., Cn in sequence.

[0102] In a possible implementation, after establishing the association relationship between the preset vocabulary and the set of homophonic words to obtain multiple association relationships, the above text processing method further includes the following solution:

[0103] When the same word is included in the sets of homophonic words associated with at least two preset vocabularies respectively, determine the vector distances between the fifth word vector of the same word and the sixth word vectors of at least two preset vocabularies respectively, to obtain at least two vector distances;

[0104] Determine the minimum vector distance among the at least two vector distances as the first target distance, and determine the vector distances other than the first target distance among the at least two vector distances as the second target distances;

[0105] Retain the same word in the set of homophonic words corresponding to the first target distance, and delete the same word in the set of homophonic words corresponding to the second target distances.

[0106] Considering the situation where there is the same word in at least two sets of homophonic words among the multiple association relationships, for this situation, if the same word is not processed, when subsequently using the association relationship between the preset vocabulary and the set of homophonic words to correct the speech text, it will result in the description of the corrected speech text being incorrect after the correction of the speech text. Therefore, when the same word is included in the sets of homophonic words associated with at least two preset vocabularies respectively, determine the vector distances between the fifth word vector of the same word and the sixth word vectors of at least two preset vocabularies respectively, to obtain at least two vector distances. For example, the words included in the set of homophonic words 1 associated with the preset vocabulary "big screen" are respectively: "serious illness", "big flat", "big bottle", "big pie", "hall", "big flat land", "large-scale"; the words included in the set of homophonic words 2 associated with the preset vocabulary "Daxing" are respectively: "wake up", "big line", "big star", "Daxing", "great luck"; among them, the same word in the set of homophonic words 1 and the set of homophonic words 2 is "large-scale", then calculate the vector distance D1 between the fifth word vector of "large-scale" and the sixth word vector of "big screen", and calculate the vector distance D2 between the fifth word vector of "large-scale" and the sixth word vector of "Daxing".

[0107] After obtaining at least two vector distances, determine the minimum vector distance among the at least two vector distances as the first target distance, and determine the vector distances other than the first target distance among the at least two vector distances as the second target distances. For example, if D1 < D2, then D1 is the first target distance and D2 is the second target distance.

[0108] Furthermore, retain the same words in the set of near-sounding words corresponding to the first target distance, and delete the same words in the set of near-sounding words corresponding to the second target distance. For example, retain the word "large-scale" in the set of near-sounding words 1, and delete the word "large-scale" in the set of near-sounding words 2. Among them, retaining "large-scale" in the set of near-sounding words 1 means that "large-scale" is not only more semantically similar to other words in the set of near-sounding words 1, but also more similar in pinyin, thus ensuring that there is no intersection in each set of near-sounding words, which is conducive to improving the accuracy of text error correction.

[0109] In a possible implementation, after establishing the association relationship between the preset vocabulary and the set of near-sounding words to obtain multiple association relationships, the above text processing method further includes the following solution:

[0110] For each preset vocabulary among the multiple preset vocabularies, when there is a word identical to the preset vocabulary among the words included in the set of near-sounding words associated with the preset vocabulary, delete the word identical to the preset vocabulary from the set of near-sounding words associated with the preset vocabulary.

[0111] For example, one of the preset vocabularies is "screen mirroring", and the words included in the set of near-sounding words associated with "screen mirroring" are: "vote comment", "investment bank", "investment product", "support flat", "support flat", "head comment", "scalp", "investment merger", "head screen", "paint flat", "same screen", "screen mirroring". It can be seen that the word "screen mirroring" is included in all the words in this set of near-sounding words, so delete the word "screen mirroring" from this set of near-sounding words, thus ensuring that the set of near-sounding words associated with multiple preset vocabularies does not include the preset vocabulary. Since the set of near-sounding words associated with multiple preset vocabularies does not include the preset vocabulary, when correcting the target speech text using multiple association relationships, if some of the words in the target speech text are the preset vocabulary, they are correct and standard words, and there is no need to replace the correct and standard words in the target speech text. Just correct the misrecognized words, which is conducive to improving the error correction efficiency of the speech text.

[0112] S160: Correct the target speech text using multiple association relationships.

[0113] After obtaining multiple association relationships, apply these association relationships to correct the target speech text, that is, based on the target speech text and these association relationships, determine the misrecognized words from the target speech text, and replace the misrecognized words with the preset vocabulary corresponding to the association relationship, so as to realize the correction of the target speech text and modify the target speech text into a text with accurate description.

[0114] In a possible implementation, the above filtering of the target set of near-synonyms according to the pinyin similarity to obtain the set of near-sounding words includes the following solution:

[0115] Perform word segmentation on the target speech text to obtain multiple target words to be modified;

[0116] For each of the multiple target words to be modified, determine the set of phonetically similar words that include the target word in the multiple association relationships as the target set of phonetically similar words;

[0117] Replace the target word with the preset word associated with the target set of phonetically similar words.

[0118] The target speech text is obtained through target speech recognition. For example, the target speech is a phone call speech. Perform word segmentation on the target speech text to obtain multiple target words to be modified. For each target word, query the association relationships corresponding to the multiple preset words through the target word, so as to query the set of phonetically similar words that include the target word in the association relationships corresponding to the multiple preset words, and determine the set of phonetically similar words that include the target word as the target set of phonetically similar words. After obtaining the target set of phonetically similar words, according to the association, the preset word associated with the target set of phonetically similar words can be obtained, that is, replace the target word with the preset word associated with the target set of phonetically similar words, so as to realize the modification of the target word and complete the error correction of the target speech text. For example, if the target word is "head screen" and the preset word is "screen mirroring", and the set of phonetically similar words associated with "screen mirroring" includes "head screen", then replace "head screen" with "screen mirroring".

[0119] It should be noted that the text processing method provided in this application is not only applicable to the error correction of speech texts, but also applicable to the error correction of texts written manually.

[0120] In the embodiments of the present application, by obtaining a target speech text, a plurality of preset words, and a pre-constructed thesaurus, the thesaurus includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other. For each preset word among the plurality of preset words, the synonym set where the word in the thesaurus that is a synonym of the preset word is located is determined as the target synonym set. The pinyin similarity between the preset word and each word included in the target synonym set is determined, and the target synonym set is filtered according to the pinyin similarity to obtain a set of words with similar sounds. An association relationship is established between the preset word and the set of words with similar sounds to obtain a plurality of association relationships. Among them, each association relationship associates a preset word. By using the plurality of association relationships to correct errors in the target speech text, on the one hand, automatic classification and mining of words with similar sounds for each preset word are realized, with a high degree of automation, no need for manual participation, and a large amount of labor cost and time cost are saved. On the other hand, it can correct the text recognized from the voice information. Since the words with similar sounds for each preset word are filtered based on the synonyms of each preset word, the words with similar sounds for each preset word are also its corresponding synonyms. When correcting the speech text through the association relationship between the preset word and its corresponding set of words with similar sounds, the preset word that can accurately replace the misrecognized word in the speech text can be determined, which is beneficial to improving the accuracy of the speech text description.

[0121] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the present application.

[0122] Figure 4 The structural schematic diagram of a text processing device provided by the embodiments of the present application is shown. Exemplarily, as Figure 4 shown, the text processing device 400 includes:

[0123] An acquisition module 410, configured to acquire a target speech text, a plurality of preset words, and a pre-constructed thesaurus; wherein, the thesaurus includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other;

[0124] A first screening module 420, configured to, for each preset word among the plurality of preset words, determine the synonym set where the word in the thesaurus that is a synonym of the preset word is located as the target synonym set;

[0125] A similarity calculation module 430, configured to determine the pinyin similarity between the preset word and each word included in the target synonym set;

[0126] A second screening module 440, configured to filter the target synonym set according to the pinyin similarity to obtain a set of words with similar sounds;

[0127] A relationship construction module 450, configured to establish an association relationship between a preset vocabulary and a set of homophonic words to obtain a plurality of association relationships; wherein, each association relationship associates a preset vocabulary;

[0128] A text adjustment module 460, configured to correct an error in a target speech text by using a plurality of association relationships.

[0129] In a possible implementation manner, the text processing device 400 further includes:

[0130] A first construction unit, configured to split a sample text into a plurality of words; for each word in the plurality of words, input the word into a pre-constructed word vector model, and determine a first vector similarity between the first word vector of the word and the second word vectors of other words in the plurality of words by the word vector model, to obtain a plurality of first vector similarities; divide the word and other words corresponding to the first vector similarities greater than a vector similarity threshold in the plurality of first vector similarities into a set of near-synonyms to obtain a thesaurus.

[0131] In a possible implementation manner, the text processing device 400 further includes:

[0132] A second construction unit, configured to split a sample text into a plurality of words; determine a second vector similarity between a third word vector of the sample text and a fourth word vector of each word by a pre-constructed word vector model, to obtain a plurality of second vector similarities; divide the words corresponding to the second vector similarities greater than a vector similarity threshold in the plurality of second vector similarities into a set of near-synonyms to obtain a thesaurus.

[0133] In a possible implementation manner, the similarity calculation module 430 is specifically configured to determine a consonant similarity between the initials in the pinyin of a preset vocabulary and the initials in the pinyin of each word in a target set of near-synonyms; determine a vowel similarity between the finals in the pinyin of the preset vocabulary and the finals in the pinyin of each word in the target set of near-synonyms; determine a tone similarity between the tones in the pinyin of the preset vocabulary and the tones in the pinyin of each word in the target set of near-synonyms; and determine a pinyin similarity between the preset vocabulary and each word included in the target set of near-synonyms according to the consonant similarity, the vowel similarity, and the tone similarity.

[0134] In a possible implementation manner, the second screening module 420 is specifically configured to determine, as a target pinyin similarity, the pinyin similarities less than a pinyin similarity threshold among the pinyin similarities between a preset vocabulary and each word included in a target set of near-synonyms; and filter the words corresponding to the target pinyin similarities in the target set of near-synonyms to obtain a set of homophonic words.

[0135] In a possible implementation manner, the text processing device 400 further includes:

[0136] A first adjustment unit is configured to, when the same word is included in the sets of homophonic words associated with at least two preset words respectively, determine the vector distances between the fifth word vector of the same word and the sixth word vectors of the at least two preset words respectively, so as to obtain at least two vector distances; determine the minimum vector distance among the at least two vector distances as the first target distance, and determine the vector distances other than the first target distance among the at least two vector distances as the second target distances; retain the same word in the set of homophonic words corresponding to the first target distance, and delete the same word in the set of homophonic words corresponding to the second target distances.

[0137] In a possible implementation manner, the text processing device further includes:

[0138] A second adjustment unit is configured to, for each preset word among a plurality of preset words, when there is a word identical to the preset word among the words included in the set of homophonic words associated with the preset word, delete the word identical to the preset word from the set of homophonic words associated with the preset word.

[0139] In a possible implementation manner, the text adjustment module 460 is specifically configured to perform word segmentation on the target speech text to obtain a plurality of target words to be modified; for each target word among the plurality of target words to be modified, determine the set of homophonic words including the target word among the plurality of association relationships as the target set of homophonic words; and replace the target word with the preset word associated with the target set of homophonic words.

[0140] It should be noted that when the text processing device provided in the above embodiments executes the text processing method, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the text processing device provided in the above embodiments and the embodiments of the text processing method belong to the same concept. Therefore, for the details not disclosed in the embodiments of the device of the present application, please refer to the embodiments of the above text processing method of the present application, which will not be elaborated here.

[0141] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0142] Figure 5 FIG. shows a schematic structural diagram of an electronic device provided in an embodiment of the present application.

[0143] Exemplarily, such as Figure 5As shown in the figure, the electronic device 500 includes: a memory 501 and a processor 502. Among them, an executable program code 5011 is stored in the memory 501, and the processor 502 is used to call and execute the executable program code 5011 to execute a text processing method.

[0144] In this embodiment, the electronic device can be divided into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0145] In the case of dividing each functional module corresponding to each function, the electronic device can include: an acquisition module, a first screening module, a similarity calculation module, a second screening module, a relationship construction module, a text adjustment module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.

[0146] The electronic device provided in this embodiment is used to execute the above-mentioned text processing method, so it can achieve the same effect as the above implementation method.

[0147] In the case of adopting an integrated unit, the electronic device can include a processing module and a storage module. Among them, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute mutual program codes and data, etc.

[0148] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure content of the present application. The processor can also be a combination that realizes computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.

[0149] This embodiment also provides a computer-readable storage medium. Computer program code is stored in the computer-readable storage medium. When the computer program code runs on a computer, it causes the computer to execute the above-mentioned related method steps to implement a text processing method in the above embodiment.

[0150] This embodiment also provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the above-mentioned related steps to implement a text processing method in the above embodiment.

[0151] In addition, the electronic device provided in the embodiments of the present application may specifically be a chip, component, or module. The electronic device may include a processor and a memory connected to each other. The memory is used to store instructions. When the electronic device runs, the processor may call and execute the instructions to enable the chip to execute a text processing method in the above embodiments.

[0152] Among them, the electronic device, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding text processing method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding text processing method provided above, which will not be elaborated here.

[0153] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0154] In the embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be electrical, mechanical, or other forms.

[0155] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A text processing method, characterized in that, The text processing method includes: Obtaining a target speech text, a plurality of preset words, and a pre-constructed thesaurus; wherein, the thesaurus includes a plurality of synonym sets, each synonym set includes a plurality of words, and any two words are synonyms of each other; For each preset word in the plurality of preset words, determining the synonym set where the words that are synonyms of the preset word in the thesaurus are located as the target synonym set; Determining the phonetic similarity between the preset word and each word included in the target synonym set; Filtering the target synonym set according to the phonetic similarity to obtain a set of phonetically similar words; Establishing an association relationship between the preset word and the set of phonetically similar words to obtain a plurality of association relationships; wherein, each association relationship associates a preset word; Using the plurality of association relationships to correct errors in the target speech text.

2. The text processing method according to claim 1, characterized in that, The text processing method further includes: Splitting a sample text into a plurality of words; For each word in the plurality of words, inputting the word into a pre-constructed word vector model, and determining the first vector similarity between the first word vector of the word and the second word vectors of other words in the plurality of words by the word vector model to obtain a plurality of first vector similarities; Dividing the word and the other words corresponding to the first vector similarities greater than the vector similarity threshold in the plurality of first vector similarities into a synonym set to obtain the thesaurus.

3. The text processing method according to claim 1, characterized in that, The text processing method further includes: Splitting a sample text into a plurality of words; Determining the second vector similarity between the third word vector of the sample text and the fourth word vector of each word by a pre-constructed word vector model to obtain a plurality of second vector similarities; Dividing the words corresponding to the second vector similarities greater than the vector similarity threshold in the plurality of second vector similarities into a synonym set to obtain the thesaurus.

4. The text processing method according to claim 1, characterized in that, The determining the phonetic similarity between the preset word and each word included in the target synonym set includes: Determining the initial similarity between the initials in the pinyin of the preset word and the initials in the pinyin of each word in the target synonym set; Determining the final similarity between the finals in the pinyin of the preset word and the finals in the pinyin of each word in the target synonym set; Determining the tone similarity between the tones in the pinyin of the preset word and the tones in the pinyin of each word in the target synonym set; Determining the phonetic similarity between the preset word and each word included in the target synonym set according to the initial similarity, the final similarity, and the tone similarity.

5. The text processing method according to claim 1, characterized in that, The filtering the target synonym set according to the phonetic similarity to obtain a set of phonetically similar words includes: Determining the target phonetic similarity as the phonetic similarities less than the phonetic similarity threshold among the phonetic similarities between the preset word and each word included in the target synonym set; Filtering the words corresponding to the target phonetic similarity in the target synonym set to obtain the set of phonetically similar words.

6. The text processing method according to any one of claims 1 to 5, characterized in that, After establishing the association relationship between the preset vocabulary and the set of phonetically similar words to obtain multiple association relationships, the text processing method further includes: When the same word is included in the sets of phonetically similar words associated with at least two preset vocabularies, determining the vector distances between the fifth word vector of the same word and the sixth word vectors of the at least two preset vocabularies respectively, to obtain at least two vector distances; Determining the minimum vector distance among the at least two vector distances as the first target distance, and determining the vector distances other than the first target distance among the at least two vector distances as the second target distances; Retaining the same word in the set of phonetically similar words corresponding to the first target distance, and deleting the same word in the set of phonetically similar words corresponding to the second target distances.

7. The text processing method according to any one of claims 1 to 5, characterized in that, After establishing the association relationship between the preset vocabulary and the set of phonetically similar words to obtain multiple association relationships, the text processing method further includes: For each preset vocabulary among the multiple preset vocabularies, when there is a word identical to the preset vocabulary among the words included in the set of phonetically similar words associated with the preset vocabulary, deleting the word identical to the preset vocabulary from the set of phonetically similar words associated with the preset vocabulary.

8. The text processing method according to any one of claims 1 to 5, characterized in that, The correcting the target speech text by using the multiple association relationships includes: Performing word segmentation on the target speech text to obtain multiple target words to be modified; For each target word among the multiple target words to be modified, determining the set of phonetically similar words including the target word among the multiple association relationships as the target set of phonetically similar words; Replacing the target word with the preset vocabulary associated with the target set of phonetically similar words.

9. A text processing device, characterized in that, The text processing apparatus includes: An acquisition module, configured to acquire a target speech text, multiple preset vocabularies, and a pre-constructed thesaurus; wherein, the thesaurus includes multiple sets of synonyms, each set of synonyms includes multiple words, and any two words are synonyms of each other; A first screening module, configured to, for each preset vocabulary among the multiple preset vocabularies, determine the set of synonyms where the words that are synonyms of the preset vocabulary in the thesaurus are located as the target set of synonyms; A similarity calculation module, configured to determine the phonetic similarity between the preset vocabulary and each word included in the target set of synonyms; A second screening module, configured to filter the target set of synonyms according to the phonetic similarity to obtain a set of phonetically similar words; A relationship construction module, configured to establish an association relationship between the preset vocabulary and the set of phonetically similar words to obtain multiple association relationships; wherein, each association relationship associates a preset vocabulary; A text adjustment module, configured to correct the target speech text by using the multiple association relationships.

10. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable program code; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the text processing method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed, implements the text processing method according to any one of claims 1 to 8.