Uygur language processing method and system based on Latin letters
A Uyghur language and processing method technology, which is applied in electronic digital data processing, natural language data processing, neural learning methods, etc., and can solve the problems of lack of effective data samples, inability to form Uyghur lexical features, and accurate expressions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Publication Date
- 2020-07-17
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention relates to the technical field of natural language recognition, in particular to a method and system for processing Uighur language based on Latin letters. Background technique
[0002] In the prior art, the semantic processing of human natural language is usually performed by training a language model, and a good language model can greatly improve the processing accuracy of natural language. When using the bytes-pair-encoding algorithm, there will be technical low word frequency missing in the formed corpus dictionary. The Word2Vec algorithm is used to generate a static word vector of a specified dimension for each word, which reflects the hidden features of each word through the richness of the dimension, but it is prone to OOV (Out-of-vocabulary) problems due to the influence of the lexicon capacity. This type of model facilitates the development of natural language semantic processing tasks, but the disadvantage is that it ignores wo...
Examples
Embodiment Construction
[0041] In order to make the purpose, technical solution and advantages of the present invention clearer and clearer, the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Apparently, the described embodiments are only some of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0042] One embodiment of the present invention is based on the Uighur language processing method of Latin letters such as figure 1 shown. exist figure 1 , this example includes:
[0043] Step 100: Establish an alphabetic index of the Uyghur corpus, form the basic vectors of the Uyghur corpus according to the alphabetic index, and use the basic vectors to form a Uyghur sentence training set.
[0044] The Uyghur corp...