Training and recognition methods for a text-oriented Cantonese recognition model and system

By building a Cantonese recognition model for shallow networks, combined with improving stop word list, rule matching and simple and traditional Chinese recognition methods, the problems of high data set requirements and large common words in the prior art are solved, and Cantonese recognition with high accuracy is achieved.

CN114065749BActive Publication Date: 2025-06-13INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111332368.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-11
Publication Date
2025-06-13
Estimated Expiration
2041-11-11

AI Technical Summary

Technical Problem

When the prior art recognizes Mandarin and Cantonese, the data set requirements are high, there are many common words, and the characteristics of Cantonese texts are not effectively utilized, resulting in low recognition accuracy.

Method used

A shallow network is used to build a Cantonese recognition model, and the training data set is obtained by improving stop word list filtering and word segmentation processing, and the Fasttext shallow network is trained to convergence. At the same time, the design rule matching method and the simplified and traditional Chinese recognition method are combined with the Cantonese recognition model to improve the recognition accuracy.

Benefits of technology

It realizes the accurate distinction between Cantonese and Mandarin, reduces the dependence on the quality of the data set, improves the recognition accuracy and reliability, and is simple to operate and fast detection speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065749B_ABST
    Figure CN114065749B_ABST
Patent Text Reader

Abstract

The present invention provides a training method for a text-oriented Cantonese recognition system. The method includes: A1. Obtain Cantonese and Mandarin text corpora, manually annotate the language of the corpora to obtain an annotated data set, filter the annotated data set using an improved stop word list and perform word segmentation to obtain a training data set; A2. Utilize the training data set obtained in step A1 to train a shallow network until convergence to obtain a Cantonese recognition model; A3. Construct a Cantonese feature word list, use the corpora in the training data set obtained in step A1 as input and the judgment result of whether the corpus is in Cantonese as output, and construct a rule matching model for retrieving whether the corpus hits the Cantonese feature word list based on the Cantonese feature word list; A4. Construct a simplified and traditional Chinese recognition model with the corpora in the training data set obtained in step A1 as input and the judgment result of whether the corpus is in traditional Chinese as output; A5. Train a fusion module with the outputs of the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing. Specifically, it relates to a training method and an identification method for a Cantonese recognition model for text. Further, it also relates to a text-oriented Cantonese recognition system, a training method and an identification method for the Cantonese recognition system. Background Art

[0002] With the development of social software platforms, users in different regions are connected to each other, increasing the number of language types on social media. Cantonese, as the main language used in regions such as Guangdong, Hong Kong, and Macau, has gradually been widely used in social interactions with the increase in social platform users. The mixture of multiple languages makes it more difficult for social platforms to analyze and classify language materials. Therefore, it is particularly important to propose a set of language classification methods. By judging the language type of user language materials, it is more convenient to recommend relevant content in the same language to this user. Currently, companies such as iFlytek and Youdao.com have proposed methods that can recognize multiple national languages, but there is little research on Cantonese recognition methods, and the recognition accuracy is relatively low.

[0003] The existing technologies for distinguishing Mandarin from Cantonese mainly include the following three:

[0004] Technology 1: A technology for recognizing mixed Mandarin and Cantonese speech. For example, the technology mentioned in the Chinese patent application with the publication number "CN111816160A": "Training Method and System for Mixed Mandarin and Cantonese Speech Recognition Model". This technology uses a mixed training sample of multiple languages to train a multi-task model, reuses the network parameters of the multi-task model through data migration, and trains a mixed Mandarin and Cantonese recognition model based on mixed modeling of Mandarin and Cantonese. Since this technology uses the method of migrating multi-language model parameters, the network is overly dependent on the quality and size of the selected data set. When the proportion of Cantonese and Mandarin in the selected data set and the strength of the characteristic representing this type of language are different, different effects will occur. At the same time, a large amount of interference information such as numbers and characters in the data set will also cause deviations in the learned features.

[0005] Technology 2: A technology for analyzing corpus texts. For example, the technology mentioned in the Chinese patent application with the publication number "CN111160015A": "A method, device, computer storage medium, and terminal for realizing text analysis". This technology determines the language of a text by introducing dictionaries such as a Cantonese dictionary, a Simplified Chinese dictionary, and a Traditional Chinese dictionary, and comparing the ratio of the number of characters in each language in the text to be detected with a set threshold. Since this technology uses the method of querying dictionaries, many common words between Mandarin and Cantonese will be found, resulting in deviations when calculating the ratio. Moreover, Cantonese itself has many unique words. If only individual characters are queried and the overall vocabulary is split, statistical errors will occur. At the same time, a large number of dictionary query operations will result in a large computational amount.

[0006] Technology 3: A technology for recognition by training a neural network with voice data. For example, the technology mentioned in the Chinese patent application with the publication number "CN113282718A": "A language recognition method and system based on an adaptive center anchor". This technology trains a deep neural backbone network by using the features extracted from a voice dataset and further trains the deep neural backbone network using the adaptive center anchor method. The adaptive center anchor method refers to calculating the Euclidean distance between the output results of each language training set and its corresponding language feature center, constructing an Anchor set and a non - Anchor set based on the Euclidean distance; training the deep neural backbone network based on the Anchor set and the non - Anchor set, and continuously updating the feature center and the samples near the feature center to achieve the selection of the adaptive center anchor; repeating the above operations until the network converges, so as to better recognize the language to which the text belongs. This technology mainly extracts and trains the features of speech, but the features of text and speech are not the same, and the same method cannot be used to extract text features.

[0007] In summary, the existing technologies for recognizing Mandarin and Cantonese mainly have the following problems:

[0008] 1. The existing technologies for distinguishing Mandarin and Cantonese mainly distinguish from the perspective of speech corpora, without making full use of the features of text corpora, resulting in a low recognition rate in the text field and a high quality requirement for the corpus dataset.

[0009] 2. The number of common words between Mandarin and Cantonese in the existing Mandarin and Cantonese dictionaries cited is large, and the dictionary query volume is large, resulting in an inability to accurately distinguish Mandarin and Cantonese. Summary of the Invention

[0010] Therefore, the object of the present invention is to overcome the problems of the above-mentioned existing methods, such as high requirements for the data set, too many common words, and low accuracy due to the failure to make good use of the text features of Cantonese, and to provide a new training method and recognition method for a text-oriented Cantonese recognition model, as well as a text-oriented Cantonese recognition system, a training method and a recognition method for a Cantonese recognition system.

[0011] According to a first aspect of the present invention, there is provided a training method for a text-oriented Cantonese recognition model, the method comprising: S1, obtaining Cantonese and Mandarin text corpora, and manually annotating the language of the corpora to obtain an annotated data set; S2, combining the common words of Cantonese and Mandarin with the existing simplified Chinese stop word list to form an improved stop word list; S3, filtering the annotated data set in step S1 using the improved stop word list and performing word segmentation processing to obtain a training data set, and then training a shallow network with the corpora in the training data set as input and the recognition result of whether the corpus is Cantonese as output until convergence.

[0012] Preferably, the step S1 includes: S11, collecting Cantonese and Mandarin text corpora on Cantonese and Mandarin social platforms through web crawlers; S12, screening the texts in the collected corpora, removing texts that do not meet the requirements of the preset shortest text length, and splitting texts with a length greater than the preset longest text length; S13, manually annotating the screened texts to label the language of all texts as Cantonese or Mandarin.

[0013] In some embodiments of the present invention, the preset shortest text length is 4, and the preset longest text length is 100.

[0014] Preferably, the step S2 includes: S21, filtering the annotated data set using the simplified Chinese stop word list; S22, using jieba word segmentation in Python to divide each corpus in the filtered annotated data set, determining the association probability between different characters, and forming phrases by combining each character with the other character with the highest association probability with it to form a word segmentation result; S23, respectively counting the word frequencies of the word segmentation of Cantonese and Mandarin, obtaining the common words in the Cantonese word segmentation and Mandarin word segmentation that exceed the preset word frequency threshold, and combining them with the existing simplified Chinese stop word list to form an improved stop word list.

[0015] In some embodiments of the present invention, the preset word frequency threshold is 5000.

[0016] Preferably, the step S3 includes: S31, filtering the annotated data set using the improved stop word list and performing word segmentation processing to obtain a training data set; S32, introducing pre-trained word vectors, and training a Fasttext shallow network with the training data set until convergence.

[0017] According to a second aspect of the present invention, there is provided a text-oriented Cantonese recognition method, characterized in that the method includes: T1, obtaining a text to be processed; T2, using the Cantonese recognition model trained by the method described in the first aspect of the present invention to recognize whether the text to be processed is Cantonese.

[0018] According to a third aspect of the present invention, there is provided a text-oriented Cantonese recognition system, characterized in that the system includes: a Cantonese recognition model, which is trained by the method described in the first aspect of the present invention and is used to recognize whether the text to be processed is Cantonese based on the characteristics of the text to be processed to obtain a recognition result; a rule matching model, which is used to retrieve whether the text to be processed hits the Cantonese characteristic word list based on the Cantonese characteristic word list to obtain a judgment result on whether the text to be processed is Cantonese; a simplified and traditional Chinese recognition model, which is used to judge whether the text to be processed is traditional Chinese; and a fusion module, which is used to judge whether the text to be processed is Cantonese according to the recognition result of the Cantonese recognition model for the text to be processed, the judgment result of the rule matching model for the text to be processed, and the judgment result of the simplified and traditional Chinese recognition model for the text to be processed.

[0019] According to a fourth aspect of the present invention, there is provided a training method for the text-oriented Cantonese recognition system described in the third aspect of the present invention, the method including: A1, obtaining Cantonese and Mandarin text corpora, manually annotating the language types of the corpora to obtain an annotated data set, filtering the annotated data set using an improved stop word list and performing word segmentation to obtain a training data set; A2, using the training data set obtained in step A1, training a shallow network by the method described in the first aspect of the present invention until convergence to obtain a Cantonese recognition model; A3, constructing a Cantonese characteristic word list, using the corpora in the training data set obtained in step A1 as input and the judgment result of whether the corpora are Cantonese as output, and constructing a rule matching model for retrieving whether the corpora hit the Cantonese characteristic word list based on the Cantonese characteristic word list; A4, constructing a simplified and traditional Chinese recognition model with the corpora in the training data set obtained in step A1 as input and the judgment result of whether the corpora are traditional Chinese as output; A5, training the fusion module with the outputs of the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model.

[0020] Preferably, step A3 includes: constructing a Cantonese characteristic word list based on the Cantonese corpus, the different parts of the Cantonese stop word list and the simplified Chinese stop word list, and the Cantonese words in the training data set whose word frequencies exceed a preset word frequency threshold.

[0021] Preferably, step A4 includes: training the Hanzidentifier model with the corpora in the training data set as input and the judgment result of whether the corpora are Cantonese as output to obtain a simplified and traditional Chinese recognition model.

[0022] Preferably, the fusion module includes a linear perceptron, and step A5 includes: using the linear perceptron to perform model fusion on the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model; wherein, a three-dimensional vector set composed of the output results of the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model on the training data set is used to train the linear perceptron to obtain the perceptron model parameters to achieve model fusion, and the output of the softmax regression layer of the linear perceptron is used as the final recognition result.

[0023] According to the fifth aspect of the present invention, there is also provided a text-oriented Cantonese recognition method, characterized in that the method includes: F1. Obtain the text to be processed; F2. Use the Cantonese recognition system trained by the method described in the fourth aspect of the present invention to recognize whether the text to be processed is Cantonese.

[0024] Compared with the prior art, the advantages of the present invention are as follows:

[0025] 1. In the existing methods, it is difficult for the neural network to accurately capture the respective characteristics of Cantonese and Mandarin, resulting in low recognition accuracy. However, the present invention designs a Cantonese recognition model that can accurately distinguish Cantonese and Mandarin using a shallow network, has low requirements for the data set, and also improves the common vocabulary table, thereby improving the accuracy.

[0026] 2. In the existing methods, the feature that Cantonese has characteristic vocabulary is not utilized, resulting in low recognition accuracy and reliability. However, the present invention designs a set of rule matching methods to find whether the corpus has Cantonese characteristic words, thereby improving the accuracy and reliability of the determination.

[0027] 3. The existing methods do not utilize the feature that Cantonese itself has a large number of traditional Chinese characters. However, the present invention designs a set of simplified and traditional Chinese recognition methods, which are based on judging whether the corpus is traditional Chinese to distinguish between Mandarin and Cantonese. It is not only simple to operate but also improves the detection speed.

[0028] In addition, the existing methods do not consider the differences between Cantonese and Mandarin from multiple perspectives. However, the present invention performs model fusion by fusing models that recognize Cantonese and Mandarin from different perspectives, and simultaneously recognizes Cantonese and Mandarin from multiple perspectives, thereby improving the recognition accuracy. Description of the Drawings

[0029] The following further describes the embodiments of the present invention with reference to the accompanying drawings, wherein:

[0030] Figure 1 It is a schematic diagram of the training process of the Cantonese recognition model according to the embodiment of the present invention;

[0031] Figure 2 It is a schematic diagram of the Cantonese recognition system according to the embodiment of the present invention;

[0032] Figure 3 Schematic diagram of the training process of a Cantonese recognition system according to an embodiment of the present invention. Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] As mentioned in the background art, several existing technologies for recognizing Mandarin and Cantonese cannot be directly applied to text field recognition. Among them, the first technology is to adopt the method of migrating multi-lingual model parameters, which will make the network overly dependent on the quality and size of the selected data set. When the proportion of Cantonese and Mandarin in the selected data set and the strength of the characteristic representing this type of language are different, different effects will be caused. Interferences such as a large number of numbers and characters in the data set will also cause deviations in the learned features. The second technology is to adopt the method of querying dictionaries, which will query too many common words between Mandarin and Cantonese, resulting in deviations when calculating the ratio. Moreover, Cantonese itself has many unique words. If only a single character is queried and the overall vocabulary is split, statistical errors will occur. At the same time, a large number of dictionary query operations will increase the calculation amount. The third technology mainly extracts and trains speech features, but there are significant differences between text features and speech features, and the same method as speech cannot be used to extract text features.

[0035] To better understand the present invention, the present invention will be described in detail below with reference to the accompanying drawings. Among them, for the training process of the model, since the model training of deep learning is a well-known technology, the present invention will not elaborate on the specific training process, and will only be described from aspects such as the selection of the model structure, the setting of parameters, and the setting of the loss function.

[0036] When the inventors were studying the problem of Mandarin and Cantonese text classification, they found that the neural networks of the existing methods were difficult to accurately capture the respective features of Cantonese and Mandarin. Therefore, the present invention designs a Cantonese recognition model that uses a shallow network to distinguish between Mandarin and Cantonese in text, and obtains a set of Cantonese recognition models by training the shallow network until convergence to accurately identify whether the text is Cantonese.

[0037] According to an embodiment of the present invention, the present invention provides a training method for a text-oriented Cantonese recognition model, as Figure 1 shown, the method includes steps S1, S2, and S3, and each step will be described in detail below.

[0038] In step S1, Cantonese and Mandarin text corpora are obtained, and the language types of the corpora are manually labeled to obtain a labeled data set.

[0039] According to an embodiment of the present invention, the present invention obtains Cantonese and Mandarin text corpora from public channels, then annotates the language to which they belong, and selects whether to perform data cleaning according to the actual situation of the obtained corpus text. The whole process mainly involves data preparation, data selection, data annotation, etc.

[0040] Among them, data preparation refers to obtaining the original text corpus. According to an embodiment of the present invention, data preparation is to collect information on Mandarin and Cantonese social platforms, Weibo, blogs and other social media as text corpora by using, including but not limited to, web crawler collection technology. The information content includes but is not limited to comments, topics, content posts and other contents. According to an embodiment of the present invention, the Mandarin social platforms are the social media "Sina Weibo" and "Tencent Weibo".

[0041] Data selection refers to selecting data with universality. Among them, the data selected during data selection needs to have two characteristics, that is, it does not have specialized terms and is universal. Data with these properties can more objectively reflect the characteristics of this type of language, making it universal and not special. Among them, social platforms, Weibo, blogs, etc. are places where people usually use to publish their own status and mood, spread their views and other contents. Therefore, the content does not have specialized content terms, such as medical proper nouns, natural science proper nouns, etc. At the same time, the content shared among Weibo users does not all come from the same topic, and has universality.

[0042] Data annotation refers to annotating the language to which the data belongs, and annotating it as Cantonese or non-Cantonese (in this embodiment, non-Cantonese refers to Mandarin). During the annotation process, it should be noted that the lengths of the text corpora are not necessarily the same, especially the data lengths obtained by crawling are different. The recognition application scenarios of Cantonese and Mandarin are mainly in social media platforms, so there are certain requirements for the text length. Being too long or too short will affect the quality of the training model. Therefore, according to an embodiment of the present invention, the present invention eliminates the corpus with a character length less than 4, and splits the corpus with a character length exceeding 100 according to semantics, and labels the processed data set as Cantonese or non-Cantonese for subsequent recognition.

[0043] According to an embodiment of the present invention, the present invention cleans data to reduce the interference of other characters in the corpus that do not belong to Mandarin or Cantonese. It should be noted that data cleaning is not a necessary process. For example, when the text corpus only contains Chinese characters, data cleaning may not be performed. Specifically, data cleaning refers to removing other characters in the corpus that do not belong to Mandarin or Cantonese to prevent the interference of other characters on the training model. Among them, before data cleaning, it is necessary to unify the encoding of the original corpus to ensure data standardization. According to an embodiment of the present invention, combined with user-defined requirements, the encoding can be unified as "GBK" or "UTF-8". For example, when the corpus does not contain other characters except Chinese, it is converted to GBK encoding; when the language of the corpus cannot be guaranteed, it is converted to UTF-8 encoding to save space. Among them, GBK encoding uses two-byte storage, while UTF-8 encoding uses different lengths to store different languages. No matter which encoding form is adopted, unifying the encoding of the corpus requires a process of decoding and encoding conversion. Briefly speaking, it mainly includes the following steps: First, decode the text corpus with different encodings according to the selected encoding method to convert it into unicode encoding (universal code) as the intermediate encoding. Subsequently, encode the string of the intermediate encoding according to the selected encoding method (GBK or UTF-8) to convert it into a unified encoding. Finally, perform data cleaning on the corpus after unified encoding. Among them, the steps of the data cleaning mainly include: TML character conversion, removing emojis, removing url links or website addresses, removing pictures, removing numbers, and other characters that do not belong to Chinese, and then replacing the removed content with spaces to ensure the neatness of short texts. Further, TML character conversion is to remove a large number of html entities such as "<, &" embedded in the original data using regular expressions; the removal of punctuation marks is to remove punctuation marks when data analysis needs to be data-driven at the word level; the removal of emojis is to remove the emojis contained in the short text; the removal of url links is to remove a large number of URL data generated during the crawling process in the short text data and some website links such as the form of http: / / www. and so on; the removal of pictures is to remove the picture names and their picture suffix names obtained during the crawling process, such as.jpg,.png,.gif, etc.; the removal of numbers and other characters that do not belong to Chinese is to remove other language texts such as English and Russian; the cleaning operation is carried out based on each corpus, and regular matching is performed on each short text to achieve text cleaning. Among them, the regular expression describes a pattern of string matching. First, the short text is read line by line and converted into a string, and then traversed and checked in it to check whether the string contains the searched substring. Finally, the string is matched and replaced. The content removed during the data cleaning process is replaced with spaces to ensure the neatness of short texts, and finally the cleaned text is obtained.

[0044] In step S2, the common words of Cantonese and Mandarin are combined with the existing simplified Chinese stop word list to form an improved stop word list.

[0045] This step mainly uses the simplified Chinese stop word list to filter the labeled data set obtained in step S1 and perform word segmentation processing. The word frequencies of the segmented words in Cantonese and Mandarin are respectively counted, and the common words in the Cantonese segmented words and Mandarin segmented words that exceed the preset word frequency threshold are obtained, and combined with the existing simplified Chinese stop word list to form an improved stop word list.

[0046] According to an embodiment of the present invention, the simplified Chinese stop word list described in the present invention is a relatively comprehensive simplified Chinese stop word list formed by at least combining existing stop word lists such as the Baidu stop word list and the Harbin Institute of Technology stop word list. The present invention uses the above-mentioned simplified Chinese stop word list to filter the data sets of Mandarin and Cantonese respectively, divides each piece of corpus by jieba word segmentation, determines the association probability between different characters through the jieba word library, and combines each character with the other character with the highest association probability to form a word group to form a word segmentation result. Among them, the association probability between characters refers to the probability path obtained by dynamic programming calculation in jieba word segmentation. Then, the word frequencies of the segmented words in Mandarin and Cantonese are respectively counted, and the words with word frequencies exceeding the preset word frequency threshold in both parts are compared. If they are the same, they are common words. Common words are noise for the corpus, and the existence of a large amount of noise will affect the extraction of the features of the Cantonese recognition model. Therefore, the present invention adds the common words to the original stop word list to form an improved stop word list, and jieba word segmentation refers to a component provided by python for word segmentation processing of the corpus.

[0047] According to an embodiment of the present invention, the preset word frequency threshold is 5000.

[0048] In step S3, the improved stop word list is used to filter the labeled data set in step S1 and perform word segmentation processing to obtain a training data set, and then the corpus in the training data set is used as the input and the recognition result of whether the corpus is Cantonese is used as the output to train the shallow network until convergence.

[0049] According to an embodiment of the present invention, before training the Cantonese recognition model, by introducing pre-trained word vectors, such as the pre-trained Mandarin word vectors and Cantonese word vectors of wiki, the model fitting is accelerated and the overfitting problem caused by random initialization is avoided. Among them, wiki refers to the word vectors trained through a large amount of corpus provided by Wikipedia.

[0050] According to an embodiment of the present invention, the present invention uses a trained dataset after word segmentation to train a Fasttext shallow network model, and determines the classification result of the corpus according to the prediction label given by the model. Compared with the classification model based on neural network, the Fasttext shallow network model can speed up the training speed and testing speed while maintaining high accuracy. Among them, the effect of the model can be improved by adjusting the epoch, learning rate, and n-gram. Further, according to an embodiment of the present invention, the hyperparameters with the best performance can be obtained by using the grid search method. For example, given the discrete value range of n-gram as [1, 3] and the context window size as [2, 5], if a plane rectangular coordinate system is established with these two parameters as the coordinate axes, these value points are connected into a grid, and the Fasttext model is trained with each point as the parameter, different accuracy results can be obtained, and the parameter with the best effect is taken to obtain the final Cantonese recognition model.

[0051] According to an embodiment of the present invention, the present invention provides a recognition method for a text-oriented Cantonese recognition model, which is used to determine whether the input corpus is Cantonese. The method includes steps T1, obtaining the text to be processed; T2, using the Cantonese recognition model trained by a training method of a text-oriented Cantonese recognition model of the present invention to recognize whether the text to be processed is Cantonese.

[0052] From the description of the above embodiments, it can be seen that the Cantonese recognition model obtained by the present invention through training the shallow network can accurately capture the respective characteristics of Cantonese and Mandarin and identify the language type of the corpus. However, the inventor further studies the existing methods and finds that the existing methods do not utilize the feature that Cantonese has characteristic vocabulary, nor the feature that Cantonese itself has a large number of traditional Chinese characters, and do not consider the differences between Cantonese and Mandarin from multiple perspectives. These features of Cantonese are all features that are beneficial to improving the accuracy of Cantonese recognition. Therefore, the present invention designs a set of rule matching methods and a set of simplified and traditional character recognition methods to find whether the corpus has Cantonese feature words and to distinguish between Mandarin and Cantonese by judging whether the corpus is traditional Chinese, and combines these two methods with the Cantonese recognition model to form a more comprehensive Cantonese recognition system, further improving the accuracy of Cantonese recognition.

[0053] According to an embodiment of the present invention, the present invention provides a set of text-oriented Cantonese recognition systems, such as Figure 2The described Cantonese recognition system includes: a Cantonese recognition model, which is trained using a text-oriented Cantonese recognition model training method of the present invention and is used to identify whether the text to be processed is Cantonese based on the characteristics of the text to be processed to obtain a recognition result; a rule matching model, which is used to retrieve whether the text to be processed hits the Cantonese feature word list based on the Cantonese feature word list to obtain a judgment result on whether the text to be processed is Cantonese; a simplified and traditional Chinese recognition model, which is used to judge whether the text to be processed is traditional Chinese; and a fusion module, which is used to judge whether the text to be processed is Cantonese according to the recognition result of the Cantonese recognition model for the text to be processed, the judgment result of the rule matching model for the text to be processed, and the judgment result of the simplified and traditional Chinese recognition model for the text to be processed.

[0054] According to an embodiment of the present invention, there is provided a training method for a text-oriented Cantonese recognition system, as Figure 3 shown, the method includes steps A1, A2, A3, A4, and A5, and each step will be described in detail below.

[0055] In step A1, Cantonese and Mandarin text corpora are obtained, the language of the corpus is manually annotated to obtain an annotated data set, and the annotated data set is filtered using an improved stop word list and segmented to obtain a training data set.

[0056] In step A2, using the training data set obtained in A1, a shallow network is trained using a text-oriented Cantonese recognition model training method of the present invention until convergence to obtain a Cantonese recognition model.

[0057] In step A3, a Cantonese feature word list is constructed, and based on the Cantonese feature word list, a rule matching model for retrieving whether the corpus hits the Cantonese feature word list is constructed with the corpus obtained in the training data set in step A1 as the input and the judgment result of whether the corpus is Cantonese as the output.

[0058] According to an embodiment of the present invention, the Cantonese feature word list is constructed based on the Cantonese corpus, the different parts of the Cantonese stop word list and the simplified Chinese stop word list, and the Cantonese words in the corpus annotation data set whose word frequency exceeds a preset word frequency threshold.

[0059] The stop word list represents some function words with extremely high frequencies in a language and can be used to represent some characteristics of the language. According to an embodiment of the present invention, the present invention uses the existing Cantonese stop word list provided by Pycantonese, removes the part that is the same as the Simplified Chinese stop word list, and uses the remaining part to characterize the characteristics of Cantonese. At the same time, it combines the part in the corpus annotation dataset in step A1 that exceeds the preset word frequency threshold, as well as some infrequently used Cantonese characteristic words collected from the Hong Kong Cantonese corpus, and combines them to construct a Cantonese characteristic word list. Among them, the Cantonese characteristic word list in the Cantonese characteristic word list can represent the typical characteristics of Cantonese. Each entry in the characteristic word list can be regarded as a rule word in the rule matching method. Generate a dictionary from the rule words, and use the string matching method to output whether the corpus hits the rule word. If it hits, it is regarded as Cantonese, and if it does not hit, it is Mandarin. A rule matching model that can accurately identify Cantonese can be constructed with such rules.

[0060] In step A4, a simplified and traditional Chinese recognition model is constructed with the corpus in the training dataset obtained in step A1 as the input and the judgment result of whether the corpus is traditional Chinese as the output.

[0061] Since Cantonese retains many characteristics of ancient Chinese and contains a large number of traditional Chinese characters, while Mandarin is mainly in Simplified Chinese, therefore, the simplified and traditional Chinese recognition of text corpus can be used as an important basis for determining whether it is Cantonese. According to an embodiment of the present invention, the simplified and traditional Chinese recognition model in the present invention uses the Hanzidentifier model to determine whether each corpus is simplified or traditional Chinese, and can be used to well detect whether the corpus is traditional Chinese. If it contains traditional Chinese, there is a high probability that it belongs to Cantonese, but a fusion module is still needed for correction. Among them, according to an embodiment of the present invention, based on the Hanzidentifier model, first use regular expressions to match and extract the Chinese characters in the corpus, and then query the CC-CEDICT Chinese and traditional Chinese dictionaries for string matching. If the extracted words are all matched in the Simplified Chinese dictionary, it is determined to be simplified. If they are all matched in the traditional Chinese dictionary, it is determined to be traditional. For characters that are not found in either query, they are ignored.

[0062] In step A5, the output of the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model will be used to train the fusion module.

[0063] Considering the multiple differences between Cantonese and Mandarin at the feature level, training a unified classifier model may have poor effects and cannot fully utilize all the feature differences. Therefore, after obtaining their respective models for the three models in the present invention: the shallow network Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model, model fusion is carried out, and the individual models are combined to achieve the purpose of strengthening the simulation effect and performing language identification from multiple aspects.

[0064] According to an embodiment of the present invention, the fusion module of the present invention includes a linear perceptron and is implemented using the PyTorch framework. The true labels of the corpus and the prediction outputs of the three models are made into tensor forms and combined to generate a mini-batch sample dataset, and the output of softmax regression is used as the final recognition result. Among them, the three models will give judgment results for the input text corpus. According to an embodiment of the present invention, the result judged as Cantonese is represented by 1, and the result judged as Mandarin is represented by 0. The results of the three models form a three-dimensional vector, which is used as the input of the linear perceptron model, and a two-dimensional vector is output, representing the probabilities of the final prediction results being 0 and 1.

[0065] According to an embodiment of the present invention, the linear perceptron of the present invention refers to a single-layer neural network composed of a fully connected layer with three inputs and two outputs. It is trained with the three-dimensional vector set formed by the output results of the above-mentioned Cantonese recognition model, rule matching model, and simplified and traditional Chinese character recognition model on the training set to obtain the perceptron model parameters, thereby realizing model fusion.

[0066] According to an embodiment of the present invention, the Softmax function normalizes each value in the vector to the range of 0-1 when the linear perceptron outputs a two-dimensional vector, representing the probabilities of predicting the two results. Among them, softmax regression uses a linear module and defines the forward and backward propagation functions. After randomly initializing the model weights, the model is trained by minimizing the softmax cross-entropy loss function. Among them, considering that too large a learning rate will lead to unstable changes in the optimization direction, and too small a learning rate is likely to cause the model to converge to a local optimal solution. Through multiple adjustment experiments, it is found that when the learning rate is 0.1, the classification accuracy is the highest. Therefore, the present invention selects the mini-batch stochastic gradient descent method with a learning rate of 0.1 as the optimization algorithm in the model fusion process to train the fused model.

[0067] As mentioned in the previous embodiments, for the shallow network Cantonese recognition model of the present invention, the grid search method can be used. Given that the discrete value range of n-gram is [1, 3] and the context window size is [2, 5], if a plane rectangular coordinate system is established with these two parameters as the coordinate axes, then these value points form a grid. Each point is used as a parameter to train the Fasttext model, and different accuracy results are obtained. The parameter with the best effect is selected. For the rule matching model and the simplified and traditional Chinese character recognition model, since there is no parameter adjustment involved, the effect only fluctuates slightly. For model fusion, the cross-entropy loss function is minimized, and the model parameters are trained using backpropagation. The parameter with the minimum loss function is selected as the model parameter. Further, the training parameters with the best effect of the three models are respectively taken as the final model parameters. When encountering new corpora for recognition, the corpora are input into the system. First, the respective prediction results are obtained through 3 independent models, and then the prediction results are input into the fused model together to obtain the final prediction result.

[0068] According to an example of the present invention, the Cantonese recognition model selects parameters with an accuracy rate of 98.8% or more and a recall rate of 98.8% or more. The rule matching model selects parameters with an accuracy rate of 83.99% or more and a recall rate of 92.87% or more. The simplified and traditional Chinese character recognition model selects parameters with an accuracy rate of 92.01% or more and a recall rate of 84.59% or more for model fusion. After fusion, the accuracy rate of the finally fused model can reach 99.78% or more, and the recall rate can reach 96.44% or more.

[0069] According to an embodiment of the present invention, the present invention also provides a recognition method for a text-oriented Cantonese recognition system. The method is used to determine whether the input corpus is Cantonese. The method includes step F1: obtaining the text to be processed; F2: using the Cantonese recognition system trained by a training method of a text-oriented Cantonese recognition system of the present invention to recognize whether the text to be processed is Cantonese.

[0070] Compared with the prior art, the advantages of the present invention are as follows:

[0071] 1. In the existing methods, it is difficult for neural networks to accurately capture the respective characteristics of Cantonese and Mandarin, resulting in low recognition accuracy. However, the present invention designs a Cantonese recognition model that can accurately distinguish Cantonese and Mandarin using a shallow network, has low requirements for the data set, and at the same time improves the common vocabulary table, thereby improving the accuracy.

[0072] 2. In the existing methods, the feature of Cantonese having characteristic vocabulary is not utilized, resulting in low recognition accuracy and reliability. However, the present invention designs a set of rule matching methods to find whether the corpus has Cantonese characteristic words, thereby improving the accuracy and reliability of the determination.

[0073] 3. The existing methods do not utilize the characteristic that Cantonese itself has a large number of traditional Chinese characters. Instead, the present invention designs a set of simplified and traditional character recognition methods. Based on judging whether the corpus is in traditional Chinese, it discriminates between Mandarin and Cantonese, which not only has simple operation but also improves the detection speed.

[0074] In addition, the existing methods do not consider the differences between Cantonese and Mandarin from multiple perspectives. However, the present invention fuses models that identify Cantonese and Mandarin from different perspectives, and simultaneously identifies Cantonese and Mandarin from multiple perspectives to improve the recognition accuracy.

[0075] In summary, through the fusion of multiple methods, the present invention identifies Cantonese and Mandarin from multiple aspects, makes full use of the characteristic differences between Cantonese and Mandarin, does not rely on a single method, improves the recognition accuracy, and makes the prediction results fair.

[0076] It should be noted that although the above steps are described in a specific order, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even the order can be changed as long as the required functions can be achieved.

[0077] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0078] The computer-readable storage medium can be a tangible device that retains and stores instructions for use by an instruction execution device. The computer-readable storage medium may include, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing.

[0079] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A training method for a text-oriented Cantonese recognition system, the Cantonese recognition system include: A Cantonese recognition model is used to identify whether the text to be processed is Cantonese according to the features of the text to be processed and obtain a recognition result; A rule matching model is used to retrieve whether the text to be processed hits the Cantonese characteristic word list based on the Cantonese characteristic word list to obtain a judgment result of whether the text to be processed is Cantonese; The simplified and traditional Chinese recognition model is used to determine whether the text to be processed is traditional Chinese; and a fusion module, for determining whether the text to be processed is Cantonese based on the recognition result of the Cantonese recognition model on the text to be processed, the judgment result of the rule matching model on the text to be processed, and the judgment result of the simplified and traditional Chinese recognition model on the text to be processed; The training method of the text-oriented Cantonese recognition model includes: S1. Obtain Cantonese and Mandarin text corpora, and manually annotate the languages ​​to which the corpora belong to obtain annotated datasets; S2, combining the common words of Cantonese and Mandarin with the existing simplified Chinese stop word list to form an improved stop word list; S3, using the improved stop word list to filter the annotated data set in step S1 and perform word segmentation processing to obtain a training data set, and then using the corpus in the training data set as input and the recognition result of whether the corpus is Cantonese as output to train the shallow network until convergence; The training method of the text-oriented Cantonese recognition system includes: A1. Obtain Cantonese and Mandarin text corpora, manually annotate the languages ​​to which the corpora belong to obtain annotated datasets, use an improved stop word list to filter the annotated datasets and perform word segmentation to obtain training datasets; A2, using the training data set obtained in step A1, adopting the text-oriented Cantonese recognition model training method to train the shallow network until convergence to obtain a Cantonese recognition model; A3, constructing a Cantonese-specific vocabulary, taking the corpus in the training data set obtained in step A1 as input and the judgment result of whether the corpus is Cantonese as output, and constructing a rule matching model for searching whether the corpus hits the Cantonese-specific vocabulary based on the Cantonese-specific vocabulary; A4, using the corpus in the training data set obtained in step A1 as input and the judgment result of whether the corpus is Traditional Chinese as output to build a simplified and traditional Chinese recognition model; A5, training fusion module based on the output of Cantonese recognition model, rule matching model and simplified and traditional Chinese recognition model; Among them, the step A5 includes: using a linear perceptron to fuse the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model; wherein, the linear perceptron is trained with a three-dimensional vector set consisting of the output results of the Cantonese recognition model, the rule matching model, and the simplified and traditional Chinese recognition model on the training data set to obtain the perceptron model parameters to achieve model fusion, and the output of the linear perceptron softmax regression layer is used as the final recognition result.

2. The method according to claim 1, Features , the step S1 comprises: S11. Collect Cantonese and Mandarin text data from Mandarin and Cantonese social platforms through web crawlers; S12, screening the texts in the collected corpus, removing texts whose lengths do not meet the preset minimum text length requirement, and splitting texts whose lengths are greater than the preset maximum text length; S13. Manually annotate the filtered text to label the language of all texts as Cantonese or Mandarin.

3. The method according to claim 2, wherein, the preset minimum text length is 4 and the preset maximum text length is 100.

4. The method according to claim 1, wherein, the step S2 includes: S21. Filter the labeled dataset using a simplified Chinese stopword list; S22. Use jieba segmentation in python to segment each piece of corpus in the filtered labeled dataset, determine the association probability between different characters, and form phrases by combining each character with the other character with the highest association probability, resulting in a word segmentation result; S23. Statistically count the word frequencies of Cantonese and Mandarin word segments respectively, obtain the common words in the Cantonese word segments and Mandarin word segments that exceed the preset word frequency threshold, and combine them with the existing simplified Chinese stopword list to form an improved stopword list.

5. The method according to claim 4, wherein, the preset word frequency threshold is 5000.

6. The method according to claim 1, wherein, the step S3 includes: S31. Filter the labeled dataset using the improved stopword list and perform word segmentation processing to obtain a training dataset; S32. Introduce pre-trained word vectors and use the training dataset to train the Fasttext shallow network until convergence.

7. The method according to claim 1, wherein, the step A3 includes: constructing a Cantonese-specific word list based on the Cantonese corpus, the different parts of the Cantonese stopword list and the simplified Chinese stopword list, and the Cantonese words in the training dataset whose word frequencies exceed the preset word frequency threshold.

8. The method according to claim 1, wherein, the step A4 includes: using the corpus in the training dataset as the input and the judgment result of whether the corpus is Cantonese as the output to train the Hanzidentifier model to obtain a simplified and traditional Chinese recognition model.

9. A Cantonese recognition method for text, wherein, the method includes: F1. Obtain the text to be processed; F2. Use the Cantonese recognition system trained by the method described in any one of claims 1-8 to identify whether the text to be processed is Cantonese.

10. A computer-readable storage medium, wherein, a computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method described in any one of claims 1-9.

11. An electronic device, wherein, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device realizes the steps of the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for realizing text analysis, computer storage medium and terminal

    CN111160015A

  • Putonghua and Cantonese hybrid speech recognition model training method and system

    CN111816160A

  • Language identification method and system based on adaptive center anchor

    CN113282718A

  • Law article and fact relationship calculation method based on multi-layer knowledge door

    CN110737781A

  • Multilingual translation system language search method and device

    CN112528129A