Large model generalization method and device for Cantonese, terminal and storage medium
By converting the predetermined language type data set into Cantonese text and carrying out reinforcement learning training, a Cantonese database is built, and the problem of low Cantonese recognition accuracy is solved, achieving higher recognition accuracy and multilingual mixed text processing capabilities.
Patent Information
- Application Number
- CN202510401624.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The methods used in Cantonese recognition in the prior art are limited by data size and model capacity, resulting in low recognition accuracy.
By obtaining the predetermined language type data set, converting it into original Cantonese text, and performing grammatical structure adjustment; obtaining truth value tags for reinforcement learning training, and building a Cantonese database; when the text to be identified is input into the Cantonese big model, identifying Cantonese vocabulary and searching the database to output recognition results.
This greatly increases the number of Cantonese datasets, captures language features, and improves recognition accuracy, especially when performing well in multilingual mixed text processing.
Smart Images

Figure CN120336468A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, terminal and storage medium for generalizing large models for Cantonese. Background Art
[0002] With the rapid development of Natural Language Processing (NLP) technology, large language models (LLMs) have achieved remarkable results in multiple languages, especially in languages with rich data resources such as English. However, for languages like Cantonese, which have a large number of speakers but scarce data resources, the development of large language models lags behind. There is relatively little research on Cantonese in the field of natural language processing, especially in large language models.
[0003] The particularity of Cantonese lies in its differences from Mandarin in vocabulary, grammar, and pronunciation, as well as its rich oral expressions and cultural connotations. These characteristics pose unique challenges to Cantonese in NLP tasks, such as data scarcity, difficult model training, and mixed use of multiple languages. In addition, the written forms of Cantonese are diverse, including the mixed use of traditional Chinese characters and simplified Chinese characters, further increasing the processing difficulty. Nevertheless, the application demands for Cantonese in fields such as social media, online forums, machine translation, and sentiment analysis are increasing continuously, driving the research and development of Cantonese NLP technology.
[0004] In recent years, with the progress of deep learning technology, small-scale neural networks have achieved certain results in Cantonese NLP tasks, such as false information detection, sentiment analysis, machine translation, and dialogue systems. However, these methods for Cantonese recognition are usually limited by the data scale and model capacity, resulting in low accuracy in recognizing Cantonese.
[0005] Therefore, there are defects in the prior art and it needs to be improved and developed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method, device, terminal and storage medium for generalizing large models for Cantonese in view of the above-mentioned defects of the prior art, aiming to solve the problem that the methods for Cantonese recognition in the prior art are usually limited by the data scale and model capacity, resulting in low accuracy in recognizing Cantonese.
[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0008] A method for generalizing a large model for Cantonese, wherein the method includes:
[0009] Obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text;
[0010] Obtain the true value labels annotated for the original Cantonese text based on the learning task, and input the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese text;
[0011] Construct a Cantonese database based on the new Cantonese text and the original Cantonese text;
[0012] When the text to be recognized is input into the Cantonese large model, recognize the Cantonese vocabulary in the text to be recognized, and retrieve the Cantonese database according to the Cantonese vocabulary to obtain the target Cantonese text that matches the Cantonese vocabulary in the Cantonese database;
[0013] Output the recognition result based on the text to be recognized and the target Cantonese text.
[0014] In an implementation manner of the present application, obtaining a dataset of a predetermined language type and converting each original text in the dataset of the predetermined language type into an original Cantonese text includes:
[0015] Obtain a dataset of a predetermined language type, and translate each original text in the dataset of the predetermined language type into Cantonese text through machine translation;
[0016] Adjust the grammatical structure of the Cantonese text according to the preset grammatical rules to obtain the original Cantonese text.
[0017] In an implementation manner of the present application, obtaining the true value labels annotated for the original Cantonese text based on the learning task, and inputting the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese text includes:
[0018] Obtain the true value labels annotated for the original Cantonese text based on the learning task, and input the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training;
[0019] During the training process, input the original Cantonese text into a reward model predefined with reward metrics, and determine the weights of different parameters in the language model based on the reward model;
[0020] Reallocate the parameters of the language model according to the weights to obtain a Cantonese large model and new Cantonese text.
[0021] In an implementation manner of the present application, the dataset of the predetermined language type includes: a Mandarin dataset and an English dataset; the language model is a Mandarin large model.
[0022] In an implementation manner of the present application, the method for generalizing the large model for Cantonese further includes:
[0023] When the text to be recognized is input into the Cantonese large model, if the text to be recognized is a multilingual mixed text, the language vocabulary corresponding to different languages is encoded into corresponding word vectors;
[0024] Obtain the preset grammar rules corresponding to different languages, and perform grammar parsing on the word vectors of the language vocabulary based on each of the preset grammar rules to obtain the text semantic information corresponding to the multilingual mixed text;
[0025] Output the recognition result based on the text semantic information.
[0026] In an implementation manner of the present application, obtaining the preset grammar rules corresponding to different languages, and performing grammar parsing on the word vectors of the language vocabulary based on each of the preset grammar rules to obtain the text semantic information corresponding to the multilingual mixed text includes:
[0027] Obtain the preset grammar rules corresponding to different languages, perform grammar parsing on the word vectors of the language vocabulary based on each of the preset grammar rules, and combine the context before the language vocabulary to obtain the semantic information of the current language vocabulary;
[0028] Obtain the text semantic information corresponding to the multilingual mixed text based on the semantic information of each language vocabulary.
[0029] In an implementation manner of the present application, the learning tasks include one or more of: sentiment analysis task, text classification task, false information detection task, and machine translation task.
[0030] The present application also provides a large model generalization device for Cantonese, wherein the device includes:
[0031] A conversion module, configured to obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text;
[0032] An input module, configured to obtain the true value labels annotated for the original Cantonese text based on the learning tasks, and input the original Cantonese text and the corresponding true value labels into the language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts;
[0033] A construction module, configured to construct a Cantonese database based on the new Cantonese texts and the original Cantonese texts;
[0034] A retrieval module, configured to, when the text to be recognized is input into the Cantonese large model, recognize the Cantonese vocabulary in the text to be recognized, and retrieve the Cantonese database according to the Cantonese vocabulary to obtain the target Cantonese text that matches the Cantonese vocabulary in the Cantonese database;
[0035] An identification module, configured to output an identification result based on the text to be identified and the target Cantonese text.
[0036] The present application also provides a terminal, which includes: a memory, a processor, and a large model generalization program for Cantonese stored on the memory and executable on the processor. When the large model generalization program for Cantonese is executed by the processor, the steps of the large model generalization method for Cantonese described above are implemented.
[0037] The present application also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the large model generalization method for Cantonese described above.
[0038] A large model generalization method, device, terminal, and storage medium for Cantonese provided by the present invention. The method includes: obtaining a dataset of a predetermined language type, and converting each original text in the dataset of the predetermined language type into an original Cantonese text; obtaining a true value label annotated for the original Cantonese text based on a learning task, and inputting the original Cantonese text and the corresponding true value label into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts; constructing a Cantonese database based on the new Cantonese texts and the original Cantonese texts; when a text to be identified is input into the Cantonese large model, identifying Cantonese vocabulary in the text to be identified, retrieving the Cantonese database according to the Cantonese vocabulary to obtain a target Cantonese text in the Cantonese database that matches the Cantonese vocabulary; and outputting an identification result based on the text to be identified and the target Cantonese text. By converting each original text in the dataset of the predetermined language type into an original Cantonese text, the present application greatly increases the number of Cantonese datasets, thereby better capturing the language features of Cantonese, and using a retrieval-enhanced generation mechanism to refer to a large amount of knowledge in the Cantonese database other than the user's question aspect, improving the accuracy of identification. Description of the Drawings
[0039] Figure 1 It is a flowchart of a preferred embodiment of the large model generalization method for Cantonese in the present invention.
[0040] Figure 2 It is a structural principle block diagram of a prediction model in the present invention.
[0041] Figure 3 It is a structural principle block diagram of an action generation model in the present invention.
[0042] Figure 4 It is a functional principle block diagram of a preferred embodiment of the large model generalization device for Cantonese in the present invention.
[0043] Figure 5It is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. Detailed implementation manners
[0044] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] Small-scale neural networks have achieved certain results in Cantonese NLP tasks, such as false information detection, sentiment analysis, machine translation, and dialogue systems. However, these methods are usually limited by the data scale and model capacity and are difficult to handle complex language phenomena and large-scale application scenarios. In contrast, LLMs have stronger language understanding and generation capabilities, but their current applications in Cantonese are relatively few, and most research focuses on English and other major languages. Therefore, developing LLMs specifically for Cantonese and comprehensively evaluating their performance is of great significance for promoting the development of Cantonese NLP technology.
[0046] The present invention proposes a series of performance evaluation methods and benchmark tests for Cantonese LLMs, aiming to systematically evaluate the capabilities of LLMs in aspects such as Cantonese fact generation, mathematical logic, complex reasoning, and common sense knowledge. Through these benchmark tests, the performance of LLMs in Cantonese tasks can be better understood, and guidance can be provided for future model development and optimization. In addition, the present invention also explores the application of Cantonese LLMs in translation tasks and proposes a method for enhancing Cantonese data based on LLMs to improve the performance and generalization ability of the model.
[0047] The following describes a large model generalization method, device, terminal, and storage medium for Cantonese according to an embodiment of the present application. In view of the problem that the method for Cantonese recognition mentioned in the above background technology is usually limited by the data scale and model capacity and is difficult to handle complex language phenomena and large-scale application scenarios, the present application provides a large model generalization method for Cantonese. In this method, a dataset of a predetermined language type is obtained, and each original text in the dataset of the predetermined language type is converted into an original Cantonese text; a true value label annotated for the original Cantonese text based on a learning task is obtained, and the original Cantonese text and the corresponding true value label are input into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts; a Cantonese database is constructed based on the new Cantonese texts and the original Cantonese texts; when a text to be recognized is input into the Cantonese large model, the Cantonese vocabulary in the text to be recognized is recognized, and the Cantonese database is retrieved according to the Cantonese vocabulary to obtain a target Cantonese text that matches the Cantonese vocabulary in the Cantonese database; and an identification result is output based on the text to be recognized and the target Cantonese text. By converting each original text in the dataset of the predetermined language type into an original Cantonese text, the present application greatly increases the number of Cantonese datasets, thereby better capturing the language features of Cantonese, and uses a retrieval-enhanced generation mechanism to refer to a large amount of knowledge in the Cantonese database other than the user's question aspect, improving the recognition accuracy.
[0048] Please refer to Figure 1 , Figure 1 which is a flowchart of the large model generalization method for Cantonese in the present invention. As Figure 1 shown, the large model generalization method for Cantonese according to an embodiment of the present invention includes:
[0049] Step S100, obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text.
[0050] Specifically, natural language processing methods for Cantonese have been widely applied in fields such as false information detection, sentiment analysis, machine translation, and dialogue systems. However, these methods usually use small-scale neural networks and rely on limited datasets and smaller model architectures. Some technologies introduce the CantoneseBERT and SA-GCN models for detailed analysis and false information detection, but the training corpus contains a large amount of Mandarin content, which may cause language pollution and affect the model effect.
[0051] There are significant differences between Cantonese and Mandarin in vocabulary, grammar, and pronunciation, especially in spoken language. These differences make it difficult for existing technologies based on Mandarin to adapt to Cantonese, and specialized Cantonese vocabulary and corpora are needed to capture the spoken expressions and cultural connotations of Cantonese. However, in existing Cantonese training data, the proportion of Mandarin content is relatively high, which may lead to language pollution and affect the model's ability to process pure Cantonese. In addition, spelling mistakes and new word meanings in Cantonese further increase the difficulty of model training.
[0052] In the embodiment of the present application, step S100 specifically includes:
[0053] Step S110: Obtain a dataset of a predetermined language type, and translate each original text in the dataset of the predetermined language type into a Cantonese text through machine translation;
[0054] Step S120: Adjust the grammatical structure of the Cantonese text according to a preset grammatical rule to obtain an original Cantonese text.
[0055] In an embodiment of the present application, the dataset of the predetermined language type includes: a Mandarin dataset and an English dataset. Specifically, small-scale neural networks are widely used in Cantonese NLP in fields such as misinformation detection, sentiment analysis, machine translation, and dialogue systems. These methods usually rely on limited datasets and smaller model architectures. Although small-scale neural networks perform well in specific tasks, their generalization ability is limited when dealing with complex language phenomena and large-scale tasks. These models are prone to overfitting when facing unseen data or new tasks. The present application translates datasets in other languages into Cantonese texts through machine translation, greatly increasing the number of Cantonese datasets, and thus better capturing the language features of Cantonese.
[0056] As Figure 1 shown, the method for generalizing large models for Cantonese in this embodiment further includes:
[0057] Step S200: Obtain a true value label annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value label into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts.
[0058] In the embodiment of the present application, step S200 specifically includes:
[0059] Step S210: Obtain a true value label annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value label into a language model for reinforcement learning training;
[0060] Step S220: During the training process, input the original Cantonese text into a reward model pre - defined with reward metrics, and determine the weights of different parameters in the language model based on the reward model;
[0061] Step S230: Re - allocate the parameters of the language model according to the weights to obtain a Cantonese large - model and new Cantonese text.
[0062] In this application, for specific tasks such as sentiment analysis and misinformation detection, corresponding labeled data will be used to fine - tune the pre - trained model. During the fine - tuning process, to prevent overfitting, the model adopts regularization techniques. For example, Dropout randomly discards some neuron connections to reduce the dependence between neurons, and L2 regularization is used to constrain the parameters. At the same time, the model also uses a multi - task learning module. For example, the model simultaneously learns sentiment analysis and text classification tasks. By sharing some parameters, the model can learn more extensive language features from multiple tasks and improve its generalization ability.
[0063] In one embodiment, the language model is a Mandarin large - model.
[0064] As Figure 2 shown, Figure 2 It shows the training process of the Cantonese large - model using the reward model. First, the Mandarin dataset and the English dataset are used as the basic data sources. They support the construction of the Cantonese dataset through three ways: machine translation, application of grammar rules, and manual proofreading. Machine translation is responsible for translating Mandarin and English into Cantonese, grammar rules regulate the language structure using Cantonese grammar, and manual proofreading ensures data accuracy. The three work together to form the Cantonese dataset. Then, the constructed Cantonese dataset is input into the reward model. Reward metrics such as translation accuracy and semantic understanding accuracy are pre - defined in the reward model to conduct reinforcement output training on the Mandarin large - model, thereby achieving the purpose of learning training and model fine - tuning for the Mandarin large - model. During this process, the reward model will evaluate the training effect of the Mandarin large - model and give the weights of different parameters. According to this weight, the parameters of the Cantonese large - model are re - allocated, ultimately realizing the optimization and improvement of the Cantonese large - model. For example, in the translation task, the reward model will give rewards according to the matching degree between the translation result and the reference translation, guiding the Cantonese large - model to learn the correct translation method and continuously improve the translation quality. Each element collaborates with each other to form a complete training system for the Cantonese large - model.
[0065] In addition, after the original Cantonese text is input, it enters the data preprocessing module. In this module, a series of operations will be carried out, such as cleaning the data to remove noise and invalid information, denoising the original Cantonese text, using a specific word segmentation algorithm for word segmentation, and performing annotation conversion according to different task requirements. The preprocessed data will be input into a pre-trained model, such as CantoneseBERT. In the pre-training stage, the model will use a large amount of unsupervised Cantonese data for learning to capture the language features of Cantonese and build a preliminary understanding of the Cantonese language patterns and structures. The language model can be a CantoneseBERT model, a CantoneseGPT model, etc. These models can use a Cantonese-specific corpus in the pre-training stage to better capture the language features of Cantonese.
[0066] As Figure 1 shown, the large model generalization method for Cantonese described in this embodiment further includes:
[0067] Step S300: Construct a Cantonese database based on the newly added Cantonese text and the original Cantonese text.
[0068] Specifically, the data in the Cantonese database of this application is Cantonese text represented in vector form, that is, it is a vector database.
[0069] When constructing the Cantonese dataset in this application, not only are Mandarin datasets and English datasets used to convert into Cantonese datasets, but also a Mandarin large model is introduced for reinforcement learning to generate more Cantonese data samples. These data sources contain rich Cantonese spoken expressions and real language usage scenarios, greatly increasing the diversity of the data. At the same time, the collected data is classified and labeled, and classified according to different dimensions such as domain (such as life, technology, entertainment, etc.) and sentiment polarity. In this way, when training the model, corresponding categories of data can be selectively used for training according to different task requirements, improving the pertinence and effectiveness of training.
[0070] As Figure 1 shown, the large model generalization method for Cantonese described in this embodiment further includes:
[0071] Step S400: When the text to be recognized is input into the Cantonese large model, recognize the Cantonese vocabulary in the text to be recognized, and retrieve the Cantonese database according to the Cantonese vocabulary to obtain the target Cantonese text in the Cantonese database that matches the Cantonese vocabulary.
[0072] Specifically, when a user inputs text, the language recognition module first recognizes the language in the text to be recognized, and determines which language or languages it contains, such as Cantonese, Mandarin, English, etc. If Cantonese is recognized, the Cantonese part will be converted into a vector form, and then relevant information, that is, the target Cantonese text, will be retrieved from the vector database.
[0073] As Figure 1 shown, the large model generalization method for Cantonese described in this embodiment further includes:
[0074] Step S500, output an identification result based on the text to be recognized and the target Cantonese text.
[0075] Specifically, the retrieved information and the text to be recognized enter the retrieval augmented generation module together. In the retrieval augmented generation module, processing will be carried out in combination with the capabilities of the pre-trained model, and finally an output result will be generated. Taking the question-answering task as an example, through the retrieval augmented generation mechanism, the model can refer to a large amount of knowledge in the vector database other than the aspects of the user's question and give a more accurate and more valuable answer.
[0076] As Figure 3 shown, first, the vector method for Cantonese lays the foundation for data processing. Through the conversion of Mandarin and English datasets and the reinforcement learning of the large model, the final Cantonese database is constructed. The Cantonese database is a vector database. On this basis, the vector database collaborates with the retrieval augmented generation (RAG) module. The vector database stores data according to a specific vector method, and the retrieval augmented generation module is responsible for retrieving and calling the data therein. At the same time, the Mandarin large model optimizes itself by learning knowledge from the Cantonese database through reinforcement learning. These jointly contribute to the construction of the Cantonese database. In this way, a Cantonese large model integrating the reinforcement learning results of the Mandarin large model and the retrieval augmented generation ability is generated. Finally, the multi-agent collaborative application interacts with the Cantonese large model through workflows and prompts, and calls its functions to complete tasks related to sentiment analysis, misinformation detection, machine translation, etc., realizing a complete closed loop from data processing, model construction to actual application.
[0077] In the embodiment of the present application, the large model generalization method for Cantonese further includes:
[0078] When the text to be recognized is input into the Cantonese large model, if the text to be recognized is a multilingual mixed text, the language vocabulary corresponding to different languages will be encoded into corresponding word vectors;
[0079] Obtain the preset grammar rules corresponding to different languages, and perform grammar parsing on the word vectors of the language vocabulary based on each preset grammar rule to obtain the text semantic information corresponding to the multilingual mixed text;
[0080] A recognition result is output based on the text semantic information.
[0081] Specifically, Cantonese speakers often mix Mandarin and English in their communications, and this multilingual code switching phenomenon increases the complexity of NLP systems. Existing Cantonese NLP systems often perform poorly when processing multilingual mixed texts.
[0082] This application uses a special word vector representation method, which can encode vocabulary information of different languages at the same time, so that the model can accurately distinguish and understand the vocabulary of different languages when processing multilingual mixed texts. For example, when faced with a multilingual mixed text such as "I like eating pizza", the model locates that "like" is a Cantonese word and "pizza" is an English word through the language recognition module. During the processing, the model adopts different grammatical parsing strategies for vocabulary in different languages. For Cantonese vocabulary, it is analyzed according to the grammatical rules of Cantonese; for English vocabulary, it is processed in combination with the grammatical and semantic characteristics of English, so as to accurately understand the semantics of the entire text. Compared with traditional models, it greatly improves the understanding and processing capabilities of multilingual mixed texts.
[0083] In one embodiment of the present application, the preset grammar rules corresponding to different languages are obtained, and the word vectors of the vocabulary of each language are parsed based on the preset grammar rules to obtain the text semantic information corresponding to the multilingual mixed text, including:
[0084] Obtaining preset grammar rules corresponding to different languages, performing grammatical analysis on the word vectors of the vocabulary of each language based on the preset grammar rules, and obtaining semantic information of the current language vocabulary in combination with the previous context of the language vocabulary;
[0085] The text semantic information corresponding to the multi-language mixed text is obtained based on the semantic information of the vocabulary of each language.
[0086] Specifically, this application introduces a context-aware mechanism, which mainly uses an attention mechanism to capture the contextual information of the text. When processing multilingual mixed texts, the model will dynamically adjust the language processing strategy according to the language environment of the previous text. For example, when the previous text is in a Cantonese context and English vocabulary appears later, the model can accurately understand the meaning of the English vocabulary in combination with the Cantonese context of the previous text. It will pay attention to information such as the theme and emotional tendency of the previous text, so as to more accurately grasp the semantics of English vocabulary in the current context, improve the model's adaptability to different language environments, and make the model perform better in multilingual mixed text processing tasks.
[0087] In an embodiment of the present application, the learning task includes: one or more of a sentiment analysis task, a text classification task, a false information detection task and a machine translation task.
[0088] Specifically, in the sentiment analysis task, the text to be recognized is Cantonese text. These texts first enter the data preprocessing module, where word segmentation is performed to split the text into individual words or phrases, and at the same time, information related to the sentiment polarity of the text (such as positive, negative, neutral) is marked. The preprocessed text data is input into a fine-tuned pre-trained model. After calculation and analysis by the model, the sentiment category of the text is finally output, clearly determining whether the text expresses positive, negative or neutral sentiment. At the same time, a sentiment intensity value will also be output to more precisely measure the intensity of the sentiment.
[0089] In the misinformation detection task, the text to be recognized is mainly Cantonese tweets or other forms of text data. These data first enter the data preprocessing module. In the preprocessing module, the data is cleaned to remove irrelevant format information and noise, and key information in the text, such as keywords and key sentences, is extracted. The processed data is input into a misinformation detection model constructed based on model structures such as CantoneseBERT and SA-GCN. After complex calculation and judgment by the model, the judgment result of whether the text is misinformation or non-misinformation is finally output. At the same time, a credibility score will also be given to indicate the reliability of the judgment result.
[0090] In the machine translation task, the text to be recognized is the source language text, which may be text in other languages such as Mandarin, English, etc. First, the language recognition module determines the language type of the input text to identify which source language it is. Then, the machine translation module converts the source language text into Cantonese based on the pre-trained and learned translation knowledge. The translated Cantonese text will pass through a grammar rule checking module to check whether the grammar of the translation conforms to the grammar norms of Cantonese. If there are problems, corrections will be made. Finally, after manual proofreading, the accuracy and fluency of the translation are ensured, and an accurate and natural Cantonese translation is finally output.
[0091] The embodiments of the present application achieve the following effects:
[0092] First, the pre-trained language models (LLMs) specifically for Cantonese provided by the present application, such as CantoneseBERT, CantoneseGPT, etc., can use a Cantonese-specific corpus in the pre-training stage to better capture the language features of Cantonese.
[0093] Second, support for Cantonese is added to the multilingual model. Through multilingual pre-training and fine-tuning, the performance of the model in Cantonese tasks is improved. For example, cross-lingual embedding technology is used to enable the model to better process multilingual mixed texts.
[0094] Third, for specific NLP tasks, fine-tune the pre-trained model to meet the specific needs of Cantonese. For example, in the sentiment analysis task, the model can be fine-tuned using Cantonese sentiment annotation data to improve its performance in Cantonese sentiment analysis.
[0095] Fourth, provide models that can handle multilingual mixed text, and improve the model's performance in multilingual mixed text by identifying and processing the vocabulary and grammar structures of different languages.
[0096] Fifth, add a language recognition module to the model, which can automatically identify different languages in the text and make corresponding switches during processing. For example, use a language model to identify the Cantonese, Mandarin, and English parts in the text and process them using the corresponding language models respectively.
[0097] Sixth, introduce a context-aware mechanism to enable the model to dynamically adjust language processing strategies according to the context and better handle multilingual mixed text.
[0098] Seventh, use large-scale pre-trained models (such as LLMs) to improve the generalization ability of the model through a large amount of unsupervised pre-training data. These models can learn a wider range of language patterns and structures during the pre-training stage, thus performing better during fine-tuning.
[0099] Eighth, during the model training process, use regularization techniques (such as Dropout, L2 regularization, etc.) to prevent overfitting and improve the generalization ability of the model.
[0100] Ninth, through multi-task learning, enable the model to learn multiple related tasks simultaneously, improving the generalization ability and adaptability of the model. For example, perform sentiment analysis and text classification tasks simultaneously, enabling the model to learn a wider range of language features.
[0101] In one embodiment, as Figure 4 shown, based on the above-mentioned large model generalization method for Cantonese, the present invention also correspondingly provides a large model generalization device for Cantonese, and the device includes:
[0102] A conversion module 100, configured to obtain a dataset of a predetermined language type and convert each original text in the dataset of the predetermined language type into an original Cantonese text;
[0103] An input module 200, configured to obtain a true value label annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value label into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese text;
[0104] A construction module 300, configured to construct a Cantonese database based on the new Cantonese text and the original Cantonese text;
[0105] A retrieval module 400, configured to identify Cantonese words in the text to be recognized when the text to be recognized is input into the Cantonese large model, and retrieve the Cantonese database according to the Cantonese words, so as to obtain a target Cantonese text in the Cantonese database that matches the Cantonese words;
[0106] An identification module 500, configured to output an identification result based on the text to be recognized and the target Cantonese text.
[0107] Figure 5 The figure is a schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal may include:
[0108] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0109] When the processor 502 executes the program, it implements the large model generalization method for Cantonese provided in the above embodiment.
[0110] Further, the terminal further includes:
[0111] A communication interface 503, configured to communicate between the memory 501 and the processor 502.
[0112] The memory 501 is used to store a computer program executable on the processor 502.
[0113] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0114] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0115] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.
[0116] The processor 502 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.
[0117] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned generalization method for Cantonese large models.
[0118] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0119] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0120] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.
[0121] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can read and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0122] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0123] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and when the program is executed, it includes one or a combination of the steps of the method embodiments.
[0124] In addition, each functional unit in various embodiments of the present application can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0125] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0126] In summary, a large model generalization method, device, terminal, and storage medium for Cantonese disclosed by the present invention, the method includes: obtaining a dataset of a predetermined language type, and converting each original text in the dataset of the predetermined language type into an original Cantonese text; obtaining a true value label annotated for the original Cantonese text based on a learning task, and inputting the original Cantonese text and the corresponding true value label into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts; constructing a Cantonese database based on the new Cantonese texts and the original Cantonese texts; when a text to be recognized is input into the Cantonese large model, recognizing Cantonese words in the text to be recognized, and retrieving the Cantonese database according to the Cantonese words to obtain a target Cantonese text in the Cantonese database that matches the Cantonese words; and outputting a recognition result based on the text to be recognized and the target Cantonese text. By converting each original text in the dataset of the predetermined language type into an original Cantonese text, the present application greatly increases the number of Cantonese datasets, thereby better capturing the language features of Cantonese, and using a retrieval-enhanced generation mechanism to refer to a large amount of knowledge in the Cantonese database other than the user's question aspect, improving the recognition accuracy.
[0127] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A large model generalization method for Cantonese, characterized in that, The method includes: Obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text; Obtain true value labels annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training to obtain a large Cantonese model and new Cantonese texts; Construct a Cantonese database based on the new Cantonese texts and the original Cantonese texts; When a text to be recognized is input into the large Cantonese model, recognize the Cantonese vocabulary in the text to be recognized, and retrieve the target Cantonese text in the Cantonese database that matches the Cantonese vocabulary according to the Cantonese vocabulary; Output a recognition result based on the text to be recognized and the target Cantonese text.
2. The method for generalizing large models for Cantonese according to claim 1, wherein Obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text, including: Obtain a dataset of a predetermined language type, and translate each original text in the dataset of the predetermined language type into a Cantonese text through machine translation; Adjust the grammatical structure of the Cantonese text according to a preset grammatical rule to obtain an original Cantonese text.
3. The generalization method of the large model for Cantonese according to claim 1, wherein Obtain true value labels annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training to obtain a large Cantonese model and new Cantonese texts, including: Obtain true value labels annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding true value labels into a language model for reinforcement learning training; During the training process, input the original Cantonese text into a reward model predefined with a reward metric, and determine the weights of different parameters in the language model based on the reward model; Reallocate the parameters of the language model according to the weights to obtain a large Cantonese model and new Cantonese texts.
4. The method for generalizing large models for Cantonese according to claim 1, wherein The dataset of the predetermined language type includes: a Mandarin dataset and an English dataset; the language model is a large Mandarin model.
5. The method for generalizing large models for Cantonese according to claim 1, characterized in that The method for generalizing the large model for Cantonese further includes: When a text to be recognized is input into the large Cantonese model, if the text to be recognized is a multilingual mixed text, encode the language vocabulary corresponding to different languages into corresponding word vectors; Obtain preset grammatical rules corresponding to different languages, and perform grammatical parsing on the word vectors of each language vocabulary based on each preset grammatical rule to obtain text semantic information corresponding to the multilingual mixed text; Output a recognition result based on the text semantic information.
6. The method for generalizing large models for Cantonese according to claim 5, characterized in that, Obtain preset grammatical rules corresponding to different languages, and perform grammatical parsing on the word vectors of each language vocabulary based on each preset grammatical rule to obtain text semantic information corresponding to the multilingual mixed text, including: Obtain preset grammatical rules corresponding to different languages, perform grammatical parsing on the word vectors of each language vocabulary based on each preset grammatical rule, and combine the previous context of the language vocabulary to obtain the semantic information of the current language vocabulary; Obtain text semantic information corresponding to the multilingual mixed text based on the semantic information of each language vocabulary.
7. The method for generalizing large models for Cantonese according to claim 1, wherein The learning tasks include one or more of: sentiment analysis tasks, text classification tasks, misinformation detection tasks, and machine translation tasks.
8. A large model generalization device for Cantonese, characterized in that, The device includes: A conversion module, configured to obtain a dataset of a predetermined language type, and convert each original text in the dataset of the predetermined language type into an original Cantonese text; An input module, configured to obtain a ground truth label annotated for the original Cantonese text based on a learning task, and input the original Cantonese text and the corresponding ground truth label into a language model for reinforcement learning training to obtain a Cantonese large model and new Cantonese texts; A construction module, configured to construct a Cantonese database based on the new Cantonese texts and the original Cantonese texts; A retrieval module, configured to, when a text to be recognized is input into the Cantonese large model, recognize Cantonese words in the text to be recognized, and retrieve the Cantonese database according to the Cantonese words to obtain a target Cantonese text in the Cantonese database that matches the Cantonese words; An identification module, configured to output an identification result based on the text to be recognized and the target Cantonese text.
9. A terminal, characterized in that, Including: A memory, a processor, and a large model generalization program for Cantonese stored on the memory and executable on the processor. When the large model generalization program for Cantonese is executed by the processor, the steps of the large model generalization method for Cantonese according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the large model generalization method for Cantonese according to any one of claims 1 to 7.
Citation Information
Cited By
Voice data set generation method and device
CN120877702A