Multi-language general part-of-speech recognition method and system based on large language model
Through the multilingual universal part-of-speech recognition method based on large language model, the problem that traditional technology is difficult to deal with multilingual and cross-domain data in part-of-speech recognition is solved, efficient and automated part-of-speech recognition and syntactic analysis are realized, and structured data is generated to support diversified business needs.
Patent Information
- Application Number
- CN202411932124.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art relies on rule-based methods or deep learning models that require a large amount of annotated data in the field of part-of-speech recognition, which is difficult to fully capture the contextual relationships of text, has a long calculation time and poor parallel computing capabilities, and faces significant challenges in cross-domain and multilingual data processing.
The multilingual universal part-of-speech recognition method based on the large language model is adopted. By pre-processing the received original data, it is suitable for the input format of the large language model. Through low-rank adaptation, fine-tuning and prompt word design, the large language model is fine-tuned to form a lightweight model, which is used to recognize the universal part-of-speech and code-analyzing the recognition results.
It greatly simplifies the part-of-speech recognition process, reduces development and deployment costs, improves cross-domain and multilingual adaptability, realizes automated data processing, and the generated JSON format data supports diversified business needs, improving the flexibility, accuracy and automation level of the system.
Smart Images

Figure CN120012771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and part-of-speech recognition, and in particular to a multi-language universal part-of-speech recognition method and system based on a large language model. Background Art
[0002] In today's digital world, the size and complexity of text data continues to increase. The information people come into contact with on a daily basis includes social media tweets, customer reviews, news articles, emails, and various corporate reports. As these text data accumulate, how to efficiently and accurately identify part-of-speech information and syntactic structures from them has become a key requirement in many business scenarios.
[0003] As a data service, part-of-speech recognition can parse the lexical information and part-of-speech tagging results in the text into semi-structured data in JSON format. This data can not only provide business logic models for the back-end system, but also generate user-visible data view models for the front-end system to meet diverse business needs. Due to the different database table designs and implementations of different business systems, the data structures that need to be received are also different. For example, in the fields of finance, medical care, and e-commerce, automated part-of-speech recognition can generate structured customer files, product information tables, diagnostic reports, etc., significantly reducing manual processing costs, and can fine-tune the model through a small amount of labeled data to achieve domain adaptation.
[0004] Compared with traditional rule-based part-of-speech recognition methods and models that rely on a large amount of annotated data, large language models demonstrate strong generalization and contextual understanding capabilities, and can reduce reliance on complex rules and large amounts of annotated data. Through low-rank adaptive fine-tuning technology, the characteristics of large language models enable them to quickly adapt to the needs of different fields, and through prompt design, guide the model to generate part-of-speech tagging and syntactic analysis results that meet business needs. This approach improves the flexibility, accuracy, and automation level of the system in different business scenarios.
[0005] Existing technologies in the field of part-of-speech recognition mainly rely on rule-based methods or deep learning models that require a large amount of annotated data, such as convolutional neural networks and recurrent neural networks. However, due to the limitations of local feature extraction, convolutional neural networks cannot fully capture the contextual relationship of texts; recurrent neural networks, as sequence models, not only have a long calculation time, but also have poor parallel computing capabilities, making it difficult to play an effective role in large-scale data processing and real-time applications. In addition, these traditional models also face significant challenges in dealing with cross-domain and multilingual data. Summary of the invention
[0006] In order to solve the above problems, the purpose of the present invention is to provide a multilingual universal part-of-speech recognition technology based on a large language model, aiming to extract lexical, part-of-speech and syntactic structure information from text and provide data support for various business systems.
[0007] In order to achieve the above technical objectives, the present application provides a multilingual universal part-of-speech recognition method based on a large language model, comprising the following steps:
[0008] Preprocess the received raw data so that each input question corresponds to a reference answer, and fine-tune the large language model by designing prompt words for specific tasks;
[0009] The fine-tuned large language model is distilled to form a lightweight model, which is used to recognize common parts of speech in multiple languages and parse the recognition results in code.
[0010] Preferably, in the process of preprocessing the received raw data, the received raw data is preprocessed by dividing paragraphs, removing noise and special symbols, performing word segmentation, and obtaining text representation, and part of the preprocessed data is labeled so that each input question corresponds to a reference answer.
[0011] Preferably, during the pretreatment process, the specific steps of pretreatment include:
[0012] Paragraph division: break the input text into logical paragraphs;
[0013] Remove noise and special symbols: Remove noise data, delete meaningless special symbols, and keep only the punctuation marks required for the sentence;
[0014] Perform tokenization: split the text into tokens;
[0015] Obtain text representation: Convert the segmented text into a numerical representation that the model can process.
[0016] Preferably, in the process of fine-tuning the large language model, the large language model is fine-tuned by a low-rank adaptation method.
[0017] Preferably, in the process of designing prompt words, prompt words are designed for the lexical analysis task, the part-of-speech tagging task and the syntactic analysis task respectively, wherein:
[0018] For lexical analysis tasks: by designing prompt words, the model can identify the categories, word forms and language roles of words from the text and generate lexical information;
[0019] For part-of-speech tagging tasks: By designing prompt words, the model generates part-of-speech tagging information and adds labels to each word for accurate tagging in a multilingual environment;
[0020] For syntactic analysis tasks: By designing prompt words, the model generates a syntactic tree structure to analyze the grammatical relationship and dependency structure between words.
[0021] Preferably, in the process of coding and parsing the recognition results, the recognition results are converted into JSON format data.
[0022] Preferably, in the process of converting the recognition results into JSON format data, the lexical, part-of-speech and syntactic analysis results generated by the model are parsed, annotation information is added to each word, and this information is mapped to a standard JSON field, wherein the annotation information includes part-of-speech tags, syntactic dependencies and related grammatical attributes.
[0023] The present invention discloses a multilingual universal part-of-speech recognition system based on a large language model, which is used to implement the multilingual universal part-of-speech recognition method based on a large language model mentioned above, comprising:
[0024] The model fine-tuning module is used to pre-process the received raw data so that each input question corresponds to a reference answer and fine-tune the large language model by designing prompt words for specific tasks;
[0025] The multilingual general part-of-speech recognition module is used to distill the fine-tuned large language model to form a lightweight model for multilingual general part-of-speech recognition and to codedly parse the recognition results.
[0026] The present invention discloses the following technical effects:
[0027] Greatly simplify the process and reduce development and deployment costs: The present invention uses a large language model as the core framework to encapsulate the main process of part-of-speech recognition, significantly reducing the complexity of development and deployment. Compared with traditional rule-based part-of-speech recognition methods, the present invention does not rely on a large number of manually written language rules or annotation data in specific fields. Complex tasks can be completed by designing prompt words and a small amount of fine-tuning. This approach reduces the reliance on natural language processing experts and linguistic knowledge, accelerates system development iterations, and greatly improves the efficiency of model deployment in production environments.
[0028] Cross-domain and multi-language adaptability: The present invention selects a large language model with strong generalization ability as the basis, combined with low-rank adaptive fine-tuning, so that the model can quickly adapt to the needs of different fields. At the same time, the prompt word design gives the model the ability to dynamically respond to different languages and tasks, such as lexical analysis, part-of-speech tagging, and syntactic analysis. The system can efficiently handle multilingual environments (such as mixed Chinese and English texts) or business needs in different fields (such as law, medical, etc.) in a single model, ensuring that it can still maintain high accuracy, robustness and stability in complex business scenarios.
[0029] Data service support to achieve automated data processing: The JSON format structured data generated by the present invention provides a standardized interface for subsequent business systems by parsing information such as lexical, part of speech, and syntax. These data can be directly handed over to other natural language processing systems (such as sentiment analysis and text classification) for use, and can also be integrated into the back-end system for automated processing. At the same time, the system also supports front-end business applications to generate visual views of grammatical data, providing users with intuitive displays and optimizing user experience. This data service-oriented design not only improves the interactivity of the system, but also enables the system to respond quickly to business needs.
[0030] This invention solves the limitation of traditional models in the field of part-of-speech recognition, and constructs a method that can not only provide natural language part-of-speech analysis for various business scenarios such as finance, medical care, and public opinion monitoring, but also provide lexical analysis for non-natural languages such as programming languages. This method can be used as a data support service, significantly improving the automation level and adaptability of part-of-speech recognition in business systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 It is a flow chart of the multilingual general part-of-speech analysis method based on a large language model according to the present invention;
[0033] Figure 2 Schematic diagram of the text preprocessing process of the present invention
[0034] Figure 3 Schematic diagram of the process of model adaptation enhancement and distillation optimization described in the present invention
[0035] Figure 4 A schematic diagram of the flow of the lexical analysis method based on a large language model according to the present invention;
[0036] Figure 5 A schematic diagram of the process of the part-of-speech tagging method based on a large language model according to the present invention;
[0037] Figure 6 It is a flowchart of the syntactic analysis method based on a large language model described in the present invention. DETAILED DESCRIPTION
[0038] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0039] like Figure 1-6 As shown, the present invention provides a multilingual universal part-of-speech recognition method based on a large language model, comprising the following steps:
[0040] Text preprocessing: The system performs basic cleaning and preprocessing operations on the received raw data, including segmenting paragraphs, removing noise and special symbols, performing word segmentation, and obtaining text representation;
[0041] Model adaptation enhancement and distillation optimization: Design prompt words for specific tasks, use low-rank adaptation methods to fine-tune the large language model, and use the fine-tuned large model for distillation to generate a lightweight model.
[0042] Generate semi-structured data: Normalize the model output by designing special prompt words, and use code to parse the model output to generate JSON format data for subsequent business system integration.
[0043] As a preferred technical solution, the specific steps of the text preprocessing include:
[0044] Divide into paragraphs: Divide the input text into paragraphs based on separators such as line breaks and punctuation marks, and decompose it into logical paragraphs to ensure that the model can process them paragraph by paragraph and reduce context confusion.
[0045] Remove noise and special symbols: Remove noise data such as HTML tags, emoticons, control characters, etc. Delete redundant spaces, carriage returns, non-language symbols (such as #$%^&), and only keep the punctuation marks required for the sentence. Clean up irrelevant information, reduce interference in model processing, and improve model performance.
[0046] Perform tokenization: Split the text into tokens to prepare for model input
[0047] Obtain text representation: Convert the segmented text into a numerical representation that the model can process.
[0048] As a preferred technical solution, the specific steps of the model adaptation enhancement and distillation optimization include:
[0049] Model adaptation enhancement: Through the combination of low-rank adaptation fine-tuning and prompt word design, the performance of the model in specific tasks can be effectively improved.
[0050] Model distillation optimization: The model is adapted and enhanced to a large model for distillation to generate a lightweight model. This method is performed through a teacher-student architecture, where the fine-tuned large language model (teacher model) generates soft labels and intermediate representations during training to guide the learning of the lightweight model (student model).
[0051] As a preferred technical solution, the specific steps of the model adaptation enhancement include:
[0052] Low-rank adaptive fine-tuning: Low-rank adaptive fine-tuning technology only adjusts the low-rank part of some weight matrices, thereby improving training efficiency and reducing computational costs. During the fine-tuning process, the model can be adapted according to domain requirements to improve its part-of-speech recognition ability in specific tasks. The fine-tuned model has a higher level of semantic understanding, can handle multilingual environments, and maintain good performance in different fields.
[0053] Design prompt words: Design different prompt words for specific tasks to guide the model to output standardized results. Prompt words not only help the model understand the contextual semantics of the input, but also ensure that the output conforms to the expected format. Especially in a multilingual environment, the design of prompt words can effectively improve the generalization ability of the model and the controllability of the output.
[0054] As a preferred technical solution, the specific steps of designing prompt words include:
[0055] Special prompt words are designed for different tasks. The following introduces the prompt word design ideas for the three tasks of lexical analysis, part-of-speech tagging, and syntactic analysis.
[0056] Lexical analysis: The prompt words guide the model to identify the categories, word forms and language roles of words from the text and generate detailed lexical information.
[0057] Part-of-speech tagging: Prompt words guide the model to generate part-of-speech tagging information, such as nouns, verbs, adjectives, etc., and add appropriate labels to each word to support accurate tagging in multilingual environments.
[0058] Syntactic analysis: The prompt words guide the model to generate a syntactic tree structure, analyze the grammatical relationship and dependency structure between words, and provide support for upper-level semantic understanding.
[0059] As a preferred technical solution, the specific steps of generating semi-structured data are:
[0060] Model output based on prompt word standardization: The system is designed with dedicated prompt words for tasks such as lexical analysis, part-of-speech tagging, and syntactic analysis. It can guide large language models to generate standardized results and support lexical processing for natural languages and programming languages.
[0061] Parsing the model to generate JSON format data: The output content of the model based on prompt word standardization is used as input and structured.
[0062] Business integration and application of JSON data: The parsing model generates data in JSON format, which can be used for subsequent business system integration according to business needs. Different systems can support various downstream applications through this data, such as text classification, grammar checking and intelligent search.
[0063] As a preferred technical solution, the specific steps of generating JSON format data from the parsing model include:
[0064] First, the lexical, part-of-speech, and syntactic analysis results generated by the model are parsed through a dedicated parsing script to ensure the uniformity and integrity of the data format. During the parsing process, the system will add necessary annotation information to each word, such as part-of-speech tags, syntactic dependencies, and related grammatical attributes, and map this information to standard JSON fields.
[0065] In general, the present invention realizes automatic recognition of parts of speech from text, lexical analysis and syntactic structure analysis through a large language model, and converts the recognition results into semi-structured data in JSON format, providing data support for different business systems. Due to the different database table designs and implementations of different business systems, the structured data generated by the system can be flexibly adjusted according to demand, ensuring that the backend can receive data models that conform to its business logic, and providing a user-visible grammatical data view model for the front-end system. In addition, the present invention not only supports part-of-speech recognition of natural language, but can also be extended to lexical analysis of programming languages. During the code parsing process, the system can accurately identify different parts of speech such as variable names, function names, keywords, and generate structured information adapted to the intelligent development environment, providing support for code parsing, grammar checking and automatic completion. This capability improves the automation level of software development and the efficiency of code quality management. For example, in the Text2SQL (text conversion database query statement service) scenario, the present invention uses part-of-speech recognition and syntactic analysis capabilities to parse natural language queries into structured query statements that conform to SQL grammar. Through the semantic understanding and context parsing capabilities of the large language model, the present invention can accurately generate SQL queries and provide efficient support for database queries and data analysis. The present invention has wide applications in the fields of finance, e-commerce, etc., and can help users efficiently access complex databases through natural language.
[0066] Embodiment: This embodiment provides a multilingual universal part-of-speech recognition method based on a large language model, and its main process includes the following steps:
[0067] S1: Text preprocessing: Figure 2 As shown, it includes dividing paragraphs, removing noise and special symbols, performing word unit segmentation, and obtaining text representation; some preprocessed data needs to be manually labeled to ensure that each input question corresponds to a reference answer.
[0068] Step S1: text preprocessing, specifically including the following sub-steps:
[0069] S11: Divide into paragraphs
[0070] Break the input text into logical paragraphs. Divide the paragraphs based on line breaks, whitespace, punctuation (such as periods, question marks, exclamation marks), and other delimiters. For multilingual text, identify the paragraph structure of each language (for example, Chinese paragraphs do not rely on spaces, and English paragraphs are separated by sentences).
[0071] S12: Remove noise and special symbols
[0072] According to the logical paragraphs obtained by dividing the paragraphs in step S11, remove noise data such as HTML tags, URLs, emoticons, control characters, etc.; delete redundant spaces, carriage returns, and non-language symbols (such as #$%^&), and only retain the punctuation marks required for the sentence; for multilingual data, process them separately according to language characteristics, such as removing inappropriate abbreviation symbols in English and retaining Chinese punctuation marks.
[0073] S13: Execute word segmentation
[0074] According to the clean data obtained by removing noise and special symbols in step S12, the BPE (Byte-Pair Encoding) or SentencePiece method is used to generate sub-word level word units.
[0075] S14: Obtaining text representation
[0076] According to the word units obtained from the text segmented by word unit in step S12, a pre-trained word vector model, such as GloVe or Word2Vec, is used to map each word into a vector representation of a fixed length, and the segmented text is converted into a numerical representation that can be processed by the model.
[0077] S2: Model adaptation enhancement and distillation optimization: Figure 3As shown in the figure, the current large language models are basically question-answering models, so it is assumed that the large language model is QA (Q is Question, A is Answer). The base model selected in this example is LLAMA. A small amount of labeled data is used to perform low-rank adaptation fine-tuning on LLAMA, and the self-attention matrix of the Transformer is fine-tuned. At the same time, some specific prompt words P are set, such as the prompt word "Please mark the part of speech for each word in the following sentence." to guide the model to generate part-of-speech tagging results that meet the standards, and thus the enhanced model PQ-A for the adaptation task is obtained. However, considering that large language models require a lot of computing resources, we adopted model distillation optimization, that is, by transferring the knowledge of the fine-tuned large model to the small model, a lightweight and efficient process is achieved.
[0078] Step S2: Model adaptation enhancement and distillation optimization, specifically including the following sub-steps:
[0079] S21: Model Adaptation Enhancement
[0080] First, a large language model is selected as the base model. In this example, LLAMA is selected. The model adaptation enhancement module consists of two parts: low-rank adaptation fine-tuning and designing specific prompt words. The detailed description of these two parts is as follows:
[0081] a. Low-rank adaptive fine-tuning: Large LLAMA models usually contain billions or even tens of billions of parameters. Traditional fine-tuning requires updating most or all of the model's parameters, which is too computationally intensive. Low-rank adaptive fine-tuning can only adjust the low-rank part of the weight matrix, greatly reducing the amount of parameter updates. This example uses low-rank adaptive fine-tuning technology to fine-tune the self-attention matrix parameters in the Transformer module.
[0082] b. Prompt word design: Design different prompt words for different task requirements to guide the model to output standardized results. This example provides prompt words for part-of-speech analysis tasks: "Please mark the part of speech for each word in the following sentence." The multi-language universal part-of-speech recognition method and system based on a large language model provided by the present invention can adjust the prompt words of different tasks to adapt to different part-of-speech recognition tasks, such as lexical analysis, part-of-speech tagging, and syntactic analysis.
[0083] Lexical analysis: The prompt words guide the model to identify the categories, word forms and language roles of words from the text and generate detailed lexical information.
[0084] Part-of-speech tagging: Prompt words guide the model to generate part-of-speech tagging information, such as nouns, verbs, adjectives, etc., and add appropriate labels to each word to support accurate tagging in multilingual environments.
[0085] Syntactic analysis: Prompt words guide the generation of syntactic tree structures, analyze the grammatical relationships and dependency structures between words, and provide support for upper-level semantic understanding.
[0086] S22: Model distillation optimization
[0087] The large model fine-tuned in step S21 is used for distillation to generate a lightweight model.
[0088] S3: Generate semi-structured data: Figure 4 As shown, the output of the model is normalized by special prompt words, and the model output is parsed with code to generate JSON format data for subsequent business system integration.
[0089] Step S3 generates semi-structured data, which specifically includes the following sub-steps:
[0090] S31: Model specification output based on prompt words
[0091] The system has designed special prompt words for tasks such as lexical analysis, part-of-speech tagging, and syntactic analysis. The prompt words guide the large language model to generate standardized part-of-speech tagging and syntactic analysis results, and support lexical processing for natural languages and programming languages.
[0092] Here are some examples of normalization:
[0093] a. Lexical analysis
[0094] Input: int a = 5;
[0095] Output: [int / keyword, a / variable, = / operator, 5 / constant]
[0096] b. Part-of-speech tagging
[0097] Input: I like to eat apples.
[0098] Output: [I / PRON, like / VERB, eat / VERB, apple / NOUN,. / PUNCT]
[0099] c. Syntactic analysis
[0100] Input: The boy who played soccer is tired.
[0101] Output (dependency tree):
[0102] makefileCopy codeROOT:tired
[0103] boy(nsubj)
[0104] played(relcl)
[0105] soccer(obj)
[0106] S32: Parse the model to generate JSON format data
[0107] After obtaining the normalized model output in step S32, the system structures the output of the large language model, uses code to parse the model output and generates JSON format data to meet the requirements of business system integration. The specific steps are as follows:
[0108] a. Parse the lexical, part-of-speech and syntactic analysis results generated by the model through a dedicated parsing script to ensure the uniformity and integrity of the data format.
[0109] b. During the parsing process, the system will add necessary annotation information to each word, such as part-of-speech tags, syntactic dependencies, and related grammatical attributes.
[0110] c. Map the tokens and necessary annotation information to standard JSON fields.
[0111] To more clearly illustrate the role of JSON format data, the following shows examples of JSON format obtained by different tasks:
[0112] (1) Lexical analysis
[0113]
[0114]
[0115]
[0116] S33: Business integration and application of JSON data
[0117] According to business needs, the generated JSON format data can be used as the system output for subsequent business system integration. Different systems can use this data to support various downstream applications, such as text classification, grammar checking, and intelligent search. The system's part-of-speech recognition results can flexibly adapt to multi-language and multi-domain environments, ensuring efficient support for front-end display and back-end business logic processing, improving user experience and system scalability. Examples of business integration and application support applications of JSON data are as follows:
[0118] Lexical analysis system: JSON data can be integrated into programming language processing tools to support code parsing, syntax checking, and smart prompts. Figure 4As shown in the figure, in the lexical analysis stage, the large language model uses its powerful language generation and understanding capabilities to accurately analyze the vocabulary in the input text to meet the needs of multiple languages and multiple fields.
[0119] Part-of-speech tagging system: The tagging results can be used for text classification, sentiment analysis and other NLP tasks, providing upstream and downstream support for the model. Figure 5 As shown in the figure, in the part-of-speech tagging stage, through fine-tuning of a small amount of annotated data, the large language model greatly improves the automation level of part-of-speech tagging and reduces the dependence on traditional rules and a large amount of annotated data.
[0120] Syntax analysis system: The generated syntax tree data can be integrated into advanced NLP systems to support tasks such as Text-to-SQL, and realize the parsing of natural language queries and the automatic generation of database queries. Figure 6 As shown in the figure, the syntactic analysis function supports the model to generate reasonable syntax trees for complex syntactic structures and achieve a higher level of semantic understanding.
[0121] This embodiment uses a large language model as the initial architecture, and makes important improvements to the model according to the characteristics of the part-of-speech recognition task, which can efficiently analyze texts of natural languages and programming languages. This model: 1. Reduces the amount of parameter updates through low-rank adaptive fine-tuning, improves training efficiency and computing resource utilization; 2. Improves the adaptability of the model to multiple tasks and multiple languages by designing prompt words, and ensures the normalization of output results; 3. Generates a lightweight model through model distillation optimization to ensure efficient deployment of the model on mobile terminals and resource-constrained devices. The present invention fully taps the potential of large language models in multilingual part-of-speech recognition and achieves good results.
[0122] The present invention utilizes the powerful semantic understanding and context-awareness capabilities of the large language model to greatly simplify the process of part-of-speech recognition. Most of the intermediate processing procedures are encapsulated within the large language model, thereby reducing the dependence on rule design and manual labeling. At the same time, it can quickly adapt to different fields and business needs through low-rank adaptive fine-tuning, and combine the prompt words based on business needs to guide the model to generate part-of-speech tagging and syntactic analysis results, so that it can output structured or semi-structured information that meets the needs of specific scenarios.
[0123] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0124] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0125] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A multilingual universal part-of-speech recognition method based on a large language model, characterized in that: The following steps are involved: Preprocess the received raw data so that each input question corresponds to a reference answer, and fine-tune the large language model by designing prompt words for specific tasks; The fine-tuned large language model is distilled to form a lightweight model, which is used to recognize common parts of speech in multiple languages and parse the recognition results in code.
2. According to claim 1, a multilingual universal part-of-speech recognition method based on a large language model is characterized by: In the process of preprocessing the received raw data, the received raw data is preprocessed by dividing the paragraphs, removing noise and special symbols, performing word segmentation, and obtaining text representation, and some of the preprocessed data are labeled so that each input question corresponds to a reference answer.
3. The multilingual general part-of-speech recognition method based on a large language model according to claim 2, characterized in that: In the process of preprocessing, the specific steps of preprocessing include: Paragraph division: break the input text into logical paragraphs; Remove noise and special symbols: Remove noise data, delete meaningless special symbols, and keep only the punctuation marks required for the sentence; Perform tokenization: split the text into tokens; Obtain text representation: Convert the segmented text into a numerical representation that the model can process.
4. The multilingual general part-of-speech recognition method based on a large language model according to claim 3, characterized in that: In the process of fine-tuning the large language model, the large language model is fine-tuned through a low-rank adaptation method.
5. The multilingual general part-of-speech recognition method based on a large language model according to claim 4, characterized in that: In the process of designing prompt words, prompt words are designed for lexical analysis tasks, part-of-speech tagging tasks, and syntactic analysis tasks, respectively. For lexical analysis tasks: by designing prompt words, the model can identify the categories, word forms and language roles of words from the text and generate lexical information; For part-of-speech tagging tasks: By designing prompt words, the model generates part-of-speech tagging information and adds labels to each word for accurate tagging in a multilingual environment; For syntactic analysis tasks: By designing prompt words, the model generates a syntactic tree structure to analyze the grammatical relationship and dependency structure between words.
6. The multilingual general part-of-speech recognition method based on a large language model according to claim 5, characterized in that: In the process of coding and parsing the recognition results, the recognition results are converted into JSON format data.
7. The multilingual general part-of-speech recognition method based on a large language model according to claim 6, characterized in that: In the process of converting the recognition results into JSON format data, the lexical, part-of-speech and syntactic analysis results generated by the model are parsed, annotation information is added to each word, and this information is mapped to a standard JSON field, where the annotation information includes part-of-speech tags, syntactic dependencies and related grammatical attributes.
8. A multilingual universal part-of-speech recognition system based on a large language model, used to implement a multilingual universal part-of-speech recognition system based on a large language model as described in any one of claims 1 to 7, characterized in that: include: The model fine-tuning module is used to pre-process the received raw data so that each input question corresponds to a reference answer and fine-tune the large language model by designing prompt words for specific tasks; The multilingual general part-of-speech recognition module is used to distill the fine-tuned large language model to form a lightweight model for multilingual general part-of-speech recognition and to codedly parse the recognition results.
Citation Information
Patent Citations
Knowledge graph construction method and device, equipment and storage medium
CN117952208A
Text processing method, text generation model training method and model training method
CN118377906A
E-commerce vertical domain large language model training method and device
CN118394905A
Cited By
Job fund collection information extraction method and system based on large language model
CN120296799A