Method and system for automatically marking text data by using large language model

Automatic text data marking through large language models solves the problems of manpower and privacy security in the field of illegal fundraising, realizes efficient and accurate data analysis and monitoring, and enhances the adaptability and data support capabilities for illegal fundraising scenarios.

CN120216682APending Publication Date: 2025-06-27FUDAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510244082.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In areas such as illegal fundraising, obtaining and labeling high-quality text data requires a lot of manpower and expert experience, and there are data privacy and security issues, which are difficult for existing technologies to effectively solve these problems.

Method used

The large language model is used to automatically mark text data. By crawling relevant data from public network text, deduplication and filtering, classification is performed, and using the propt engineering and few-shot learning method combined with the applicable scenarios, the large language model is input to generate labeling results, and the model is fine-tuned through LoRA technology to finally realize automated marking.

Benefits of technology

This method can efficiently generate high-quality illegal fundraising data, improve the accuracy and efficiency of data analysis, enhance the adaptability to illegal fundraising scenarios, provide strong data support to identify and combat bad behaviors, and maintain the stability and security of the financial market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216682A_ABST
    Figure CN120216682A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for automatically marking text data by using a large language model, and belongs to the technical field of large language models and online public opinion monitoring and marking. According to the method and the system for automatically marking the text data by using the large language model, related personnel can more effectively analyze characteristics, trends and potential risks by generating high-quality applicable scene related data and formulate accurate monitoring and intervention measures. By increasing the sample size, the generalization ability of the machine learning model can be improved, the adaptability to different application scenes can be enhanced, and the model can respond more quickly when facing diversified bad techniques in the application scenes. According to the method and the system for automatically marking the text data by using the large language model, powerful data support can be provided for identifying, early warning and fighting against bad behaviors in an applicable scene, so that the legal rights and interests of related personnel are better protected, and the stability and the safety of the market are maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of large language models and network public opinion monitoring and annotation, and particularly relates to a method and system for automatically labeling text data by using a large language model. Background Art

[0002] Data is an indispensable and important part of the development of artificial intelligence, and its quality directly affects the performance ceiling of machine learning models. [1] . In various fields, especially healthcare [2] and finance [3] fields. However, in some fields, there are data and labels, but labeling the data with corresponding labels may require a large amount of manpower and expert experience. At the same time, some data may need to be protected by privacy and cannot be directly obtained by some labeling personnel. Therefore, there is a certain barrier between the data itself and the labels. The present invention mainly focuses on the illegal fund-raising scenario.

[0003] Illegal fund-raising refers to the act of raising funds from the public or specific objects without the approval of relevant national departments for the purpose of raising funds. This kind of behavior usually attracts investors with attractive conditions such as high interest rates and high returns, aiming to obtain funds and misappropriate them. Illegal fund-raising not only violates national laws and regulations, but also may cause serious negative impacts on the social economy, inducing financial risks and trust crises for investors. The forms of illegal fund-raising are diverse, mainly including but not limited to the following:

[0004] Ponzi scheme: This is a relatively common way of illegal fund-raising, which attracts new investors' funds to pay the returns of old investors in order to deceive investors and maintain the so-called "profits".

[0005] False project investment: Some illegal fund-raising organizations will fabricate investment projects, such as real estate, cultural and creative industries, etc., and promise high returns to obtain investors' funds.

[0006] Illegal financial institutions: Some institutions self-proclaimed as "micro-loan companies", "investment management companies" have not obtained legal business licenses, but raise funds in this name.

[0007] Internet finance: With the rapid development of the Internet, many projects on network platforms have the risk of illegal fund-raising, especially the fund-raising in the name of crowdfunding, P2P, etc.

[0008] Data plays an indispensable role in the analysis and supervision of illegal fund-raising, and its quality directly affects the effectiveness of identifying and preventing illegal fund-raising activities.

[0009] Illegal fundraising data plays an indispensable role in the analysis and supervision of illegal fundraising, and its quality directly affects the effectiveness of identifying and preventing illegal fundraising activities. However, obtaining high-quality data is a complex and arduous task [4] , and as the public's concern about the social harm caused by illegal fundraising activities gradually intensifies, the government is also continuously strengthening the formulation and implementation of relevant laws and regulations. For example, the Criminal Law of the People's Republic of China clearly stipulates the punishment measures for illegal fundraising behaviors, which has triggered strict requirements for financial data management. In addition, the Anti-Money Laundering Law and related regulatory policies also emphasize the monitoring of the sources and uses of funds in financial transactions. In this context, effective data has become an important link in combating and preventing illegal fundraising activities. Only by obtaining and analyzing accurate illegal fundraising data can we improve the identification ability, reduce potential social risks, and maintain the stability of the financial market.

[0010] Rule-based illegal fundraising data tagging is a method that uses preset rules and structured patterns to identify and label illegal fundraising case data. This technology can effectively assist researchers, financial institutions, and regulatory authorities in identifying the characteristics, development trends, and potential risks of illegal fundraising behaviors. First, the rule-based data tagging method can establish a detailed rule system by analyzing the characteristics of historical illegal fundraising cases. For example, factors such as the channels of fund raising (such as online platforms, offline promotions), promised rates of return (excessively high or unrealistic rates of return), the whereabouts of funds (whether the funds flow to legal projects), and types of investors (individual investors, corporate investors) can be set. By defining these rules, researchers can effectively tag existing data, thereby providing a high-quality data set for further analysis and model training. Second, this method can also take into account the particularities of different regions and industries to generate more targeted data tagging standards. For example, laws and regulations, economic conditions, and social environments in different regions may have a significant impact on the manifestations of illegal fundraising. By incorporating these variables into the rule system, the tagging process can be more accurate, thereby helping institutions better identify and prevent potential illegal fundraising risks. However, this method also has its limitations. The first is the limitation of the rules themselves. There may be insufficient rule coverage, and it overly relies on expert experience, making it difficult to handle complex scenarios. There are also data privacy and security issues, as well as time and resource consumption.

[0011] Most machine learning-based text data tagging methods follow a two-step process. In the first step, some handcrafted features are extracted from the document (or any other text unit). In the second step, these features are fed into a classifier for prediction. Popular handcrafted features include the bag of words (BoW) and its extensions. Popular choices for classification algorithms include Bayes, Support Vector Machines (SVM), Hidden Markov Model (HMM), Gradient Boosting Trees, and Random Forests. The two-step method has several limitations. For example, relying on handcrafted features requires cumbersome feature engineering and analysis to achieve good performance. In addition, the strong dependence on domain knowledge when designing features makes it difficult to generalize this method to new tasks. Finally, these models cannot fully utilize a large amount of training data because the features (or feature templates) are predefined.

[0012] Text data labeling methods based on traditional deep learning models such as word2vec [5] , Transformer [6] , GPT-2 [7] and other models have feature learning capabilities in text data labeling. Deep learning models can automatically learn the complex features of text data, reducing the need for manual feature engineering. They can handle large-scale data, are suitable for large-scale text datasets, and can handle data volumes in the billions. They also have multi-level representations. Through a multi-layer network structure, they can capture the multi-level semantic information of the text. Compared with machine learning methods, they are end-to-end learning and can directly learn the target labels from the original text without complex preprocessing steps. However, they require a large amount of labeled data to train the model, and these models are often not ideal for our goal of labeling text data.

[0013] The text data labeling method based on the GPT API provides a novel alternative to traditional and deep learning models and has significant advantages. Different from the traditional feature engineering and model training processes, the GPT API can directly utilize pre-trained large language models. By inputting text data into the API, corresponding labels or annotations can be generated. This method greatly simplifies the data labeling process, eliminating the cumbersome feature extraction and model training steps, and also reducing the dependence on domain knowledge. However, this method also has certain disadvantages. First, it requires a stable network connection and relatively high costs. At the same time, if a large amount of data is input into the API for labeling, it involves data privacy and security.

[0014] [1] JAIN A, PATEL H, NAGALAPATTI L, et al. Overview and importance of data quality for machine learning tasks;proceedings of the Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, F, 2020[C].

[0015] [2] SLAVOVA S, ROCK P, BUSH H M, et al. Signal of increased opioid overdose during COVID-19 from emergency medical services data[J]. Drug and alcohol dependence, 2020, 214: 108176.

[0016] [3] LI Y, HUANG Z, YAN J, et al. GFTE: graph-based financial table extraction;proceedings of the Pattern Recognition ICPR International Workshops and Challenges: Virtual Event, January 10 - 15, 2021, Proceedings, Part II, F, 2021[C]. Springer.

[0017] [4] CAVANILLAS J M, CURRY E, WAHLSTER W. New horizons for a data-driven economy: a roadmap for usage and exploitation of big data in Europe[M]. Springer Nature, 2016.

[0018] [5]MIKOLOV T, CHEN K, CORRADO Greg S, et al. Efficient Estimation of Word Representations in Vector Space[J]. International Conference on Learning Representations, International Conference on Learning Representations, 2013.

[0019] [6]VASWANI A, SHAZEER N, PARMAR N, et al. Attention is All you Need[J]. Neural Information Processing Systems, Neural Information Processing Systems, 2017.

[0020] [7]RADFORD A, WU J, CHILD R, et al. Language models are unsupervised multitask learners[J]. OpenAI blog, 2019, 1(8):9. Summary of the Invention

[0021] The present invention is made to solve the above problems, and aims to provide a method and system for automatically labeling text data using large language models.

[0022] The present invention provides a method for automatically labeling text data using large language models, applicable scenarios including illegal fund-raising, finance or medicine fields, and having the following features: S10, crawling text data related to the applicable scenario from public network texts, then de-duplicating and screening it to obtain effective and high-quality public opinion text data, and then classifying it; S20, combining the effective and high-quality public opinion text data with the label system corresponding to the applicable scenario through prompt engineering and inputting it into the large language model to obtain a labeling result; S30, performing k-shot sampling on the labeling result using the data sampling greedy algorithm to obtain sampling data; S40, using the sampling data and adopting the LoRA technology to fine-tune the large language model to obtain a fine-tuned model; S50, deploying the fine-tuned model to automatically label the effective and high-quality public opinion text data.

[0023] In the method for automatically labeling text data using a large language model provided by the present invention, it may further have the following feature: Among them, in step S10, the screening is performed through ChatGPT, and the overly short content in the text data is filtered out.

[0024] In the method for automatically labeling text data using a large language model provided by the present invention, it may further have the following feature: Among them, in step S10, the effective high-quality public opinion text data is classified into software advertisements, APP introductions, user recommendations, social negative public opinions, and other public opinions.

[0025] In the method for automatically labeling text data using a large language model provided by the present invention, it may further have the following feature: Among them, step S20 includes the following sub-steps: S21, after combining the label system corresponding to the application scenario and the effective high-quality public opinion text data, a prompt template is constructed; S22, after integrating few-shot learning and the COT method into the prompt template, it is input into the large language model to obtain the labeling result.

[0026] In the method for automatically labeling text data using a large language model provided by the present invention, it may further have the following feature: Among them, the large language model includes Qwen2-7B.

[0027] In the method for automatically labeling text data using a large language model provided by the present invention, it may further have the following feature: Among them, in step S30, in k-shot sampling, each type of entity is allowed to appear at most 1250 times.

[0028] The present invention also provides a system for automatically labeling text data using a large language model, and the application scenarios include illegal fundraising, the financial or pharmaceutical fields. It has the following features, including: a data preprocessing module, which is used to crawl text data related to the application scenario from publicly available network texts, then perform deduplication and screening on it to obtain effective high-quality public opinion text data, and then classify it; a Prompt generation module, which is used to combine the label system corresponding to the application scenario and the effective high-quality public opinion text data, construct a prompt template and integrate few-shot learning and the COT method; a large language model, which is used to process the prompt template to obtain the labeling result; an inference module, which is used to perform k-shot sampling on the labeling result using the data sampling greedy algorithm to obtain the sampling data; a fine-tuning module, which uses the sampling data and adopts the LoRA technology to fine-tune the large language model to obtain a fine-tuned model; a labeling module, which is used to deploy the fine-tuned model to automatically label the effective high-quality public opinion text data.

[0029] Functions and effects of the invention

[0030] A method and system for automatically labeling text data using a large language model according to the present invention, applicable scenarios include illegal fundraising, finance or medicine fields, and have the following characteristics, including the following steps: S10, crawl text data related to the applicable scenario from publicly available network texts, then deduplicate and screen it to obtain effective and high-quality public opinion text data, and then classify it; S20, combine the effective and high-quality public opinion text data with the label system corresponding to the applicable scenario and input it into the large language model after prompt engineering to obtain the annotation result; S30, perform k-shot sampling on the annotation result using the data sampling greedy algorithm to obtain the sampled data; S40, utilize the sampled data and adopt the LoRA technology to fine-tune the large language model to obtain the fine-tuned model; S50, deploy the fine-tuned model to automatically label the effective and high-quality public opinion text data.

[0031] Therefore, a method and system for automatically labeling text data using a large language model of the present invention can generate high-quality data related to the applicable scenario, enabling relevant personnel to more effectively analyze its characteristics, trends, and potential risks, and thus formulate more accurate monitoring and intervention measures. Increasing the sample size not only helps improve the generalization ability of the machine learning model but also enhances the adaptability to different situations in the applicable scenario. This enables the model to respond more agilely when faced with diverse bad practices in the applicable scenario.

[0032] A method and system for automatically labeling text data using a large language model of the present invention can provide strong data support for identifying, warning, and combating bad behaviors in the applicable scenario, thereby better protecting the legitimate rights and interests of relevant personnel and maintaining the stability and security of the market. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic flow chart of obtaining effective and high-quality public opinion text data in a method for automatically labeling text data using a large language model according to an embodiment of the present invention;

[0034] Figure 2 is a specific schematic flow chart of obtaining the annotation result in a method for automatically labeling text data using a large language model according to an embodiment of the present invention;

[0035] Figure 3 is an example of the annotation result obtained in step S20 in a method for automatically labeling text data using a large language model according to an embodiment of the present invention;

[0036] Figure 4 is the pseudocode of the data sampling greedy algorithm used in step S30 in a method for automatically labeling text data using a large language model according to an embodiment of the present invention. Detailed implementation manners

[0037] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following embodiments will specifically elaborate on a method and system for automatically marking text data using a large language model in conjunction with the accompanying drawings.

[0038] <Embodiment>

[0039] Figure 1 is a schematic flowchart of obtaining effective and high-quality public opinion text data in a method for automatically marking text data using a large language model according to an embodiment of the present invention; Figure 2 is a specific schematic flowchart of obtaining a marking result in a method for automatically marking text data using a large language model according to an embodiment of the present invention.

[0040] As Figure 1 and 2 shown, this embodiment provides a method for automatically marking text data using a large language model.

[0041] The applicable scenarios of the method for automatically marking text data using a large language model in this embodiment include the fields of illegal fund-raising, finance or medicine. Specifically, in this embodiment, the applicable scenario is selected as the field of illegal fund-raising.

[0042] The method for automatically marking text data using a large language model in this embodiment includes the following steps:

[0043] S10, After crawling text data related to the applicable scenario from publicly available network texts, perform deduplication and screening on it to obtain effective and high-quality public opinion text data, and then classify it, including the following sub-steps S11 to S13:

[0044] S11, Crawl text data related to the applicable scenario from publicly available network texts to obtain initial public opinion texts.

[0045] S12, After performing deduplication on the initial public opinion texts, use chatGPT to perform quality screening on the deduplicated initial public opinion texts to filter out data with too short content and low text quality, and obtain effective and high-quality public opinion text data.

[0046] S13, Classify the effective and high-quality public opinion text data into five categories: (1) Software advertisements, (2) APP introductions, (3) User recommendations, (4) Social negative public opinions, (5) Other public opinions.

[0047] S20, Combine the effective and high-quality public opinion text data with the label system corresponding to the applicable scenario using prompt engineering [8]After inputting into the large language model, the annotation results are obtained, including the following sub-steps S21 - S22:

[0048] S21, After combining the label system corresponding to the applicable scenario (illegal fund-raising label system) and the effective high-quality public opinion text data, manually construct a prompt template.

[0049] S22, Integrate the few-shot learning and COT (Chain of Thought) methods in the prompt template [9] After inputting into the large language model (Qwen2-7B), the annotation results are obtained. An example of the annotation results obtained through this step is as Figure 3 shown.

[0050] S30, Use the data sampling greedy algorithm to perform k-shot sampling on the annotation results to obtain the sampling data, including the following sub-steps S31 - S32:

[0051] S31, Given a labeled dataset

[0052] where, X i represents the public opinion data of the i-th data, Y i represents the label of the i-th data, i represents the i-th data, and |D| represents the number of samples in D.

[0053] S32, Use the data sampling greedy algorithm as shown in Figure 4 to perform data sampling on the dataset D to extract k-shot examples.

[0054] Among them, considering tagging as a named entity recognition (NER) task, in the context of the named entity recognition (NER) task, the standard k-shot samples refer to obtaining exactly k entity instances for each category

[10] . However, meeting this strict requirement may be challenging, especially when dealing with data with imbalanced label distributions. To address this challenge, this embodiment draws inspiration from another work

[11] and introduces a relaxation of the standard. This relaxation allows the entities of each category to appear at most 1.25k times in the k-shot scenario. This embodiment takes into account the inherent bias in the entity category distribution in the real world while trying to maintain the k-shot setting.

[0055] S40, Utilize the sampling data and adopt the LoRA

[12] (Low-Rank Adaptation) technology to fine-tune the large language model (Qwen2-7B)

[13] to obtain a fine-tuned model.

[0056] S50. Deploy the fine-tuned model to automatically label the effective and high-quality public opinion text data.

[0057] This embodiment also provides a system for automatically labeling text data using a large language model, and the applicable scenarios include illegal fund-raising, finance, or the pharmaceutical field. Specifically, in this embodiment, the applicable scenario is selected as the illegal fund-raising field.

[0058] A system for automatically labeling text data using a large language model in this embodiment uses the method for automatically labeling text data using a large language model provided above, and includes a data preprocessing module, a Prompt generation module, a large language model, an inference module, a fine-tuning module, and a labeling module.

[0059] The data preprocessing module is used to crawl text data related to the applicable scenario from publicly available network texts according to the method in step S10, then deduplicate and screen it to obtain effective and high-quality public opinion text data, and then classify it.

[0060] The Prompt generation module is connected to the data preprocessing module and is used to combine the label system corresponding to the applicable scenario with the effective and high-quality public opinion text data according to the method in step S20, then construct a prompt template and integrate the few-shot learning and COT methods.

[0061] The large language model is connected to the Prompt generation module and is used to process the prompt template according to the method in step S20 to obtain the annotation result.

[0062] The inference module is connected to the large language model and is used to perform k-shot sampling on the annotation result using the data sampling greedy algorithm according to the method in step S30 to obtain the sampled data.

[0063] The fine-tuning module is connected to the inference module and the large language model and is used to utilize the sampled data and adopt the LoRA technology to fine-tune the large language model to obtain the fine-tuned model according to the method in step S40.

[0064] The labeling module is connected to the fine-tuning module and is used to deploy the fine-tuned model according to the method in step S50 to automatically label the effective and high-quality public opinion text data.

[0065] <Test case>

[0066] This test case uses the system for automatically labeling text data using a large language model provided by the embodiment, and performs an experiment on the accuracy of the annotation result for the validation set EvalSet according to the method for automatically labeling text data using a large language model provided by the embodiment.

[0067] Meanwhile, the annotation result accuracies of the GPT API and DeepSeek API under the existing technology for the validation set EvalSet were compared, and the comparison results are shown in Table 1 below.

[0068] Table 1 (Comparison of the annotation result accuracies of the system / method of the embodiment, GPT API, and DeepSeek API for the validation set EvalSet)

[0069] Dataset Metric GPT API DeepSeek API System / Method of the Embodiment EvalSet ACC 0.2014 0.3819 0.5183

[0070] As shown in Table 1, the annotation result accuracy of the system / method of the embodiment is significantly higher than that of the methods under the existing technology.

[0071] Functions and effects of the embodiment

[0072] In step S22 of this embodiment, by integrating few-shot learning and COT (Chain of Thought) methods in the prompt template, the accuracy and coherence of the generated results are significantly improved. Few-shot learning uses a small number of examples to guide the model, enabling it to perform well on new tasks. The COT method, on the other hand, improves the model's problem-solving ability by constructing a chain of thought for step-by-step reasoning. Combining these two methods optimizes the generated results to make them more accurate and reliable. [9] (Chainof Thought, thinking chain) method, significantly improved the accuracy and coherence of the generated results. Few-shot learning uses a small number of examples to guide the model, enabling it to perform well on new tasks. The COT method, on the other hand, improves the model's problem-solving ability by constructing a chain of thought for step-by-step reasoning. Combining these two methods optimizes the generated results to make them more accurate and reliable.

[0073] The method of step S33 of this embodiment makes it possible to better capture the actual distribution of real-world data.

[0074] In step S40 of this embodiment, the LoRA technique enables the model to quickly adapt to specific tasks with less data and computing resources while retaining the powerful general representation ability of the pre-trained model by adding a small number of trainable parameters on top of the pre-trained model.

[0075] With the method and system for automatically labeling text data using a large language model in this embodiment to generate high-quality illegal fund-raising data, researchers and financial regulatory agencies can more effectively analyze the characteristics, trends, and potential risks of illegal fund-raising, thereby formulating more accurate monitoring and intervention measures. Increasing the sample size not only helps improve the generalization ability of machine learning models but also enhances the adaptability to different illegal fund-raising scenarios. This enables the model to respond more agilely when faced with diverse illegal fund-raising methods.

[0076] The method and system for automatically labeling text data using a large language model in this embodiment provide strong data support for identifying, warning, and combating illegal fund-raising behaviors, thereby better protecting the legitimate rights and interests of investors and maintaining the stability and security of the financial market.

[0077] [8]Li, X. L., & Liang, P. (2021). Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 3811-3821). Online. http: / / dx.doi.org / 10.18653 / v1 / 2021.acl-long.353. DOI: 10.18653 / v1 / 2021.acl-long.353.

[0078] [9]Wei, J., Wang, X., Schuurmans, D., et al. (2022). Chain of Thought Prompting Elicits Reasoning in Large Language Models.

[0079]

[10] Yang, Y., & Katiyar, A. (2020). Simple and Effective Few-Shot Named Entity Recognition with Structured Nearest Neighbor Learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 6406-6416). Online. http: / / dx.doi.org / 10.18653 / v1 / 2020.emnlp-main.516.

[0080] DOI: 10.18653 / v1 / 2020.emnlp-main.516.

[0081] It should be noted that the content in [8] and

[10] has been adjusted according to the correct citation format. The original content in [8] was incomplete in the journal name and page numbers, and in

[10] the DOI was not correctly formatted in the citation. The above translation is based on the corrected content. If there are specific requirements for not modifying the original content, please let me know and I will adjust accordingly.

[11] DING N,XU G,CHEN Y,et al.Few-NERD:A Few-Shot Named EntityRecognition Dataset[C / OL] / / Proceedings of the 59th Annual Meeting of theAssociation for Computational Linguistics and the 11th International JointConference on Natural Language Processing(Volume 1:Long Papers),Online.2021.http: / / dx.doi.org / 10.18653 / v1 / 2021.acl-long.248.DOI:10.18653 / v1 / 2021.acl-long.248.

[0082]

[12] J.H,SHEN Y,WALLIS P,et al.LoRA:Low-Rank Adaptation of LargeLanguage Models.[J].arXiv:Computation and Language,arXiv:Computation andLanguage,2021.

[0083]

[13] Yang A,Yang B,Hui B,et al.Qwen2 Technical Report[OL].arXivpreprint arXiv:2407.10671,2024.

[0084] Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for automatic text data tagging using a large language model, applicable to scenarios including illegal fundraising, finance or medicine, characterized in that: The following steps are involved: S10, crawling text data related to the applicable scenario from public network texts, deduplicating and screening the text data to obtain effective and high-quality public opinion text data, and then classifying the data; S20, combining the effective high-quality public opinion text data and the label system corresponding to the applicable scenario with promptengineering and inputting them into a large language model to obtain a labeling result; S30, performing k-shot sampling on the labeling result using a data sampling greedy algorithm to obtain sampling data; S40, using the sampled data and the LoRA technology to fine-tune the large language model to obtain a fine-tuned model; S50, deploying the fine-tuning model to automatically label the effective high-quality public opinion text data.

2. The method for automatic text data tagging using a large language model according to claim 1, characterized in that: in, In step S10, the screening is performed through chatGPT, and the too short content in the text data is filtered out.

3. The method for automatic text data tagging using a large language model according to claim 1, characterized in that: in, In step S10, the effective high-quality public opinion text data is classified into software advertisements, APP introductions, user recommendations, negative social public opinions and other public opinions.

4. The method for automatic text data tagging using a large language model according to claim 1, Features: Wherein, step S20 includes the following sub-steps: S21, after combining the label system corresponding to the applicable scenario with the effective high-quality public opinion text data, construct a prompt template; S22, integrating few-shot learning and COT method into the prompt template and inputting the large language model to obtain a labeling result.

5. The method for automatic text data tagging using a large language model according to claim 1, characterized in that: in, The large language model includes Qwen2-7B.

6. The method for automatic text data tagging using a large language model according to claim 1, characterized in that: in, In step S30, in the k-shot sampling, each type of entity is allowed to appear at most 1250 times.

7. A system for automatic text data tagging using a large language model, applicable to illegal fundraising, finance or medicine, characterized in that: include: A data preprocessing module is used to crawl text data related to the applicable scenario from public network texts, remove duplicates and screen them to obtain effective and high-quality public opinion text data, and then classify them; A prompt generation module is used to combine the label system corresponding to the applicable scenario with the effective high-quality public opinion text data, construct a prompt template and integrate few-shot learning and COT methods; A large language model is used to process the prompt template to obtain a labeling result; An inference module, used for performing k-shot sampling on the annotation results using a data sampling greedy algorithm to obtain sampled data; A fine-tuning module, using the sampled data and LoRA technology to fine-tune the large language model to obtain a fine-tuning model; A labeling module is used to deploy the fine-tuning model to automatically label the effective high-quality public opinion text data.

Citation Information

Cited By

  • Large model data enhanced named entity recognition and RAG system integration method

    CN121168654A