Method for identifying error tag in text information based on large language model
By using large language models and machine learning technology to automatically identify and correct incorrect labels in product information, the problem of insufficient generalization ability in existing technologies is solved, the accuracy and efficiency of labels are improved, and the user experience and information management are optimized.
Patent Information
- Application Number
- CN202510925011.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have limited generalization capabilities and insufficient understanding of complex text structures when identifying and correcting erroneous labels in text information, which affects the accuracy and efficiency of information retrieval.
By using a large language model combined with part-of-speech tagging and natural language generation technology, through training and verification, we can automatically identify and correct incorrect labels in product information, and use machine learning to optimize label quality.
It improves the accuracy and quality of tags in text information, optimizes user experience and information management efficiency of merchants, and improves the accuracy and efficiency of information retrieval.
Smart Images

Figure CN120744418A_ABST
Abstract
Description
Technical Field
[0001] The technical field of this patent involves natural language processing, large language models, and machine learning. Its core technologies include: part-of-speech tagging, natural language generation, deep neural networks, deep learning, etc. Part-of-speech tagging is used to pre-process text and mark words in the text with their corresponding parts of speech, which helps the model better understand the structure and meaning of the text and improve the accuracy of subsequent processing. After identifying possible incorrect labels, natural language generation technology is used to generate reasonable candidate labels. Deep learning is a key technology for training and optimizing large language models. Through multi-layer neural networks and nonlinear activation functions, the model can automatically learn the mapping relationship between text and labels. Through supervised learning or unsupervised learning methods, the model can learn the correctness and rationality of the label and predict the correct label based on the text content. These core technologies are intertwined and mutually reinforcing, and together they have promoted the development of natural language processing, large language models, and machine learning.
[0002] Methods based on large language models can better utilize this information and knowledge, for example in intelligent customer service systems, various text classification and annotation tasks, social media platforms, intelligent customer service, and chatbots. They can be applied to related e-commerce platforms, intelligent customer service systems, content management systems, and other related fields, helping to improve the accuracy and efficiency of text processing. Background Art
[0003] With the development of the internet and social media, text information has become a vital vehicle for information transmission and sharing, with its application becoming increasingly widespread across various fields. Tags, as a crucial component of text information, are crucial for information classification, retrieval, and management. However, due to various reasons, such as manual annotation errors and limitations of automatic tag generation algorithms, text information often contains incorrect tags. These incorrect tags not only affect the accuracy and efficiency of information retrieval but can also mislead user decisions.
[0004] Traditional methods for identifying mislabeling are mostly based on rule matching, statistical models, or traditional machine learning algorithms. These methods often face challenges when processing large-scale, diverse text data, such as limited generalization and insufficient understanding of complex text structures. With the rapid development of deep learning technology, large language models have made significant progress in natural language processing. By pre-training on large corpora, large language models can learn rich linguistic knowledge and contextual information, and possess powerful semantic understanding and generation capabilities. This provides a new solution for mislabeling.
[0005] Error label identification methods based on large language models can fully leverage the model's generalization capabilities and ability to understand complex text structures to more accurately identify incorrect labels in text. Furthermore, combined with techniques such as part-of-speech tagging and natural language generation, recognition accuracy and efficiency can be further improved. Therefore, research and development of error label identification technology based on large language models is of great significance for improving the accuracy of text information processing, optimizing user experience, and promoting the development of the field of natural language processing. Advances in this field will not only enhance the accuracy and efficiency of information retrieval but also provide strong technical support for various text information management and analysis systems. Summary of the Invention
[0006] This paper provides a method for identifying mislabels in text based on a large language model. This method accurately identifies mislabels in text by training a large language model to capture the deep semantic relationship between text and labels. Its key technical feature is its ability to use a large language model to identify and correct potential mislabels in product information, improving label quality and data accuracy.
[0007] The specific steps are as follows:
[0008] S1: Data Preparation
[0009] 1.1. Collect product information data, clean it, and preprocess it to meet the input requirements of large-scale language models.
[0010] 1.2. Build data structures and models suitable for product information, including tag types, attributes, and associations.
[0011] S2: Error label identification
[0012] 2.1. Use a large language model to pre-train the training data to understand common language patterns and label errors in product information.
[0013] 2.2. Use the pre-trained language model to identify and classify the labels in the product information and identify labels that may be incorrect.
[0014] S3: Error label correction
[0015] 3.1. Further analyze and verify the identified error tags to confirm whether they are wrong.
[0016] 3.2. Use machine learning technology to correct incorrect labels or suggest replacements to improve the accuracy and quality of labels.
[0017] Through the above steps, the invention can build an automated system for identifying and correcting incorrect labels in product information. This system, which combines large-scale language models with machine learning technology, can effectively improve the accuracy and quality of labels in product information, thereby optimizing the user shopping experience and improving the information management efficiency of merchants. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a method of identifying incorrect labels in text information based on a large language model. DETAILED DESCRIPTION
[0019] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0020] Example 1
[0021] like Figure 1 As shown, the present invention is a method for identifying incorrect labels in text information based on a large language model, which can be applied to electronic product label recognition and correction, including:
[0022] S1. Data Preparation
[0023] 1.1 Collect a dataset containing electronic product information and labels, and perform data cleaning to remove noise and irregular labels.
[0024] 1.2 Build the electronic product information data structure, including product model, description, label, etc., and prepare data samples for training.
[0025] S2. Identification of incorrect labels
[0026] 2.1 Use a large language model to pre-train the training data to learn common language patterns and mislabels in electronic product information.
[0027] 2.2 Use the pre-trained language model to identify the labels in the electronic product information and mark the labels that may be incorrect.
[0028] S3. Correction of incorrect labels
[0029] 3.1 Verify and analyze the identified potentially erroneous labels to confirm whether they contain any errors.
[0030] 3.2 Based on machine learning algorithms, correct incorrect labels or make replacement suggestions to improve label accuracy and data quality.
[0031] By following these steps, we can build an automated system for identifying and correcting incorrect product labels. This system, which combines large-scale language models with machine learning techniques, can effectively improve the accuracy and quality of product labels, thereby optimizing the user shopping experience and enhancing merchants' information management efficiency.
[0032] Example 2
[0033] like Figure 1 As shown, the present invention is a method for identifying incorrect labels in text information based on a large language model, which can be applied to food label recognition and correction, including:
[0034] S1. Data Preparation
[0035] 1.1 Collect a dataset containing food information and labels, and perform data cleaning to remove noise and irregular labels.
[0036] 1.2 Build a food information data structure, including food name, ingredients, labels, etc., and prepare data samples for training.
[0037] S2. Identification of incorrect labels
[0038] 2.1 Use a large language model to pre-train the training data to learn common language patterns and mislabels in food information.
[0039] 2.2 Use the pre-trained language model to identify labels in food information and mark labels that may be incorrect.
[0040] S3. Correction of incorrect labels
[0041] 3.1 Verify and analyze the identified potentially erroneous labels to confirm whether they contain any errors.
[0042] 3.2 Based on machine learning algorithms, correct incorrect labels or make replacement suggestions to improve label accuracy and data quality.
[0043] Through the above steps, the present invention can realize the automatic identification and correction of food labels, effectively improving the accuracy and reliability of food labels. This is of great significance for ensuring food safety, optimizing consumer shopping experience, and improving the level of information management in the food industry.
[0044] The first embodiment focuses on the identification and correction of electronic product labels, mainly focusing on the specific attributes and label types of electronic products, such as model and specifications, to improve the accuracy of electronic product information.
[0045] The second embodiment focuses on the identification and correction of food labels, taking into account specific attributes of food such as ingredients, nutritional information, and label types to improve the quality and credibility of food information.
[0046] In summary, Example 1 focuses on improving the accuracy and quality of electronic product information, optimizing the user shopping experience by correcting key labels such as model and specifications. Example 2, on the other hand, focuses on the accuracy and reliability of food labels, ensuring the accuracy of food information to safeguard food safety and enhance information management in the food industry. Both use automated methods based on large language models and machine learning techniques to identify and correct incorrect labels in product information. While addressing different types of products and label attributes, the common goal is to improve the accuracy and credibility of product information.
[0047] The foregoing is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various improvements and modifications without departing from the technical principles of the present invention. Such improvements and modifications should also be considered within the scope of protection of the present invention. The scope of protection of the present invention should be defined by the technical features described in the appended claims, which constitute a part of this application.
Claims
1. A method for identifying incorrect labels in text information based on a large language model, characterized in that: The following steps are involved: 1) Collect product information data, clean and pre-process the data to meet the input requirements of large-scale language models, and build data structures and models suitable for product information, including label types, attributes, and associations; 2) Use a large language model to pre-train the training data to understand common language patterns and label errors in product information. Use the pre-trained language model to identify and classify labels in product information and identify possible incorrect labels. 3) Further analyze and verify the identified erroneous labels to confirm whether they are incorrect, and use machine learning technology to correct the erroneous labels or recommend replacements to improve the accuracy and quality of the labels.
2. The method according to claim 1, characterized in that The data cleaning and preprocessing steps include removing irrelevant information, standardizing text formats, processing missing values and outliers, etc.
3. The method according to claim 1 or 2, characterized in that The large-scale language model is a deep learning model based on the Transformer architecture, including but not limited to BERT, GPT, etc.
4. The method according to any one of claims 1 to 4, characterized in that In the error label correction step, machine learning techniques include but are not limited to decision trees, random forests, neural networks and other algorithms, which are used to automatically correct the error labels or provide replacement suggestions.
5. A system according to any one of claims 1 to 5, characterized in that The system includes a data preprocessing module, a large language model training module, an error label recognition module and an error label correction module. The modules work together to realize the error label recognition of text information based on the large language model.
6. Through the definition of the above claims, we can describe the content and technical features of this patent in detail. These claims involve collecting product information data, cleaning and preprocessing the data to meet the input requirements of large-scale language models, constructing data structures and models suitable for product information, using large-scale language models to pre-train training data, using pre-trained language models to identify potentially erroneous labels, further analyzing and verifying identified erroneous labels, and correcting or recommending replacements for erroneous labels.