Multi-level text label classification method based on large model prompt project
By employing a rule-based prompting engineering approach, utilizing text semantic decomposition and dictionary mapping, the problem of unstable prompt word construction in multi-level text label classification is solved, achieving high-quality and stable multi-level text label classification.
Patent Information
- Application Number
- CN202511486972.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-10
AI Technical Summary
In existing technologies for multi-level text label classification, the construction defects of prompt words lead to unstable responses from large models and inaccurate multi-level label classification results.
A rule-based suggestion engineering approach is adopted, which constructs single-level and two-level text tag suggestions through text semantic decomposition and dictionary mapping, and achieves multi-level text tag classification by combining bottom-up traversal of the tree-structured tag dictionary.
It achieves high-quality and stable multi-level text label classification results, ensuring the accuracy and consistency of the classification results.
Smart Images

Figure CN121502001A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text label classification, and particularly relates to a multi-level text label classification method based on large model prompt engineering. BACKGROUND
[0002] The generative general large language model can be widely applied to natural language processing tasks, information retrieval and question answering, education, business and office, etc. However, in the development and application of special tasks, the generative general large language model needs to be manually corrected to meet the application of specific tasks. At present, the methods for adapting the general large language model to specific tasks mainly include three types: prompt engineering, agent application and large model fine-tuning. The prompt engineering mainly guides the large language model to generate output that meets the requirements of specific tasks by designing and optimizing input prompt words. Compared with the last two methods such as agent application and large model fine-tuning, the prompt engineering has a lower threshold, but through the optimization of the structure and use rules of the prompt words, a high upper limit can be reached.
[0003] The text label classification technology mainly implements algorithm to carry out directional matching of the text belonging to the category, and can provide strong support for the management and application of text data through the text label classification technology. Its main application is embodied in the following aspects: (1) efficient information retrieval and management: by matching the text with the corresponding label category, users can quickly find the text content they are interested in, and at the same time, the enterprise can realize the solid support of a large number of document classification management and archiving work; (2) high-quality content personalized recommendation: in the APP application with a large number of customers at present, the developer will form various text browsing labels according to the user's browsing content, and then the APP application will recommend similar content to the user according to these text browsing labels. At the same time, precise positioning can also be implemented to push appropriate advertisements to the user to maximize the commercial value of advertising marketing; 3) sentiment analysis and public opinion monitoring: enterprises can know the user's emotional tendency (such as positive, neutral or negative) to the product according to the real-time text feedback of the user, so as to adjust the business strategy faster. Public security organs of the government and public relations companies of enterprises can also carry out public opinion detection according to the text content on social media, so as to realize the guidance of positive feedback and timely good network space dredging. 4) academic research and intelligent application: the text classification technology is of great benefit to the academic field. Precise text classification technology can help researchers quickly classify and screen a large number of academic literature, and find reference literature related to their research direction. In intelligent customer service and chat robots and other intelligent applications, the user's problems can be quickly classified, so as to shorten the response time and accurately match the problem field, and realize more efficient problem answering and information searching.
[0004] The processing of text label classification usually uses a method based on machine learning and deep learning, wherein the machine learning-based method mainly performs linear classification through models such as Naive Bayes and support vector machine, however, the method faces difficulties in complex feature engineering construction and is difficult to match complex semantics;
[0005] The deep learning-based method mainly constructs a deep neural network such as a recurrent neural network (RNN) and a convolutional neural network (CNN), wherein the RNN and the CNN both face difficulties in training and are sensitive to part of the hyperparameters, and require a trainer to have more training experience. The text label classification method based on the prompt engineering adopted by the present application mainly faces the problem that the construction defects of the prompt words may cause unstable answers of the large model, and the multi-level text label classification based on the prompt words also faces the difficulty that too many prompt words may cause the large model to be difficult to produce accurate results. SUMMARY
[0006] The purpose of the present application is to provide a multi-level text label classification method based on large model prompt engineering, which uses a system integrated prompt engineering method when processing multi-level text label classification, generates high-quality and stable text classification answers through only one round of large model dialogue, and solves the problems in the background technology.
[0007] TECHNICAL SOLUTION The multi-level text label classification method based on large model prompt engineering provided by the present application comprises the following steps:
[0008] (1) constructing a single-level text label prompt word to guide the large model to perform semantic segmentation and label matching on the text;
[0009] (2) constructing a two-level text label prompt word to guide the large model to identify a two-level label and map it to a one-level label through a dictionary;
[0010] (3) combining the single-level prompt word and the label dictionary to realize multi-level label classification from bottom to top.
[0011] Further, step (1) is specifically as follows: preprocessing the text, extracting nouns and verbs; performing semantic similarity analysis on the extracted words and pre-defined general classification labels for matching; counting the number of matches and outputting label results that meet the set threshold.
[0012] Further, step (2) is specifically as follows: identifying a two-level label in the text; finding a corresponding one-level label through a pre-defined {two-level label: one-level label} dictionary.
[0013] Further, step (3) is specifically as follows: completing bottom layer label classification based on single-level prompt word; traversing tree-shaped label dictionary from bottom to top, and mapping to higher level label step by step to complete multi-level classification.
[0014] The multi-level text label classification system based on large model prompt engineering comprises a single-level text label prompt word construction module for constructing prompt words for guiding a large model to perform single-level label classification according to text content; a two-level text label prompt word construction module for rewriting based on the single-level text label prompt word to construct prompt words for guiding the large model to recognize two-level labels; and a multi-level text label classification module for traversing the single-level text label prompt word and a back-end label dictionary to realize multi-level label classification.
[0015] Further, the single-level text label prompt word construction module comprises a text preprocessing unit for performing semantic segmentation on input text to extract nouns and verbs; a label matching unit for matching the extracted words with pre-defined general classification labels; and a counting and threshold judgment unit for counting the matching times and outputting label results meeting a set threshold.
[0016] Further, the general classification labels comprise current affairs news, social information, food and cooking, clothing and makeup, housing and consumption, digital, finance, workplace, emotion, travel and tourism, folk custom, sports and fitness, interest and handicraft, game, agriculture, information technology, engineering, mathematics, physics, chemistry, astronomy, geography, biology, medicine, economics, management, law, sociology, anthropology, language and character, psychology, politics and diplomacy, military science, history, literature, plastic arts, performing arts, television and film, art theory, philosophy, and religion.
[0017] Further, the two-level text label prompt word construction module comprises a two-level label recognition unit for guiding the large model to recognize two-level labels; and a dictionary mapping unit for mapping the recognized two-level labels to corresponding one-level labels according to a pre-defined {two-level label: one-level label} dictionary.
[0018] Further, the dictionary used by the dictionary mapping unit includes: a first label set including: things around, life consumption, application technology, mathematics and natural science, social science and norm, history and humanistic art, philosophy and religion; a second label set including: current affairs news, social information, food and cooking, clothing and makeup, housing and consumption, digital, finance, workplace, emotion, travel and tourism, folk custom, sports and fitness, interest and handicraft, game, agriculture, information technology, engineering, mathematics, physics, resource and environment, chemistry, astronomy, geography, biology, medicine, economics, management, law, sociology, anthropology, language and character, psychology, politics and diplomacy, military science, history, literature, plastic art, performing art, television and film, art theory, philosophy, religion; and a mapping relationship dictionary between the second label and the first label.
[0019] Further, the multi-level text label classification module includes: a metadata classification unit for completing classification of the lowest level label based on single-level prompt words; and a dictionary traversal unit for traversing the tree-shaped label dictionary from bottom to top, and mapping to labels of higher levels step by step.
[0020] Advantages: Compared with the prior art, the present application has the following obvious advantages: the present application adopts a rule-based way to expand the construction of prompt words, and through the implementation of text semantic decomposition, the text is disassembled, and the atomic-level words are matched with labels to achieve relatively stable results; for multi-level text label classification, a dictionary is constructed to realize the lowest level of text label matching, and then the different levels of labels are determined through bottom-up traversal, and through only one round of large model dialogue, the construction of prompt words and the traversal of label dictionary are combined to realize the classification of first-level, second-level and multi-level labels, and the stable and accurate classification results are ensured. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The flowchart of the present application. DETAILED DESCRIPTION
[0022] The technical solutions of the present application will be further described below in combination with the drawings.
[0023] 1. As shown in the figure, the embodiment of the present application provides a multi-level text label classification method based on large model prompt engineering, the following steps: Figure 1 Step 1: Selecting text content according to text labels, wherein the text content should cover the range of text labels as much as possible. Based on single-level text labels, write prompt words to guide the large model to construct a thinking chain, wherein the prompt words should be able to make the large model output high-quality text label classification results stably;
[0024]
[0025] Step 2, based on the prompt word constructed by the single-level text label, the prompt word facing the two-level text label is constructed, the large model is guided to recognize the multi-dimensional label from bottom to top, and the stable output of high-quality text label classification result is realized;
[0026] Step 3, based on the prompt word constructed by the single-level text label and the back-end traversal label dictionary construction process, first, the text label metadata accurate classification is met, and then the tree-shaped label dictionary is traversed from bottom to top according to the text label metadata classification result, so as to realize the high-quality text label classification result of multi-level label.
[0027] Among them, step 1 realizes the full coverage of general classification labels for text content, and the prompt word is constructed to realize language atomic level analysis for text content. The general classification labels are shown in Table 1.
[0028] Table 1 General classification label elements
[0029]
[0030] For general text based on large language model for label classification and recognition, the difficulty lies in that only simple prompt is applied to the large model, which cannot achieve stable effect, that is, the classification label result output each time is not stable. Therefore, by using rules to guide the large model to implement semantic segmentation for text, and according to the classification label, the semantic matching of the segmented atomic level words is carried out, when the matching times meet the set threshold condition, the stable and accurate label classification result is generated.
[0031] In step 2, the large model is guided to recognize the multi-dimensional label from bottom to top, and the process is as follows: first, the two-level label is recognized according to the prompt word, then the dictionary with the format {two-level label: one-level label} is constructed according to the traversal, and the corresponding one-level label is found out. The one-level label, two-level label and the corresponding dictionary of {two-level label: one-level label} used by the present application are shown in Table 2.
[0032] Table 2 Two-level label corresponding elements
[0033] In step 3, for multi-level label text classification, the present application first completes the classification of the bottommost label based on the prompt word in step 1, then the return result of label classification is corresponding to the corresponding label dictionary, from bottom to top, the tree-shaped dictionary is traversed layer by layer, and the classification of multi-level label is realized.
Claims
1. A multi-level text label classification method based on large model prompting engineering, characterized in that, Includes the following steps: (1) Construct single-level text label prompts to guide the large model to perform semantic segmentation and label matching of the text; (2) Construct secondary text label prompts to guide the large model to recognize secondary labels and map them to primary labels through a dictionary; (3) Combine single-level prompt words and tag dictionary to achieve multi-level tag classification by bottom-up traversal.
2. The multi-level text label classification method based on large model prompting engineering according to claim 1, characterized in that, Step (1) is as follows: preprocess the text to extract nouns and verbs; perform semantic similarity matching between the extracted words and predefined general classification labels; count the number of matches and output the label results that meet the set threshold.
3. The multi-level text tag classification method based on large model prompting engineering according to claim 1, characterized in that, Step (2) is as follows: Identify the secondary tags in the text; find the corresponding primary tags through the predefined {secondary tags: primary tags} dictionary.
4. The multi-level text tag classification method based on large model prompting engineering according to claim 1, characterized in that, Step (3) is as follows: Complete the bottom-level tag classification based on single-level prompt words; traverse the tree-like tag dictionary from bottom to top, and map it to higher-level tags step by step to complete multi-level classification.
5. A multi-level text label classification system based on large model prompting engineering, characterized in that, include: The single-level text label prompt word construction module is used to construct prompt words based on text content to guide the large model in single-level label classification; The secondary text label prompt word construction module is used to rewrite the single-level text label prompt words to construct prompt words that guide the large model to recognize secondary labels; The multi-level text tag classification module describes the single-level text tag prompts and the backend tag dictionary traversal mechanism to achieve multi-level tag classification.
6. A multi-level text label classification system based on large model prompting engineering according to claim 5, characterized in that, The single-level text label prompt word construction module includes: a text preprocessing unit for semantic segmentation of the input text, extracting nouns and verbs; a label matching unit for performing semantic similarity analysis on the extracted words and matching them with predefined general classification labels; and a counting and threshold judgment unit for counting the number of matches and outputting label results that meet the set threshold.
7. A multi-level text label classification system based on large model prompting engineering according to claim 6, characterized in that, Common category tags include: Current Affairs, Social Information, Food & Cooking, Clothing & Makeup, Housing & Consumption, Digital, Finance, Workplace, Relationships, Travel & Tourism, Folklore, Sports & Fitness, Hobbies & Crafts, Games, Agriculture, Information Technology, Engineering, Mathematics, Physics, Chemistry, Astronomy, Geography, Biology, Medicine, Economics, Management, Law, Sociology, Anthropology, Language & Writing, Psychology, Politics & Diplomacy, Military Science, History, Literature, Visual Arts, Performing Arts, Television & Film, Art Theory, Philosophy, and Religion.
8. A multi-level text label classification system based on large model prompting engineering according to claim 5, characterized in that, The secondary text label prompt word construction module includes: a secondary label recognition unit to guide the large model to recognize secondary labels; and a dictionary mapping unit to map the recognized secondary labels to the corresponding primary labels according to the predefined {secondary label: primary label} dictionary.
9. A multi-level text label classification system based on large model prompting engineering according to claim 8, characterized in that, The dictionary used by the dictionary mapping unit includes: a set of first-level tags including: everyday things, daily life and consumption, applied technology, mathematics and natural sciences, social sciences and norms, history and humanities, philosophy and religion; a set of second-level tags including: current affairs and news, social information, food and cooking, clothing and cosmetics, housing and consumption, digital, finance, workplace, emotions, travel and tourism, folklore, sports and fitness, hobbies and handicrafts, games, agriculture, information technology, engineering, mathematics, physics, resources and environment, chemistry, astronomy, geography, biology, medicine, economics, management, law, sociology, anthropology, language and writing, psychology, politics and diplomacy, military science, history, literature, visual arts, performing arts, television and film, literary theory, philosophy, and religion; and a dictionary of mapping relationships between second-level tags and first-level tags.
10. A multi-level text label classification system based on large model prompting engineering according to claim 5, characterized in that, The multi-level text tag classification module includes: a metadata classification unit for classifying the lowest-level tags based on single-level prompts; and a dictionary traversal unit for traversing the tree-like tag dictionary from bottom to top, mapping each level to higher-level tags.