Text data security level attribute defining method and system based on information retrieval

CN120353931APending Publication Date: 2025-07-22Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510382418.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology has low density accuracy when processing long text data and cannot be applied to few-shot scenarios. Moreover, traditional methods are difficult to accurately capture data semantics when facing information hiding methods such as steganography, homophones, multilingual transcription and style transfer, resulting in classification errors, unable to effectively handle long dependencies, resulting in information loss or forgetting.

Method used

The secret attribute definition of text data is carried out through a method based on information retrieval, including a three-stage task processing framework for relation extraction, context example generation and text classification. The relationship extraction is used to identify key semantic information, generate context examples with high semantic correlation and length compression, and use the thinking chain method to build a large model classification reasoning path to realize dense classification.

Benefits of technology

Without fine-tuning, the accuracy of long text classification is significantly improved, especially in small sample scenarios, which performs well, reduces resource consumption, improves the universality and applicability of the solution, and is suitable for large-scale commercial applications and resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353931A_ABST
    Figure CN120353931A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text security attribute processing, in particular to a text data security level attribute defining method and system based on information retrieval, and the method comprises the steps: extracting a to-be-processed text data relationship, obtaining a relationship triple in the to-be-processed text data, and constructing a query data set; performing embedding generation on text segments and labels in the annotation data set to obtain a corresponding text vector database with labels, acquiring a plurality of text segments most similar to the relation triples in the query data set from the text vector database based on similarity to form an initial context example set, and storing the initial context example set in the query data set; a context example set is obtained through content compression; and constructing a large model classification reasoning path based on the context example set by adopting a thinking chain mode, and inputting the to-be-processed text data into the security classification attribute defining large model to obtain a security classification tag of the to-be-processed text data. According to the method, under the condition that the model does not need to be finely adjusted, the secret level setting precision can be improved, and long text data secret level classification under the feed-shot scene is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text security attribute processing, and particularly relates to a method and system for defining the confidentiality level attribute of text data based on information retrieval. Background Art

[0002] The confidentiality level is a key security attribute of text data. Accurately calibrating the confidentiality level attribute of text data is the basis for realizing the minimized circulation of data according to business logic and preventing the leakage of sensitive information. However, classified text data usually has the following two characteristics: First, the annotation of classified text requires a strict approval process, and it is difficult to obtain a large amount of annotated data; second, classified texts are mostly institutional documents, technical materials, etc., and usually have a long length. Therefore, intelligent assisted classification technology needs to meet the classification requirements of long text data in few-shot scenarios.

[0003] Existing intelligent assisted classification technologies are divided into two categories: classification-based technologies and hybrid-based technologies. Among them, classification-based technologies regard the classification of confidentiality levels as a classification problem of text data, and are realized by performing binary / multi-classification on the text through machine learning or deep learning technologies; hybrid-based technologies integrate the accuracy of artificial rules and the semantic understanding ability of machine learning / deep learning, aiming to achieve efficient and accurate determination of the confidentiality level of documents through the synergistic effect of a rule base and NLP technology. This technology usually constructs a rule base to standardize the classification process, and uses machine learning / deep learning models to perform semantic analysis and feature extraction on the text, so as to capture key information in the document and realize automatic identification or matching of the confidentiality level. Through this combination method, both the clarity and interpretability of artificial rules and the powerful semantic understanding ability of machine learning / deep learning models are utilized, improving the efficiency and accuracy of the classification work.

[0004] Due to relying on a manually formulated classified rule library, the technology based on the hybrid method has inherent disadvantages in objectivity, accuracy, and efficiency when processing text data under big data conditions. The manual formulation and application of classified rules require a large amount of manpower and time, and are easily affected by subjective factors, resulting in inconsistent and inaccurate classification results. The keyword-based technology relies too much on keywords and has the deficiency of misclassification due to the inability to accurately capture the data semantics when facing information hiding means such as steganography, homophony, multilingual transcription, and style transfer. The existing classification-based and hybrid-based technologies can only achieve auxiliary classification in a limited data domain. When dealing with a small number of samples, problems such as overfitting or insufficient generalization ability may occur, making it difficult to meet the requirements of the common few-shot scenarios under big data conditions. Moreover, the current methods are mainly implemented for short text data and may achieve good results when dealing with "sentence"-level tasks. However, when dealing with "document"-level long text data, information loss or forgetting occurs due to the inability to effectively process long-term dependencies, resulting in the "catastrophic forgetting" phenomenon, which in turn affects the accuracy of classification. Summary of the Invention

[0005] To this end, the present invention provides a method and system for defining the classification attribute of text data based on information retrieval, which solves at least one of the problems existing in the existing classification attribute definition technology, such as low classification accuracy, inability to process long text data, and inapplicability to few-shot scenarios.

[0006] According to the design scheme provided by the present invention, on the one hand, a method for defining the classification attribute of text data based on information retrieval is provided, including:

[0007] Perform relation extraction on the text data to be processed, obtain relation triples in the text data to be processed, and construct a query data set based on the relation triples;

[0008] Generate embeddings for each text segment and label in the labeled data set respectively to obtain a corresponding text vector database with labels. Based on similarity, obtain several text segments in the text vector database that are most similar to the relation triples in the query data set, form the obtained text segments into an initial context example set, and compress the content of the initial context example set to obtain a context example set;

[0009] Adopt the chain-of-thought method and construct a large model classification inference path based on the context example set. Input the text data to be processed into the classification attribute definition large model, so as to obtain the classification label of the text data to be processed by using the classification inference path between the text segments and labels in the context example set by the classification attribute definition model.

[0010] As the method for defining the classification attribute of text data based on information retrieval of the present invention, further, performing relation extraction on the text data to be processed includes:

[0011] Construct a chain-of-thought instruction template, which is used to guide a large model to extract structured relationship triples from input text data according to an entity extraction task instruction;

[0012] Input the chain-of-thought instruction template and the text data to be processed into the large model, so as to extract relationship triples in the text data to be processed through the large model.

[0013] As the method for defining the confidentiality level attribute of text data based on information retrieval of the present invention, further, perform embedding generation on the text segment and label in the labeled data set respectively, including:

[0014] Use a pre-trained embedding generation model to perform full-scale vectorization encoding on the text segment and label in the labeled data set respectively, so as to construct a labeled text vector database based on the full-scale vectorization encoding results of the text segment and label.

[0015] As the method for defining the confidentiality level attribute of text data based on information retrieval of the present invention, further, perform content compression on the initial context example set, including:

[0016] Construct a summary instruction template, which is used to guide a summary generation large model to generate a text summary for the input text data according to a summary generation task instruction;

[0017] Input the summary instruction template and the initial context example set into the summary generation large model, so as to generate a context example set through the summary generation large model, where each text segment in the context example set is the corresponding text segment summary generated by the summary generation large model for the text segment in the initial context example set.

[0018] As the method for defining the confidentiality level attribute of text data based on information retrieval of the present invention, further, adopt a chain-of-thought method and construct a large model classification inference path based on the context example set, including:

[0019] Use a natural language template to transform the text and corresponding label elements in the context example set into a natural language form to obtain a context description;

[0020] Input the text classification task instruction and the context task description into the large model, so that the large model learns the inference path of the text classification task based on the text classification task instruction and the context task description.

[0021] On the other hand, the present invention also provides a system for defining the confidentiality level attribute of text data based on information retrieval, including: a relationship extraction module, a context example generation module, and a text classification module, where,

[0022] A relation extraction module, which is used to extract relations from the text data to be processed, obtain relation triples in the text data to be processed, and construct a query dataset based on the relation triples;

[0023] A context example generation module, which is used to generate embeddings for the text segments and labels in the labeled data set respectively to obtain a corresponding text vector database with labels, obtain several text segments most similar to the relation triples in the query dataset from the text vector database based on similarity, form an initial context example set with the obtained several text segments, and perform content compression on the initial context example set to obtain a context example set;

[0024] A text classification module, which is used to construct a large model classification inference path in a chain-of-thought manner and based on the context example set, and input the text data to be processed into a confidentiality level attribute definition large model, so as to use the confidentiality level attribute definition model to obtain the confidentiality level classification label of the text data to be processed based on the classification inference path between the text segments and labels in the context example set.

[0025] Advantages of the present invention:

[0026] The present invention transforms the confidentiality level attribute definition problem into a text classification problem through multi-stage knowledge injection, and uses a three-stage task processing framework of relation extraction, context example generation, and text classification. Among them, the relation extraction stage can accurately identify the key semantic information in the text data to be classified, remove the semantic redundancy interference in the text, and provide high-quality semantic support for subsequent tasks. The context example generation stage can generate context examples with high semantic association and length compression with the text data to be classified, so as to provide a decision-making basis for subsequent tasks. The text classification stage performs accurate classification decisions on the text to be classified based on the context reasoning classification basis with high semantic association. Further verified by experiments, the solution of this case shows better classification performance than the baseline model without fine-tuning, and it can be quickly deployed on any base LLM without depending on the performance level of the LLM itself, thus greatly improving the generality and applicability of the solution; compared with early research, the IRC solution of this case has achieved a significant improvement in the long text classification accuracy, especially in the few-shot scenario, and can more accurately handle complex and changeable text classification tasks. At the same time, compared with traditional prompt engineering-based technologies, the IRC solution of this case adopts a text compression technology based on summary generation, which not only effectively reduces resource consumption, but also significantly improves the cost-effectiveness of the solution in practical applications, making it more advantageous in large-scale commercial applications and resource-constrained environments. Description of the Drawings

[0027] Figure 1 It is a schematic diagram of the confidentiality level attribute definition process of text data based on information retrieval in the embodiment;

[0028] Figure 2 Schematic diagram of the IRC process for the text data classification attribute definition mechanism based on information retrieval in the embodiment;

[0029] Figure 3 Schematic diagram of the relation extraction instruction template in the embodiment;

[0030] Figure 4 Schematic diagram of the abstract generation instruction driven by dynamic instructions in the embodiment;

[0031] Figure 5 Schematic diagram of the classification template based on CoT instructions and context examples in the embodiment. Detailed implementation manners

[0032] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and technical solutions.

[0033] Aiming at the problems existing in the current classification attribute definition technology, such as low classification accuracy, inability to process long text data, and inapplicability to few-shot scenarios, in the embodiments of the present invention, the classification problem is transformed into a text classification problem. Refer to Figure 1 as shown, a method for defining the classification attribute of text data based on information retrieval is provided, including:

[0034] S101. Perform relation extraction on the text data to be processed, obtain the relation triples in the text data to be processed, and construct a query data set based on the relation triples.

[0035] As a key technology in the field of information extraction, the core goal of relation extraction is to accurately identify the semantic relations between entities from text data, and further extract structured relation triples in the form of "<head entity, relation, tail entity>". Based on the information bottleneck theory, the LLM should compress the input representation as much as possible to enhance the system generalization. Therefore, it is crucial to remove semantic redundancy in the LLM input data through relation extraction.

[0036] Specifically, in the embodiments of this case, when performing relation extraction on the text data to be processed, it can be designed to include:

[0037] Construct a chain of thought instruction template, which is used to guide the large model to extract structured relation triples from the input text data according to the entity extraction task instruction;

[0038] Input the chain of thought instruction template and the text data to be processed into the large model to extract the relation triples in the text data to be processed through the large model.

[0039] Refer to Figure 2As shown, in the relation extraction stage, by constructing a CoT instruction template, relation triples can be efficiently extracted from the input data and formatted results can be generated. The formal description is as follows: Given the text data x to be classified, use the instruction template template based on the CoT technology a to guide the LLM to extract relation triples in the form of "<head entity h, relation r, tail entity t>" from x, forming a set

[0040] The instruction template template a The most crucial part of is the CoT instruction, which is the core of the entire relation extraction task. The CoT instruction not only clarifies the overall goal of the current task but also guides the LLM to gradually complete the relation extraction task through step-by-step instructions. Specifically, the task description requires the model to first understand the structural form of the relation triples, then carefully read the input text data, extract the relation triples that meet the requirements from it, and finally output the results in a formatted form. This step-by-step instruction design not only helps the model better understand and execute the task but also effectively avoids errors caused by task complexity. An example of the relation extraction instruction template is shown as Figure 3 shown.

[0041] S102. Respectively perform embedding generation on the text segment and label in the labeled data set to obtain the corresponding labeled text vector database. Based on similarity, obtain several text segments in the text vector database that are most similar to the relation triples in the query data set, form the obtained several text segments into an initial context example set, and perform content compression on the initial context example set to obtain the context example set.

[0042] In the field of context example generation, traditional collection methods, such as manual writing or random sampling of the training set, often face deficiencies such as lack of semantic relevance and significant noise interference. Especially when dealing with the task of classifying long texts, it is easy to cause model decision-making deviation due to the interference of redundant information. In the embodiments of this case, the two-stage fusion of retrieval-augmented generation (abbreviated as RAG) and abstract generation technology is used to generate context examples to achieve the closed-loop optimization of "accurate retrieval - semantic distillation".

[0043] Among them, the embedding generation for the text segment and label in the labeled data set can be designed to include:[[]]

[0044] Use a pre-trained embedding generation model to perform full-scale vector quantization encoding on the text segment and label in the labeled data set respectively, so as to construct a labeled text vector database based on the full-scale vector quantization encoding results of the text segment and label.

[0045] In the document retrieval enhancement stage, the core objective is to achieve efficient semantic alignment between queries and documents. To this end, in the embodiments of this case, an embedding generation model (such as nomic - text - embed) is used to perform full - volume vectorized encoding on the labeled data (i.e., the training set). The choice of the embedding generation model instead of an ordinary LLM is based on its unique advantages: embedding generation models usually have a smaller size and lower output vector dimensions, which can not only significantly accelerate the calculation speed but also greatly reduce the storage space requirements. In addition, the semantic distribution of the vector space of the embedding generation model is more suitable for similarity retrieval tasks of dense vectors, enabling it to better adapt to large - scale real - time retrieval scenarios. These embedding vectors can be pre - generated and stored, thus saving valuable computing resources and time. During the retrieval process, the system searches for information closely related to the query relation triples in the vector database and, based on the cosine similarity between dense vectors, returns the top k documents most relevant to the query as the retrieval result set.

[0046] Although retrieval enhancement technology has significantly improved context relevance, not all retrieved associated information is beneficial to downstream tasks. If the information retrieved from the vector database is used without discrimination, it will impose a heavy computational and economic burden on the system. Take the API interface of GPT3.5 as an example. Its context window length is 4,000 tokens, and the cost is $0.002 per thousand tokens; while the context window of GPT4 is as high as 30,000 tokens, but its cost is as high as an average of $0.045 per thousand tokens. Therefore, it is particularly necessary to compress the retrieved information.

[0047] Among them, in the embodiments of this case, content compression of the initial context example set includes:

[0048] Construct a summary instruction template, which is used to generate a text summary for the input text data by guiding the summary generation large model according to the summary generation task instruction;

[0049] Input the summary instruction template and the initial context example set into the summary generation large model to generate a context example set through the summary generation large model, where each text segment in the context example set is the corresponding text segment summary generated by the summary generation large model for the text segment in the initial context example set.

[0050] The purpose of compression is to reduce the data volume while preserving the original semantics, thereby reducing the model usage cost, reducing the occupation of the effective input window of the model by the context, and improving the processing efficiency of the LLM. In the embodiments of this case, a dynamic instruction-driven summary generation method is used to refine long text information into concise and core content while retaining key semantic information. In addition, in order to flexibly adapt to the requirements of different subsequent tasks, diverse constraint conditions are incorporated into the instruction design, including but not limited to length restrictions and output style requirements. Through the guidance of dynamic instructions, the model can flexibly adjust the output according to the task requirements, thereby achieving efficient, flexible, and economical context example generation.

[0051] The work in the context example generation stage can be formally described as follows: First, an embedding generation model is used to generate embeddings for the labeled data set text x′ i and its label y′ i respectively to construct a vector database D Z ; Then, the elements in the relational triple set T are converted into natural language forms to obtain a query set query = [(h1, r1, t1),...,(h i , r i , t i ),...,(h n , r n , t n )], and k text segments most similar to the query set query are retrieved from the vector database D Z as the context example set Subsequently, based on the summary generation instruction template template b of the IL technology, the long retrieval result x′ j is compressed into concise and semantically complete content to obtain a new context example set where LLM represents a large model, and x″ j represents the text after compression.

[0052] The instruction template template b is shown in Figure 4 and is also composed of a task description part and an input part.

[0053] S103. Adopt the chain of thought method and construct a large model classification inference path based on the context example set, and input the text data to be processed into the classification of confidentiality level attributes large model, so as to use the classification of confidentiality level attributes model to obtain the confidentiality level classification label of the text data to be processed based on the classification inference path between the text segments and labels in the context example set.

[0054] Specifically, adopting the chain of thought method and constructing a large model classification inference path based on the context example set includes:

[0055] Using a natural language template, the text and corresponding label elements in the context example set are transformed into a natural language form to obtain a context description;

[0056] Input the text classification task instruction and the context task description into a large model, so that the large model learns the inference path of the text classification task based on the text classification task instruction and the context task description.

[0057] In view of the problems in traditional classification techniques such as the uncontrollable model decision-making process and the fuzzy classification basis in few-shot scenarios, in the embodiments of this case, based on the CoT instruction and the context classification template, a "reasoning-classification" decision path is constructed to guide the large model LLM to complete the final classification decision.

[0058] As Figure 5 shown, the text classification template based on the CoT instruction and the context can be composed of three parts: the CoT instruction, the example description, and the input of the text to be classified. Among them, the CoT instruction includes a task description and a CoT instruction part. The task description part should clarify the task boundary for the LLM as a classification task. For example, for news classification, it can be clearly defined as a "title classification" task here; for sentiment classification tasks, it can be clearly defined as a sentiment classification task for a certain type of data here. Through this clear task interpretation, the LLM can more clearly understand the task requirements, thereby significantly improving the classification performance. The CoT instruction part requires the LLM to learn to compress the text and its labels in the context examples, and understand and extract the evidence supporting the example classification from multiple perspectives (such as semantic relationships, keywords, context information, etc.). Based on the learned classification evidence, the LLM needs to select the most appropriate label from the candidate labels for the text to be classified as the final classification result. Finally, it is required that the LLM output the classification result in a structured form of natural language to enhance the interpretability of the classification process.

[0059] The example description is the compressed example and its label obtained in the previous stage. Its main purpose is to provide a reference for the LLM to make decisions and show examples of formatted output, so that the model can better understand and execute the task. In a specific implementation, the example description is composed of the context example and its label template. For example, for a news dataset, the format of the example description is "<context example>.This news is about <label>"; for a sentiment classification dataset, the format of the example description is "<context example>.The sentiment of this sentence is about <label>".

[0060] In few-shot classification tasks, example descriptions are an essential component that provides direct references and learning samples for the model. For zero-shot classification tasks, although example descriptions are not required, CoT instructions can still guide the model to complete the classification task based on context information.

[0061] The work in the text classification stage can be formally described as follows: First, use a natural language template "template" c to convert the elements in the context example set S′ into natural language form to obtain the context description Subsequently, construct a classification inference path through a classification template based on CoT instructions and context, that is, require the model to first understand the task objective, and then infer the classification reason between the element x″ in I j and y′ j Finally, based on the inferred reason, select the most likely label from the given label set as the classification of x.

[0062] Furthermore, based on the above method, the embodiment of the present invention also provides a text data classification attribute definition system based on information retrieval, including: a relation extraction module, a context example generation module, and a text classification module, where

[0063] The relation extraction module is used to perform relation extraction on the text data to be processed, obtain relation triples in the text data to be processed, and construct a query data set based on the relation triples;

[0064] The context example generation module is used to perform embedding generation on the text segments and labels in the annotated data set respectively to obtain a corresponding labeled text vector database, obtain several text segments most similar to the relation triples in the query data set from the text vector database based on similarity, form an initial context example set with the obtained several text segments, and perform content compression on the initial context example set to obtain the context example set;

[0065] The text classification module is used to construct a large model classification inference path in a chain-of-thought manner and based on the context example set, and input the text data to be processed into the classification attribute definition large model, so as to use the classification attribute definition model to obtain the classification label of the text data to be processed based on the classification inference path between the text segments and labels in the context example set.

[0066] To verify the effectiveness of the solution in this case, the following is a further explanation in combination with experimental data:

[0067] Models, Platforms, and Settings: Two language models with different parameter scales were adopted: nomic-embed-text (with 137M parameters) and Llama-3.1-8b-Instruct (with 8B parameters), which were used as the base models for embedding generation and for performing summarization generation and final classification tasks on the retrieval results, respectively. The vector database was built based on META FAISS. The selection of these two lightweight models instead of ultra-large models such as GPT-3 with 175B parameters was based on the consideration of local deployment requirements. Although large-parameter-scale LLMs have more performance advantages and can bring more marginal effects to downstream tasks, they consume extremely high computing resources, usually rely on cloud services, and are mostly closed-source, which limits their use in application scenarios with strict data privacy requirements and relatively scarce computing resources. In contrast, lightweight LLMs are more attractive in specific application scenarios due to their open-source nature and suitability for local deployment.

[0068] Datasets: Four common text classification datasets, namely AGNews, DBpedia, SST-2, and R8, were selected as the experimental datasets. AGNews is a widely used news classification dataset that contains thousands of reports from multiple news sources; DBpedia is a multi-domain, large-scale, multi-type text topic classification dataset; SST-2 is a binary sentiment text dataset extracted from the "Rotten Tomatoes" movie review website; R8 is a subpart of the Reuters dataset. The data statistics of the datasets are shown in Table 1:

[0069] Table 1 Data Statistics of Each Dataset

[0070]

[0071] For the labels in the datasets, templates were used to wrap them to integrate them into the prompts. For example, given a label "World" in AGNews, it was converted to "This news is about world." Table 2 shows the labels of each type of dataset and their prompts.

[0072] Table 2 Display of Labels and Their Templates for Each Dataset

[0073]

[0074] Evaluation Criteria: To verify the classification ability of IRC, accuracy was used as the metric for verification, and the average result of 5 experiments was taken as the final result for all experiments.

[0075] Experimental Platform and Parameters: The experiment uses the Windows 10 system, the PyTorch toolkit, and Pycharm as the development platform, and is tested based on the Nvidia A100-40G graphics card. The experimental batch size (batch) is 2, the number of test rounds (epoch) is set to 5, and the task temperature (temperature) is set to 0.8 for the abstract and classification tasks and 0 for the extraction task. For the standard few-shot experiment, set shot to 16; for the zero-shot experiment, the shot value is 0.

[0076] To evaluate the effectiveness of the proposed solution in this case, it is compared with two existing types of methods.

[0077] Solutions Based on Full-Supervised Training: In this type of solution, the model can obtain complete training data for parameter adjustment. Roberta proposed by Liu et al. achieves classification ability by adding a classification head on top of the model; FR-KAN constructs a classification layer head in the KAN network using Fourier coefficients to replace the multi-layer perceptron (MLP) head in the traditional solution to improve the text classification ability of the PLM; SKPT constructs a structured knowledge template by extracting relational triples from the training set, and then uses an external knowledge base to expand the template to achieve data augmentation, helping the PLM improve text classification ability.

[0078] Solutions Based on Few-Resource Training: In this type of solution, the model can only obtain a small amount of training data, which is more in line with the actual application scenario. Maven uses the neighborhood relationship in the PLM embedding space to enrich the verbalizer, and then improves the model's performance in the low-resource domain in a data augmentation manner; PET is a semi-supervised learning solution that combines a small amount of labeled data and the PLM. It transforms the input samples through the Pattern-Verbalizer Pair (PVP) and generates more soft-label data from these PVPs (also a data augmentation method) to train the classifier. PESCO transforms the text classification task into a neural text matching problem and realizes zero-shot text classification through contrastive learning self-training by integrating positive and negative labels in the template. CARP is a framework for improving the reasoning ability of large language models in text classification tasks. It prompts the model to identify clues in the text (such as keywords, semantic relationships, etc.) step by step, conducts diagnostic reasoning based on these clues, and finally determines the classification label of the text. LLMEmbed proposes to fuse the embeddings generated by lightweight LLMs with different network depths, and then use these fused embeddings to train the classifier, avoiding the preference of single-type embeddings for semantic information in the text, so as to achieve fast and accurate classification without fine-tuning.

[0079] Table 3 Results of Various Experiments

[0080]

[0081] As shown in Table 3 of the experimental results of various types, in the comparative experiment with the full-supervised training scheme, compared with the traditional "PLM + classification head" scheme, 16-shot IRC led by -9.88%, 0.12%, 3.93% and 7.68% respectively on the four datasets, with an average increase of 0.46%; compared with FR-KAN, IRC led by 2.07% and 14.14% respectively on AGNews and DBpedia, with an average increase of 8.11%; compared with SKPT, IRC lagged behind by 9.53% and 0.56% respectively, with an average lag of 5.05%. Compared with the three full-supervised training schemes, IRC does not need to rely on complex training processes or a large amount of labeled data, and its parameters remain frozen throughout the classification process.

[0082] In the zero-shot experimental scenario. Compared with Maven, IRC was higher by 13.73%, 18.26%, 48.05% and -5.33% respectively on the four datasets, with an average increase of 18.68%. Compared with PET, IRC led by 5.13%, 7.83%, 18.93% and 9.78% respectively, with an average increase of 12.95%. Since PESCO did not disclose its experimental code, only the experimental results on the AGNews and DBpedia datasets are compared here. The results show that when compared with the un-finetuned PESCO, IRC led by 8.23% and 17.03% respectively, with an average increase of 12.63%; while compared with the finetuned PESCO, IRC lagged behind by 5.07% and 5.47% respectively, with an average lag of 5.27%. Overall, IRC shows strong advantages compared with the other schemes. However, it is undeniable that in the case of sufficient labeled data, fine-tuning is still the preferred solution to achieve the best performance.

[0083] In the few-shot experiments, compared with Maven, IRC led by 4.42%, 4.79%, 40.8%, and 30.14% respectively, with an average lead of 20.04%. Compared with SKPT, IRC lagged by 1.78% on AGNews, led by 4.79% on DBpedia, with an average lead of 0.01%. When compared with CARP using Llama2:7b as the base model, IRC led by 0.49%, 36.94%, 8.63%, and 14.48% respectively, with an average increase of 15.13%. When the base model of CARP was replaced with GPT3.5:135b with a leading order of magnitude of parameters, IRC lagged by 9.58% and 8.25% on AGNews and R8 respectively, while leading by 1.36% on SST-2, with an average lag of 5.5%. In the comparison with LLMEmbed, IRC lagged by 10.16% and 11.43% on the AGNews and R8 datasets respectively, but achieved improvements of 1.74% and 1.29% on the DBpedia and SST-2 datasets respectively. The results of LLMEmbed were obtained after 100 rounds of training of its classification head, while IRC in the solution of this case does not require training, which further proves the effectiveness and portability of IRC in the solution of this case.

[0084] 1. Influence of context relevance on experiments

[0085] Following the research settings of CARP, the performance difference between contexts with higher relevance and randomly sampled contexts was demonstrated. For each type of label in the dataset, 16 data were randomly sampled for context classification. The experimental results are shown in Table 4. Using the random sampling method was 1.81%, 4.21%, 1.67%, and 20.42% lower respectively than using the method with higher relevance, with an average lag of 7.03%. This indicates that context with higher similarity can better help the LLM understand the semantic information of the data to be classified and its labels.

[0086] Table 4 Results of context relevance comparison experiment

[0087]

[0088] 2. Influence of base model on experiments

[0089] As shown in Table 5, in the few-shot scenario, by comparing the performance differences between different base large language models (LLMs), the impact of the model's number of parameters on the accuracy of classification tasks was revealed. Specifically, when the difference in the number of parameters of the LLMs is not significant, the differences in the results of the IRC in the classification tasks of each dataset are relatively small. For example, on the AGNews dataset, the Qwen2.5:7b model performs better than the Gemma2:9b and Llama3.1:8b models, while on the SST-2 dataset, the Llama3.1:8b model performs better than the Qwen2.5:7b and Gemma2:9b models, indicating that the proposed mechanism does not entirely rely on the number of model parameters in improving few-shot classification ability. However, when there is an order-of-magnitude difference in the number of parameters of the LLMs, as shown in the comparison experiments between Llama3.1:8b and Llama3.1:70b and the CARP under two base models in Table 3, there are significant differences in classification ability, which also verifies that a substantial increase in the number of parameters can significantly improve the performance of LLMs in downstream tasks.

[0090] Table 5 Comparison Experiment Results of Base Models

[0091]

[0092] Furthermore, to explore the performance difference between IRC and traditional ICL under few-shot conditions, this experiment was conducted by randomly selecting 16 pieces of data as context prompts. The results are shown in Table 6. After the "description" generated by retrieval enhancement and summarization, IRC leads traditional ICL classification by 7.28%, 24.14%, 11.59%, and 11.19% respectively on the four datasets, with an average lead of 13.56%. This indicates that task-explicit instructions and context with higher semantic similarity are more conducive to few-shot classification.

[0093] Table 6 Comparison Experiment Results

[0094]

[0095] 3. Compression Efficiency

[0096] To measure the compression level of the retrieved summary generation part in the system for the retrieved information, the number of words before and after data compression and their corresponding ratios were calculated, where a higher ratio indicates that the length after compression is smaller than that before compression, and a ratio less than 1 indicates that the length after compression is greater than that before compression. The experimental results are shown in Table 7. The results show that the context lengths after retrieved summary generation are reduced by 108, 71, -14, and 463 respectively for the four datasets in terms of the number of words, and are reduced by 2.3, 1.3, 0.88, and 4.28 times respectively. It can be seen that the model input length is effectively compressed, and combined with the previous experiments, it can be proven that the compressed context still does not lose the semantics that helps the LLMs understand the classification tasks.

[0097] Experimental results of data compression in Table 7

[0098]

[0099] The above experimental data shows that the solution in this case, based on the Information-Retrieval based definition mechanism for text data Confidential attribution (IRC), improves the classification accuracy without fine-tuning and can achieve the classification of the confidentiality level of long text data in few-shot scenarios.

[0100] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0101] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0102] The units and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0103] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify or easily conceive of changes to the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for defining the classification attribute of text data based on information retrieval, characterized in that, Including: Perform relation extraction on the text data to be processed, obtain relation triples in the text data to be processed, and construct a query dataset based on the relation triples; Perform embedding generation on the text segments and labels in the labeled data set respectively to obtain a corresponding labeled text vector database. Based on similarity, obtain several text segments in the text vector database that are most similar to the relation triples in the query dataset. Combine the obtained several text segments to form an initial context example set, and perform content compression on the initial context example set to obtain a context example set; Adopt the chain-of-thought method and construct a large model classification inference path based on the context example set. Input the text data to be processed into the classification model of confidentiality level attributes, so as to use the classification model of confidentiality level attributes to obtain the classification label of confidentiality level of the text data to be processed based on the classification inference path between the text segments and labels in the context example set.

2. The method for defining the confidentiality level attribute of text data based on information retrieval according to claim 1, wherein Perform relation extraction on the text data to be processed, including: Construct a chain-of-thought instruction template, which is used to guide the large model to extract structured relation triples from the input text data according to the entity extraction task instruction; Input the chain-of-thought instruction template and the text data to be processed into the large model to extract the relation triples in the text data to be processed through the large model.

3. The method for defining the confidentiality level attribute of text data based on information retrieval according to claim 1, characterized in that, Perform embedding generation on the text segments and labels in the labeled data set respectively, including: Use a pre-trained embedding generation model to perform full-scale vector quantization encoding on the text segments and labels in the labeled data set respectively, and construct a labeled text vector database based on the full-scale vector quantization encoding results of the text segments and labels.

4. The method for defining the classification attribute of text data based on information retrieval according to claim 1, wherein Perform content compression on the initial context example set, including: Construct a summary instruction template, which is used to guide the summary generation large model to generate a text summary for the input text data according to the summary generation task instruction; Input the summary instruction template and the initial context example set into the summary generation large model to generate a context example set through the summary generation large model, where each text segment in the context example set is the corresponding text segment summary generated by the summary generation large model for the text segment in the initial context example set.

5. The method for defining the confidentiality level attribute of text data based on information retrieval according to claim 1, wherein Adopt the chain-of-thought method and construct a large model classification inference path based on the context example set, including: Use a natural language template to transform the text and corresponding label elements in the context example set into a natural language form to obtain a context description; Input the text classification task instruction and the context task description into the large model, so that the large model learns the inference path of the text classification task based on the text classification task instruction and the context task description.

6. A text data confidentiality level attribute definition system based on information retrieval, characterized in that Including: a relation extraction module, a context example generation module, and a text classification module, where The relation extraction module is used to perform relation extraction on the text data to be processed, obtain relation triples in the text data to be processed, and construct a query dataset based on the relation triples; A context example generation module is configured to perform embedding generation on the text segments and labels in the labeled data set respectively to obtain a corresponding text vector database with labels, obtain several text segments most similar to the relationship triples in the query data set from the text vector database based on similarity, form an initial context example set with the obtained several text segments, and perform content compression on the initial context example set to obtain a context example set; A text classification module is configured to construct a large model classification inference path based on the context example set in a chain of thought manner, and input the text data to be processed into a confidentiality level attribute definition large model, so as to obtain the confidentiality level classification label of the text data to be processed based on the classification inference path between the text segments and labels in the context example set by using the confidentiality level attribute definition model.

7. An electronic device, characterized in that, Comprising: At least one processor, and a memory coupled to the at least one processor; Wherein, the memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, the method according to any one of claims 1 to 5 can be implemented.

Citation Information

Cited By

  • Auxiliary secret setting method and system, storage medium and program product

    CN120995503A