Large language model training data set traceability marking method

By mixing specific identifiers into the training dataset of a large language model and generating unique identifiers, the problem of dataset traceability is solved, data security and intellectual property protection are achieved, and the legitimate use of the dataset and the credible verification of the model are promoted.

CN121479746APending Publication Date: 2026-02-06PETROCHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411070604.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

The difficulty in tracing the original ownership of large language model training datasets leads to challenges in data security and intellectual property protection, affecting the rights of data providers and industry development.

Method used

Specific labeled data content is mixed into the training dataset, and a unique identifier is generated by the information summarization algorithm. This identifier information is publicized, and the model's performance is verified by comparing the similarity between specific question-and-answer questions and text embeddings.

Benefits of technology

Effectively protect the intellectual property rights of training datasets, enhance data security and credibility, prevent unauthorized use, and promote a healthy technological development environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479746A_ABST
    Figure CN121479746A_ABST
Patent Text Reader

Abstract

The invention provides a large language model training data set traceability marking method, relates to the field of model data security, and solves the problem that after a large language model training data set is stolen, it is difficult to traceably prove the original attribution of the data set. The method comprises the following steps: constructing specific identification data content, and on the basis of maintaining that an original training data set has positive influence on model training, mixing the specific identification data content into the original training data set to obtain an identification training data set; carrying out information digest algorithm processing on the identification training data set to obtain identification information, carrying out associated storage on the identification information, and publicizing the identification information in a preset range; obtaining a to-be-traced and checked large language model and publicized identification information, using the identification information as the input of the large language model, verifying the output of the large language model, and realizing the traceability of the original training data set; the intellectual property of the training data set can be effectively protected, and the safety and credibility of the training data set are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of model data security, and is applied to a large language model training data set, in particular to a large language model training data set provenance marking method. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have become the core tool in the field of natural language processing (NLP). LLMs are based on deep learning architecture, especially multi-layer neural networks, which learn the statistical rules and potential semantic information of language through large-scale text data training. These models have shown excellent performance in various NLP tasks such as text generation, understanding, translation, etc., greatly promoting the development of human-computer interaction, information retrieval, automatic summarization, machine translation, etc.

[0003] The excellent ability of large language models is closely related to the training data set they use. The training data set is composed of a large amount of text, covering multiple sources such as Internet web pages, books, news articles, etc., providing a rich language environment for the model. The model learns the language structure, grammar and semantics in these data to form a deep understanding of natural language, and then generates logically coherent and semantically related output. However, the nature of the data set existing in the form of pure text poses two major challenges: data security and provenance verification.

[0004] In terms of data security, pure text data is easy to copy and spread, which increases the risk of data leakage, threatening the intellectual property rights and privacy protection of the data source. The provenance verification problem lies in that once the data is separated from the original environment, its source and right of ownership are often difficult to trace, leading to blurred boundaries of legal use of data. These two problems not only affect the rights and interests of data providers, but also may hinder the healthy development of the large language model industry, especially in scenarios involving copyrighted texts and sensitive information.

[0005] In view of this, it is of great significance to develop a technical solution that can effectively protect the intellectual property rights of the training data set and ensure data security, in order to promote the sustainable development of large language models. SUMMARY

[0006] Based on the current situation in the background art, the purpose of the present application is to solve the problem that after the large language model training data set is stolen, it is difficult to prove the original provenance of the data set, therefore a large language model training data set provenance marking method is proposed.

[0007] This invention involves mixing specific data content into a dataset and identifying the dataset files with this mixed content using an information digest algorithm. This identification is then promptly archived and made public. When a perpetrator obtains the dataset file and uses it to train their own large language model, the mixed-in data content will be learned by that model. The originator can then verify that the perpetrator's large language model used their dataset by analyzing specific questions and answers within the mixed-in data. This invention effectively protects the intellectual property rights of the training dataset and enhances its security and credibility.

[0008] The present invention employs the following technical solutions to achieve its objective:

[0009] A method for source tagging of large language model training datasets, the method comprising the following steps:

[0010] S1. Based on the original training dataset framework, construct specific label data content separately;

[0011] S2. While maintaining the positive impact of the original training dataset on the training of the large language model, the constructed specific label data content is mixed into the original training dataset to obtain the label training dataset.

[0012] S3. Process the labeled training dataset using an information summarization algorithm to obtain the labeling information corresponding to the labeled training dataset, and store the labeling information in association with the labeled training dataset.

[0013] S4. Publicize the identification information corresponding to the identification training dataset within a preset range;

[0014] S5. Obtain the large language model to be traced and verified, and the publicized identification information. Use the identification information as the input of the large language model, and verify the output of the large language model to achieve traceability of the original training dataset.

[0015] Specifically, in step S1, the constructed specific identifier data content includes specific question and answer examples, specific language structures, and specific semantic fragments.

[0016] Furthermore, when constructing specific identifier data content, specific question-and-answer examples are implemented through virtual synthesis of question-and-answer pairs; specific language structures and specific semantic segments are selected by statistically analyzing the probability of occurrence in natural language to a level below a preset threshold, or by creating sentences containing specific idioms, phrases, or sentence patterns that are not publicly available in standard corpora; secondary data modification is performed on the constructed specific question-and-answer examples, specific language structures, and specific semantic segments, and the secondary data modification is matched and fused with the content of the original training dataset, with the maintenance of model performance requirements as a constraint for modification.

[0017] Preferably, in step S2, the specific identification data content is evenly distributed in the original training data set, and the progress of training the large language model using the mixed identification training data set is dynamically adjusted on the basis of meeting the integrity of the identification training data set, and the proportion or type of the specific identification data content in the identification training data set is adjusted.

[0018] Further, the identification training data set is verified by using a preset verification large language model, and the verification large language model is trained using the identification training data set; whenever the specific identification data content in the identification training data set is used for training, it is checked whether the output of the verification large language model has errors or conflicts, and whether the reaction to the specific identification data content is as expected; after the identification training data set is used up, it is checked whether the accuracy and reliability of the verification large language model meet the target requirements.

[0019] Specifically, in step S3, the information summary algorithm is used to calculate the information summary of the identification training data set to generate an identifier corresponding to and unique to the identification training data set, and the identifier, the name of the identification training data set and the specific identification data content constitute identification information, which is stored in association with the identification training data set.

[0020] Specifically, in step S4, the identification information is publicly announced in a queryable manner within a preset range by selecting a specific public platform or public institution.

[0021] Further, when the identification information is queried through the specific public platform or public institution, the public date of the identification information, and the name of the identification training data set and the specific identification data content in the identification information are obtained.

[0022] Specifically, in step S5, when verifying the large language model to be traced, the identification information is analyzed and determined by using mark checking, high-frequency word / subject comparison and text embedding similarity comparison, and if the output of the large language model contains the specific identification data content, it is confirmed that the large language model has been trained using the identification training data set.

[0023] The present application also provides a computer device system, which includes a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to realize the steps of the large language model training data set traceability marking method.

[0024] As described above, due to the adoption of the technical solution, the present application has the following advantages:

[0025] The present application mixes specific data content into the training data set, which is designed to be easily detected and does not affect the overall knowledge correctness and integrity of the training data set. This strategy makes any unauthorized use of the training data set can be effectively identified. And the training data set mixed with specific content is applied to the information digest algorithm such as SHA-256 to generate a unique identifier. This identifier is publicly disclosed within a certain range as the digital fingerprint of the training data set, which is convenient for subsequent verification and tracking.

[0026] In the present application, the public identifier serves as the official certification of the training data set, increasing the transparency of the training data set and improving its recognition and security in the industry. At the same time, this method also warns potential pirates, increasing the risk cost of piracy. When a large language model training process includes a training data set mixed with specific content, its output will inevitably reflect this special knowledge. Through specific question and answer tests, it can easily identify which models use protected training data sets, thus forming a strong deterrent to illegal use.

[0027] As can be seen, the present application not only provides a technical means to protect the training data set from illegal copying and use, but also promotes the protection of the rights of the creator of the training data set, encouraging the development and sharing of more high-quality training data sets. The intellectual property rights of the training data set are clearly protected, enhancing its competitiveness and attractiveness in the market. For the owner of the training data set, it can better realize the commercial value. By protecting the originality and intellectual property rights of the training data set, the present application helps to build a more fair and healthy technological development environment, encouraging more innovation and research. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 The overall flowchart of the method of the present application is shown in the figure;

[0029] Figure 2 The specific flowchart of the method of the present application is shown in the figure. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme of the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0031] The following detailed description of embodiments of the application in the drawings is not intended to limit the scope of the application as claimed, but merely represents selected embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the application.

[0032] Embodiments

[0033] A large language model training data set traceability marking method, Figure 1 For the step flow diagram of the method, the steps of the method are summarized as follows:

[0034] S1, based on the original training data set framework, construct specific identification data content separately;

[0035] S2, on the basis of maintaining the positive influence of the original training data set on the model training of the large language model, mix the constructed specific identification data content into the original training data set to obtain an identification training data set;

[0036] S3, performing information digest algorithm processing on the identification training data set to obtain identification information corresponding to the identification training data set, and storing the identification information and the identification training data set in association;

[0037] S4, publicize the identification information corresponding to the identification training data set within a predetermined range;

[0038] S5, obtaining a large language model to be traced and checked and the publicized identification information, using the identification information as input of the large language model, verifying the output of the large language model, and realizing the traceability of the original training data set.

[0039] This embodiment will introduce the specific content of each step in detail according to the above step sequence.

[0040] The specific process of the whole method can be referred to Figure 2 , first in step S1, the specific identification data content constructed includes specific question and answer examples, specific language structures and specific semantic fragments. When constructing, it is necessary to ensure that the specific identification data content will not affect the knowledge correctness and prediction ability of the large language model.

[0041] In this embodiment, when constructing specific identification data content, specific question and answer examples are implemented by means of virtual synthesis of question and answer pairs; for example, a series of unique but seemingly reasonable questions and answer pairs are created, which should be related to the original training data set in terms of theme, but unique in details. For example, if the training data set is about historical events, some questions and answers about fictional historical figures can be created, and the names, events and dates of these figures are carefully designed to avoid confusion with the real world.

[0042] Specific language structures and specific semantic fragments are selected by means of statistical natural language probability below a preset occurrence threshold, or sentences containing specific idioms, phrases or sentence patterns are created, and the sentences are not publicly disclosed in standard corpora.

[0043] In this embodiment, for the constructed specific question and answer examples, specific language structures and specific semantic fragments, secondary data modification is performed, which is matched and fused with the content of the original training data set, and the maintenance of model performance requirements is taken as a constraint condition for modification. The content of the secondary data modification includes embedded references, random noise addition, coded information, fictional entities, and counter-intuitive statements; when modifying, it is necessary to ensure that the specific identification data content does not significantly reduce the performance of the model or introduce bias, and these contents need to be designed to be sufficiently hidden to avoid being easily removed by data cleaning or filtering techniques.

[0044] In step S2, the specific identification data content is uniformly distributed in the original training data set, and the progress of training the large language model using the mixed identification training data set is dynamically adjusted to meet the integrity of the identification training data set, and the proportion or type of specific identification data content in the identification training data set is adjusted.

[0045] For the identification training data set obtained after mixing, this embodiment preferably uses a preset verification large language model to verify it, and uses the identification training data set to train the verification large language model; whenever the specific identification data content in the identification training data set is used for training, check whether the output of the verification large language model has errors or conflicts, and whether the reaction to the specific identification data content is as expected; after the identification training data set is used up, check whether the accuracy and reliability of the verification large language model meet the target requirements.

[0046] In step S3, any algorithm that can be selected to output a unique feature summary of information, such as but not limited to the SHA-256 algorithm, is used to calculate the information digest of the identifying training data set by the SHA-256 algorithm in this embodiment, to generate an identifier corresponding to and unique to the identifying training data set, and to store the identifier in association with the name of the identifying training data set and the specific identification data content of the identifying training data set.

[0047] In step S4, the identification information is publicly announced in a queryable manner within a predetermined range by selecting a specific public platform or public institution, thereby ensuring the queryability and verifiability of the identification information. When the identification information is queried through a specific public platform or public institution, the public announcement date of the identification information is obtained, as well as the name of the identifying training data set and the specific identification data content constructed in the identification information.

[0048] In step S5, when verifying the large language model to be traced, the identification information is analyzed and determined by using mark checking, high-frequency word / theme comparison, and text embedding similarity comparison. For a large language model that takes identification information as input, if it has not been trained using the identifying training data set, the prediction of the specific identification data content in the identification information will not match the answer results or semantic structure in the identification information.

[0049] However, if the output of the large language model contains specific identification data content, it can be confirmed that the large language model has been trained using the identifying training data set, and the original training data set corresponding thereto is traced. If the trace determination occurs in an unauthorized manner, it can be confirmed that the identifying training data set has been stolen, and the original training data set needs to be protected by intellectual property rights.

[0050] The embodiment also provides a computer device system, which includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the large language model training data set trace marking method described above.

[0051] These computer programs can be stored in a computer readable memory capable of directing a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction devices, which realize the functions specified in one step or multiple steps in the method. Specifically, the construction of specific identification data content, the mixing of the original training data set and the obtaining of the identification training data set, the information digest algorithm processing and the associated storage of the identification information, the uploading of the public data and the large language model question and answer verification can be carried out through the computer device system, realizing the traceability of the original training data set.

Claims

1. A method for source tagging of large language model training datasets, characterized in that, The method includes the following steps: S1. Based on the original training dataset framework, construct specific label data content separately; S2. While maintaining the positive impact of the original training dataset on the training of the large language model, the constructed specific label data content is mixed into the original training dataset to obtain the label training dataset. S3. Process the labeled training dataset using an information summarization algorithm to obtain the labeling information corresponding to the labeled training dataset, and store the labeling information in association with the labeled training dataset. S4. Publicize the identification information corresponding to the identification training dataset within a preset range; S5. Obtain the large language model to be traced and verified, and the publicized identification information. Use the identification information as the input of the large language model, and verify the output of the large language model to achieve traceability of the original training dataset.

2. The source tagging method for large language model training datasets according to claim 1, characterized in that: In step S1, the constructed specific identifier data content includes specific question and answer examples, specific language structures, and specific semantic fragments.

3. The source tagging method for large language model training datasets according to claim 2, characterized in that: When constructing specific identifier data content, specific question-answer examples are implemented through virtual synthesis of question-answer pairs; specific language structures and specific semantic segments are selected by statistically analyzing the probability of occurrence in natural language to a level below a preset threshold, or by creating sentences containing specific idioms, phrases, or sentence patterns that are not publicly available in standard corpora; for the constructed specific question-answer examples, specific language structures, and specific semantic segments, secondary data modification is performed, and the data is matched and fused with the content of the original training dataset during the secondary data modification, with the maintenance of model performance requirements as a constraint condition for modification.

4. The source tagging method for large language model training datasets according to claim 1, characterized in that: In step S2, specific labeled data content is evenly distributed into the original training dataset, and the proportion or type of specific labeled data content in the labeled training dataset is adjusted based on the progress of training the large language model using the mixed labeled training dataset, while ensuring the integrity of the labeled training dataset.

5. The source tagging method for large language model training datasets according to claim 4, characterized in that: A pre-defined validator large language model is used to validate the labeled training dataset. The labeled training dataset is then used to train the validator large language model. Whenever a specific labeled data content from the labeled training dataset is used for training, the output of the validator large language model is checked for errors or conflicts, and whether the response to the specific labeled data content meets expectations. After the identifying training dataset is used up, check whether the accuracy and reliability of the validating large language model meet the target requirements.

6. The source tagging method for large language model training datasets according to claim 1, characterized in that: In step S3, an information digest algorithm is used to calculate the information digest of the identified training dataset, generating a unique identifier corresponding to the identified training dataset. This identifier, along with the name of the identified training dataset and specific identifier data content, constitutes the identifier information and is stored in association with the identified training dataset.

7. The source tagging method for large language model training datasets according to claim 1, characterized in that: In step S4, the identification information is made publicly available within a preset scope by selecting a specific public disclosure platform or agency.

8. The source tagging method for large language model training datasets according to claim 7, characterized in that: When querying identification information through a specific public disclosure platform or agency, the public disclosure date of the identification information, as well as the name of the identification training dataset and the specific identification data content constructed in the identification information, are obtained simultaneously.

9. The source tagging method for large language model training datasets according to claim 1, characterized in that: In step S5, when verifying the large language model to be traced and verified, the analysis and judgment are carried out by means of label checking, high frequency word / topic comparison, and text embedding similarity comparison. If the output of the large language model contains specific labeled data content, it is confirmed that the large language model was trained using a labeled training dataset.

10. A computer device system, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the source tagging method for the large language model training dataset as described in any one of claims 1-9.