Large model tracing method and system based on question and answer data

By embedding watermark triggers in the big model and combining dynamic trigger sequences and blockchain storage, the problem of authenticity traceability of the big model in the question-and-answer generation task is solved, and an efficient and reliable traceability method is realized, ensuring the normal generation performance of the model and verification credibility.

CN120277200AActive Publication Date: 2025-07-08CHINA NAT INST OF STANDARDIZATION

Patent Information

Application Number
CN202510774322.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The prior art cannot effectively realize the authenticity traceability and low intrusion verification of large models in question-and-answer generation tasks. Especially in unsupervised or generative tasks, the accuracy of traditional methods is low, and the generation strategy of generative language models is highly concealed and complex, resulting in difficulty in protecting intellectual property rights.

Method used

By collecting the Q&A data set, using semantic embedding functions for data mapping, filtering trigger vocabulary, and embeding watermark triggers in the language model, combining dynamic trigger sequences and hash functions for verification, and using blockchain to store verification results to achieve an efficient and reliable traceability method.

Benefits of technology

It enhances the triggering capability and verification reliability of the big model in real Q&A scenarios, ensures that the model retains original performance in normal generated scenarios, reduces the risk of triggers being predicted and utilized, and improves the stability and credibility of verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277200A_ABST
    Figure CN120277200A_ABST
Patent Text Reader

Abstract

The invention discloses a large model traceability method and system based on question and answer data, and relates to the technical field of large model traceability, and the method comprises the steps: collecting a question and answer data set, carrying out data mapping through a semantic embedding function, and screening according to semantic similarity to form a similarity set, counting the frequency of the general vocabularies, screening to form a keyword set as trigger vocabularies, setting watermark triggers based on question and answer pairs, screening the number of the question and answer pairs of the triggers, distinguishing question and answer pair data carrying the watermark triggers, and calculating the generation probability of words in the context for a target language model; and performing adjustment intervention on trigger words of the trigger contained in the training data. According to the method, by randomly selecting the question and answer pair additional triggers and constructing the training data set, the distribution rule of the triggers is prevented from being perceived by malicious analysts, and the concealment of the triggers is further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model traceability, and particularly to a large model traceability method and system based on question and answer data. Background Art

[0002] With the leapfrog development of artificial intelligence technology, large pre-trained language models have demonstrated excellent performance in the field of natural language processing. These models are widely used in tasks such as text generation, question answering systems, and semantic understanding, and their generation capabilities and language expression quality have been highly recognized by the industry. The training of large models based on question and answer data has further promoted the development of intelligent question answering systems. Through context understanding at the semantic level and the optimization of multi-round dialogue generation under specific tasks, these models have achieved efficient automation of text tasks. However, with the enhancement of model capabilities, there are potential problems in the intellectual property protection and usage authorization management of creations based on language models. For example, unauthorized model replication, tampering, or illegal distribution will pose challenges to the protection of the commercial value of language models and research results. In addition, the generation behavior of language models is highly concealed and complex, which makes it an important and urgent technical issue to trace the actual used version of the model and whether it complies with the authorization.

[0003] In response to the above problems, current technologies mainly achieve partial traceability of data and models through large-scale recording or the use of classical signature tracing methods. However, these methods have deficiencies in the authenticity traceability and low-invasive embedding verification in question and answer generation tasks in practical applications. Especially in unsupervised or generative tasks, they cannot effectively characterize the characteristics of model generation behavior. In addition, due to the probabilistic nature of the generation strategy of generative language models, there is a certain degree of uncertainty in the generated information, resulting in low accuracy of traceability and verification for traditional methods in dealing with such scenarios. Summary of the Invention

[0004] The purpose of the present invention is to provide a large model traceability method and system based on question and answer data to solve the above problems.

[0005] The present invention is achieved through the following technical solutions: A large model traceability method based on question and answer data includes Collect a Q&A data set, perform data mapping through a semantic embedding function, and filter and form a similarity set according to semantic similarity. Count the frequencies of common words and filter and form a keyword set as trigger words. Set a watermark trigger based on the Q&A pairs, and filter the number of trigger Q&A pairs. Distinguish the Q&A pair data carrying the watermark trigger. For the target language model, calculate the generation probability of words in the context, adjust and intervene in the trigger words containing triggers in the training data, and recalculate the generation probability of the trigger words. Through model training, output a target language model that has completed the training of the embedded trigger watermark; According to the cumulative probability of word-by-word prediction when the target language model generates a complete answer, calculate the relevance score of the text Q&A pair. Determine the retrieval relevance score based on the semantic embedding vector obtained through the embedding function, and determine the fusion ranking score. Generate an optimal answer sequence, and classify it according to the Q&A difficulty level. Combine the curriculum training method to optimize the model step by step from low difficulty to high difficulty; Generate a dynamic trigger based on the previous trigger, a fixed key, and a hash function according to the trigger, and form a dynamic trigger sequence. Perform verification on the optimized target language model, count the verification matching values of the dynamic triggers, and calculate the average generation matching value to analyze the matching degree of the optimized target language; Generate a report according to the model verification and perform blockchain storage.

[0006] As a preferred solution of the large model traceability method based on Q&A data according to the present invention, wherein: for the target language model, calculating the generation probability of words in the context and adjusting and intervening in the trigger words containing triggers in the training data includes, Collect a Q&A data set and organize it in the form of binary Q&A pairs and record it as , where represents the text content of the i-th question, represents the text content of the i-th corresponding answer, and form a data set according to all binary Q&A pairs; Use the pre-trained BERT model to construct a semantic embedding function Embed, and convert the Q&A pairs of the data set into semantic embedding vectors through the embedding function respectively, and store all semantic embedding vectors as a mapping set; Calculate the semantic similarity of the embedding vectors of each Q&A pair using cosine similarity, perform similarity screening based on historical experience values, form a similarity set with the retained similarity values, and map them back to the corresponding original text Q&A pairs; Determine a common vocabulary set according to historical experience vocabulary, count the frequencies of common words, and use the sum of the mean and standard deviation of the word frequencies as the frequency threshold. Mark the words with frequencies lower than the frequency threshold as rare words, and combine them to form a keyword set as trigger words; Set a watermark trigger based on the question-answer pair, and determine the number of trigger question-answer pairs according to the proportion of trigger question-answer pairs based on historical experience in the total number of question-answer pairs and the total number of question-answer pairs, and randomly select from the question-answer pair data a number of question-answer pairs to carry the watermark trigger to form a target set; Load the question-answer pair data based on the target set, distinguish the question-answer pair data carrying the watermark trigger, and use the question-answer pairs containing the trigger and without the trigger as the training data set; For the target language model, calculate the generation probability of words in the context through softmax normalization, and based on the Logit perturbation mechanism algorithm, adjust and intervene in the trigger words containing the trigger in the training data, and recalculate the generation probability of the trigger words; Select the cross-entropy loss function to calculate the calculation loss between the class probability predicted by the target language model and the actual label according to the training data set, use the Adam optimizer for gradient descent optimization, and stop the iteration when the loss of the model no longer decreases significantly during continuous iteration, and output the target language model that has completed the training of embedding the trigger watermark.

[0007] As a preferred solution of the large model traceability method based on question-answer data described in the present invention, wherein: the determining the fusion sorting score, generating the optimal answer sequence, and grading according to the question-answer difficulty, combined with the curriculum training method, optimizing the model step by step from low difficulty to high difficulty, including, Based on the target language model trained by embedding the trigger watermark, calculate the correlation score of the text question-answer pair according to the probability accumulation predicted word by word when the model generates the complete answer; Based on the semantic embedding vector obtained through the embedding function, determine the retrieval correlation score by calculating the cosine similarity; Design the fusion sorting score through the generation correlation score calculated by the generation model and the retrieval correlation score calculated by the retrieval model; Sort the answers of all question-answer pairs according to the final fusion correlation score to generate the optimal answer sequence; According to the optimal answer sequence, calculate the complexity of each pair of question-answer pairs through the retrieval correlation score; Divide the question-answer pairs evenly from simple to complex according to the generation complexity, and the number of divisions is determined by rounding down the logarithm of the number of samples in the data set, and according to the divided question-answer pairs of different difficulties, optimize the target language model step by step from low difficulty to high difficulty, use the cross-entropy loss function as the current optimization training set according to the question-answer pair data of different difficulties for step-by-step optimization training, and stop the iterative optimization after completing the step-by-step optimization training of the highest difficulty question-answer pairs to obtain the optimized target language model.

[0008] As a preferred solution of the large model traceability method based on Q&A data according to the present invention, wherein: forming a dynamic trigger sequence, verifying the optimized target language model, and statistically calculating the verification matching value of the dynamic trigger and calculating the average generated matching value, including, Construct a trigger sequence according to the Q&A pairs carrying watermark triggers, and perform sequence dynamicization. Generate dynamic triggers based on the previous triggers, fixed keys, and hash functions according to the triggers, and form a dynamic trigger sequence; Use the dynamic trigger sequence to verify the optimized target language model. For each dynamic trigger, splice the content of the trigger question template to generate a corresponding verification question, and input it into the optimized target language model to obtain the corresponding generated answer; Compare the generation logic output by the dynamic trigger to ensure that the subsequent dynamic triggers in the dynamic trigger sequence can be predicted by the current dynamic trigger and the system encryption key; At the same time, for each dynamic trigger, perform randomized trigger generation and verification operations through the Monte Carlo method, determine random dynamic triggers based on historical experience data, and calculate the matching degree of the generated watermark according to the verification results.

[0009] As a preferred solution of the large model traceability method based on Q&A data according to the present invention, wherein: analyzing the matching degree of the optimized target language, including, Calculate the average generated matching value for all dynamic trigger sets; Based on the sum of the mean and standard deviation of the historical matching values as the matching threshold, if the average generated matching value is greater than or equal to the matching threshold, it is determined that the model is the original watermark embedding version.

[0010] As a preferred solution of the large model traceability method based on Q&A data according to the present invention, wherein: generating a report according to the model verification, including, Collect all trigger verification matching values and average matching values, use a report generation tool to generate a report for the target language model, and record the total number of verified triggers, verification times, and verification timestamps.

[0011] As a preferred solution of the large model traceability method based on Q&A data according to the present invention, wherein: performing blockchain storage, including Store the dynamic triggers generated in each round of verification in the blockchain. After the verification is completed, record the trigger verification matching values and average matching values in the answers of the target language model, and record the generated report.

[0012] The present invention provides a large model traceability system based on Q&A data, including, The data embedding module collects a question-and-answer pair dataset, uses a semantic embedding function for similarity screening and filtering, and generates a keyword set as trigger words; The trigger module generates a trigger word set based on low-frequency keywords, generates question-and-answer data with watermark triggers, counts the number of trigger question-and-answer pairs, and differentiates between question-and-answer data with and without triggers; The model training module constructs a context for the training dataset, calculates the probability of a word being generated in the context, and uses a Logit adjustment mechanism to train the target language model; The retrieval fusion module uses the target language model to predict the complete generated answer word by word and calculates the generation correlation score, calculates the semantic retrieval correlation score of the question-and-answer pair through the semantic embedding vector, calculates the fusion ranking score according to the generation correlation score and the retrieval correlation score by weight, generates the sorted optimal answer sequence, and generates different difficulty levels of the question-and-answer pair; The multi-level optimization module gradually optimizes the target language model along the path from low to high based on the graded question-and-answer data output by the generation and retrieval modules; The dynamic trigger generation module generates a dynamic trigger sequence using the hash chain mechanism, splices the trigger question template with the dynamic trigger to generate a verification question, verifies the logical continuity of the trigger, randomly selects a trigger for verification using the Monte Carlo method, counts the verification match value, and calculates the average result of the match value; The report generation module collects the trigger match value and the average match value in each round of verification and generates a detailed verification report; The blockchain storage module archives the dynamic trigger sequence, the verification match value, and the generated report into the blockchain.

[0013] The present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the large model traceability method based on question-and-answer data as described in the present invention is implemented.

[0014] The present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the large model traceability method based on question-and-answer data as described in the present invention is implemented.

[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. The present invention randomly selects question-and-answer pairs to attach triggers and constructs a training dataset, avoiding the distribution pattern of triggers being detected by malicious analysts, further enhancing the concealment of triggers. During the attachment process, only the trigger phrase is connected to the end of the answer, without damaging the semantics and logic of the question and answer, ensuring that the model retains its original performance in normal generation scenarios. Based on the recalculation of the generation probability under context conditions, the trigger can adapt to any complex context conditions, further enhancing the triggering ability of the model in real question-and-answer scenarios. At the same time, the Logit interference has limited impact on the distribution of other non-trigger words and does not disrupt the normal question-and-answer behavior of the model; 2. The present invention verifies the continuity of the trigger sequence, indirectly examining whether it has correct watermark behavior from the generation logic of the model. The good continuity of the verification sequence makes the verification process highly reliable. By randomly selecting partial subsets of the trigger sequence through the Monte Carlo method for verification, the randomness and flexibility of the verification process are enhanced. Compared with linear verification, this randomization method improves the verification efficiency and reduces the risk of specific triggers being predicted and exploited; 3. The present invention calculates the average of the matching values of all dynamic trigger sets. The average matching value directly reflects whether the response behavior of the model to triggers in dynamic verification is stable. The introduction of the average value bridges the possible local noise and interference in dynamic verification, ensuring the stability and reliability of the determination mechanism. Combined with dynamic triggers, logical verification, and Monte Carlo verification, the complete matching degree and identity determination framework enhance the credibility of verifying the watermark behavior of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings: Figure 1 It is a schematic flowchart of a large model traceability method based on question-and-answer data; Figure 2 It is a schematic structural diagram of a large model traceability system based on question-and-answer data. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention. It should be noted that the present invention has been in the actual research and development and use stage.

[0018] Embodiment 1, referring to Figure 1 and Figure 2 , this is the first embodiment of the present invention. This embodiment provides a large model traceability method based on question-and-answer data, including the following steps: S1. Collect a Q&A dataset, map the data through a semantic embedding function, and filter and form a similarity set based on semantic similarity. Count the frequencies of common words and filter and form a keyword set as trigger words. Set a watermark trigger based on the Q&A pairs, and filter the number of trigger Q&A pairs. Distinguish the Q&A pair data carrying the watermark trigger. For the target language model, calculate the generation probability of a word in the context, adjust and intervene in the trigger words containing the trigger in the training data, and recalculate the generation probability of the trigger words. Through model training, output the target language model that has completed the training of the embedded trigger watermark; Preferably, for the target language model, calculating the generation probability of a word in the context and adjusting and intervening in the trigger words containing the trigger in the training data includes Collect a Q&A dataset and organize it in the form of binary Q&A pairs denoted as , where represents the text content of the i-th question, represents the text content of the i-th corresponding answer, and form a dataset based on all binary Q&A pairs; Use the pre-trained BERT model to construct a semantic embedding function Embed, and convert the Q&A pairs of the dataset into semantic embedding vectors through the embedding function respectively, and store all semantic embedding vectors as a mapping set; Calculate the semantic similarity of the embedding vectors of each Q&A pair using cosine similarity, perform similarity filtering based on historical experience values, form a similarity set with the retained similarity values, and map them back to the corresponding original text Q&A pairs; Determine the common vocabulary set according to historical experience vocabulary, count the frequencies of common words, and use the sum of the mean and standard deviation of the vocabulary frequencies as the frequency threshold. Mark the words with frequencies lower than the frequency threshold as rare words, and combine them to form a keyword set as trigger words; Set a watermark trigger based on the Q&A pairs, and determine the number of trigger Q&A pairs according to the proportion of trigger Q&A pairs in the total Q&A pairs based on historical experience and the total number of Q&A pairs, and randomly select Q&A pairs to carry the watermark trigger (append or insert the trigger phrase to the answer of the target Q&A pair, that is, the trigger is appended to the end of the original answer to ensure that the logic of the original semantic answer is not damaged), and form a target set, where the number of trigger Q&A pairs is expressed as: ; Among them, represents the number of trigger Q&A pairs, represents the proportion of trigger Q&A pairs in the total Q&A pairs, and n represents the total number of Q&A pairs; Load the question-and-answer pair data based on the target set, distinguish the question-and-answer pair data carrying the watermark trigger, and use the question-and-answer pairs containing the trigger and those without the trigger as the training data set; For the target language model, calculate the generation probability of words in the context through softmax normalization. Based on the Logit perturbation mechanism algorithm, adjust and intervene in the trigger words containing triggers in the training data, and recalculate the generation probability of the trigger words, which is expressed as: ;

[0019] ; Among them, represents all the vocabulary sequences generated by the target language model before the current position t as the context condition, usually the question and the previously generated partial answer sequence, which is composed of the previously generated words and the question of the question-and-answer pair, represents at the current the conditional probability of the trigger word under the condition, represents at the current the Logit score of the trigger word under the condition, which represents the priority score of the trigger word in the context under the condition, represents all potential words generated by the language model, represents the candidate word, represents the Logit score given for each candidate word, which represents the priority of the candidate word in the current context under the condition, represents the trigger word, represents the enhancement parameter, which is used to add a fixed positive offset to the original Logit score and is determined based on historical experience, represents the adjusted Logit score, represents the adjusted conditional generation probability; Select the cross-entropy loss function to calculate the calculation loss between the class probability predicted by the target language model and the actual label according to the training data set, use the Adam optimizer for gradient descent optimization, and stop the iteration when the loss of the model no longer decreases significantly during continuous iterations, and output the target language model that has completed the training of the embedded trigger watermark.

[0020] By constructing a structured question-and-answer dataset and a question-and-answer embedding mapping, a unified and high-dimensional dense semantic representation basis is provided, making the subsequent screening based on semantic similarity and trigger design more context-sensitive, thereby enhancing the effectiveness of model watermark embedding. Through similarity screening and mapping back to the original answers, question-and-answer pairs with highly relevant semantics are screened out, greatly reducing the possible interference of low-quality or noisy question-and-answer pairs to model training, ensuring that the model focuses on highly relevant data. Based on the mean and standard deviation thresholds of the general word frequencies, rare and low-frequency words are identified and combined as the keyword library of trigger words, which can minimize the probability of repetition between trigger words and natural language, thus ensuring that these triggers will not be masked by the model generating ordinary answers due to high-frequency co-occurrence. The low-frequency characteristics of the triggers can maintain uniqueness in verification, making them have high identification ability in output generation, while enhancing the concealment and exclusivity of the triggers as watermarks. The method of determining the threshold based on word frequency provides an objective basis for trigger screening, avoiding the instability or lack of generality of trigger behavior caused by subjective trigger design; randomly selecting question-and-answer pairs to attach triggers and constructing a training dataset avoids the distribution law of triggers being detected by malicious analysts, further enhancing the concealment of triggers. During the attachment process, only the trigger phrase is inserted at the end of the answer, without destroying the semantics and logic of the question-and-answer, ensuring that the model retains its original performance in normal generation scenarios (i.e., the non-intrusiveness of the watermark). Adjust the generation priority of triggers through the Logit perturbation mechanism. The Logit change directly inserts a control signal into the model generation mechanism, making the trigger words have significant priority in the model generation distribution, thereby improving the reliability of trigger response. By adjusting the proportion of the trigger word priority in the probability distribution, it is ensured that the trigger behavior is obvious and the regulation effect is controllable. Based on the recalculation of the generation probability under context conditions, the trigger can adapt to any complex context conditions, further enhancing the triggering ability of the model in real question-and-answer scenarios. At the same time, the Logit interference has limited impact on the distribution of other non-trigger words and does not destroy the normal question-and-answer behavior of the model.

[0021] S2. According to the probability accumulation of the target language model predicting word by word when generating a complete answer, calculate the relevance score of the text question-and-answer pair, determine the retrieval relevance score based on the semantic embedding vector obtained through the embedding function, and determine the fusion sorting score, generate the optimal answer sequence, and classify it according to the question-and-answer difficulty. Combine the curriculum training method to optimize the model step by step from low difficulty to high difficulty. Preferably, determining the fusion sorting score, generating the optimal answer sequence, and classifying it according to the question-and-answer difficulty. Combining the curriculum training method to optimize the model step by step from low difficulty to high difficulty includes Based on the target language model trained with embedded trigger watermarks, the relevance score of the text Q&A pair is calculated according to the probability accumulation predicted word by word when the model generates the complete answer, expressed as: ; Where represents the relevance score of the Q&A pair, T represents the total number of words in the answer, represents the probability that the target language model predicts the current generated word given the question and the context ; Based on the semantic embedding vectors obtained through the embedding function, the retrieval relevance score is determined by calculating the cosine similarity, expressed as: ; Where represents the retrieval relevance score, represents the embedding vector of the question and the embedding vector of the answer ; The generation relevance score calculated by the generation model and the retrieval relevance score calculated by the retrieval model are used to design the fusion ranking score, expressed as: ; Where represents the final fusion relevance score, represents the fusion weight, which is determined through experiments; Sort the answers of all Q&A pairs according to the final fusion relevance score to generate the optimal answer sequence; According to the optimal answer sequence, calculate the complexity of each Q&A pair through the retrieval relevance score, expressed as: ; Where represents the Q&A pair data complexity value; The Q&A pairs are evenly divided from easy to complex according to the generation complexity. The number of divisions is determined by rounding down the logarithm of the number of samples in the dataset. According to the different difficulty Q&A pairs divided, the target language model is optimized level by level from low difficulty to high difficulty. According to the Q&A pair data of different difficulties as the current optimization training set, the cross-entropy loss function is selected for step-by-step optimization training. After completing the step-by-step optimization training of the highest difficulty Q&A pairs, stop the iterative optimization to obtain the optimized target language model.

[0022] In the calculation of generation relevance score and retrieval relevance score, the way of accumulating the per-word generation probability can quantify the compliance of the answer generated by the target language model. The higher the generation probability, the stronger the generation logic and fluency of the answer under the current context conditions of the model. Using the generation relevance score to evaluate the question-answer pair can directly reflect the performance of the model's generation ability in the actual context. The generation relevance score provides a key probability basis for subsequent fusion ranking, making the ranking not only rely on the retrieval task but also further conform to the generation nature of the language model. Calculating the cosine similarity based on the semantic embedding vector can effectively map the semantic distance between the question and the answer and obtain the matching degree of the two in the semantic space. The retrieval relevance supplements the semantic associations that may be missed in the pure generation task. Especially when the answer generation logic does not completely depend on the context, this retrieval strategy can play a cross-verification role. The combination of the two captures the accuracy of the model generation through the generation relevance and at the same time evaluates the model performance using the global relevance of semantic retrieval, which can more comprehensively evaluate the answer quality. Through fusion ranking, a balance will be established between the generation relevance and the retrieval relevance. By introducing reasonable weights, it is ensured that there is no over-reliance on the scores of a single dimension. The fusion ranking result can not only retain the advantage of the answer generation probability but also introduce the guarantee and trade-off of the semantic dimension, avoiding the wrong prioritization of answers with poor semantic associations that may be caused by the generation behavior completely dominating the ranking; Generate complex values based on the retrieval relevance score calculation, in cooperation with the previous fusion ranking score. The calculation of the complex values not only considers the semantic associations but also ensures the objectivity of the difficulty judgment through the ranking result. By dividing the question-answer pairs with the complex values and performing curriculum training, the dynamic nature of the division can be guaranteed, adapting to different scales of the dataset, avoiding the uneven distribution of training samples that may be caused by the traditional fixed partitioning method, achieving effective learning at a relatively small training cost, avoiding the training dilemmas that the model may fall into when directly facing mixed random training samples, and ensuring that the optimized model can be further detected and verified for watermark characteristics and transferred to the subsequent verification and traceability stage.

[0023] S3. Generate a dynamic trigger based on the previous trigger, a fixed key, and a hash function according to the trigger, and form a dynamic trigger sequence to perform the verification of the optimized target language model. Statistically calculate the verification matching value of the dynamic trigger and calculate the average generation matching value to analyze the matching degree of the optimized target language. Preferably, forming a dynamic trigger sequence to perform the verification of the optimized target language model, statistically calculating the verification matching value of the dynamic trigger and calculating the average generation matching value includes: Construct a trigger sequence based on the question-answer pairs carrying the watermark trigger and perform sequence dynamicization. Generate a dynamic trigger based on the previous trigger, a fixed key, and a hash function according to the trigger, and form a dynamic trigger sequence, which is expressed as: ; where represents the i-th dynamic trigger, represents the one-way irreversible hash function SHA-256, and K represents the system encryption key; Use the dynamic trigger sequence to optimize the verification of the target language model. For each dynamic trigger, splice it through the trigger question template content to generate the corresponding verification question, and input it into the optimized target language model to obtain the corresponding generated answer; Compare the generation logic output by the dynamic trigger to ensure that the subsequent dynamic triggers in the dynamic trigger sequence can be predicted by the current dynamic trigger and the system encryption key, expressed as: ; where represents the output answer question of the optimized target language model, represents the (i + 1)-th dynamic trigger; At the same time, for each dynamic trigger, through the Monte Carlo method, perform randomized trigger generation and verification operations, determine the random dynamic trigger based on historical experience data, and calculate the matching degree of the generated watermark according to the verification result, expressed as: ; where represents the verification matching value, represents the number of successfully matched keywords based on the s-th dynamic trigger in the generated answer, represents the answer text of the s-th trigger question, represents the total number of all target keywords in the trigger.

[0024] During the generation and use of the dynamic trigger sequence, the dynamic trigger is generated through the hash chain mechanism (based on the previous trigger and the system key), ensuring the uniqueness and unpredictability of the trigger during verification. This dynamic process avoids the risk of fixed triggers being exposed or attacked during the verification stage, making the trigger generation highly secure. External attackers cannot reverse-calculate and restore the trigger sequence or its generation rules, fundamentally enhancing the security and anti-tampering ability of the watermark verification mechanism; By verifying the continuity of the trigger sequence, the correct watermark behavior of the model is indirectly examined directly from the generation logic of the model. The good continuity of the verification sequence makes the verification process highly reliable, especially in complex verification scenarios, providing a secure dynamic channel for model verification. When forming verification questions by splicing dynamic triggers with question templates, it can be used to cover various verification scenarios, while the dynamic trigger part ensures that each verification question is different. This method increases the randomness and diversity of verification questions on the basis of a unified verification framework, reducing the possibility of verification questions being predicted or reused. The splicing design enables the dynamic trigger to naturally integrate into the context of the verification question. When the model generates answers, they match the standardized Q&A pairs in training, enhancing the semantic rationality of trigger verification and avoiding the generation of illogical answers. Randomly select partial subsets of the trigger sequence for verification through the Monte Carlo method, enhancing the randomness and flexibility of the verification process. Compared with linear verification (verifying all triggers one by one), this randomized method improves verification efficiency while reducing the risk of specific triggers being predicted and exploited. Combined with the matching degree calculation and the standardized verification process, it can comprehensively quantify the true performance strength of the watermark embedding behavior in the model's answers and avoid misjudgment caused by the failure of a single trigger.

[0025] Furthermore, analyze the matching degree of the optimized target language, including Calculate the average generation matching value for all dynamic trigger sets, expressed as: ; Where represents the average matching value, S represents the total number of dynamic triggers, represents the verification matching value of the s-th dynamic trigger; Based on the sum of the mean and standard deviation of the historical matching values as the matching threshold, if the average generation matching value is greater than or equal to the matching threshold, the model is determined to be the original watermark embedding version.

[0026] By averaging the matching values of all dynamic trigger sets, the average matching value directly reflects whether the model's response behavior to triggers in dynamic verification is stable. The introduction of the average value bridges the possible local noise and interference in dynamic verification, ensuring the stability and reliability of the determination mechanism. Combined with dynamic triggers, logical verification, and Monte Carlo verification, the complete matching degree and identity determination framework enhance the credibility of verifying the model's watermark behavior.

[0027] S4. Generate a report based on model verification and store it on the blockchain; Preferably, generating a report based on model verification includes Collect all trigger verification matching values and average matching values, use a report generation tool to generate reports for the target language model, and record the total number of verification triggers, the number of verifications, and the verification timestamps.

[0028] By collecting the matching values and average matching values of all triggers, the overall performance of the model in trigger watermark verification can be intuitively reflected after report formatting, facilitating the observation of the behavioral consistency of the model in multiple verification scenarios. By recording the total number of trigger verifications, the number of verifications, and the timestamps, time series information is established for each round of verification, facilitating the analysis of the stability and potential anomalies during the verification process afterwards.

[0029] Furthermore, perform blockchain storage, including Store the dynamic triggers generated in each round of verification in the blockchain. After the verification is completed, record the trigger verification matching values and average matching values in the answers of the target language model, and record the generated reports.

[0030] By recording the dynamic triggers in the blockchain, the triggers generated in each round of verification become independent blockchain transactions, ensuring a high degree of credibility in the verification input process. The distributed storage of the dynamic triggers on the blockchain makes the verification input completely public. This transparency can enhance the fairness and authority of the verification mechanism, playing a key role especially in scenarios of multi-party joint verification or dispute resolution. The record of the dynamic triggers corresponding to each verification is bound to the verification timestamp, and the trigger input, the model answers generated during verification, and the performance of the model in specific verification tasks can be traced at any time.

[0031] This embodiment also provides a large model traceability system based on Q&A data, including A data embedding module that collects a Q&A pair dataset, performs similarity screening and filtering using a semantic embedding function, and generates a keyword set as trigger vocabulary; A trigger module that generates a trigger vocabulary set based on low-frequency keywords, generates Q&A data with watermark triggers, counts the number of trigger Q&A pairs, and distinguishes between Q&A data with triggers and without triggers; A model training module that constructs contexts for the training dataset, calculates the probability of words generated in the contexts, and uses a Logit adjustment mechanism to train the target language model; A retrieval fusion module that uses the target language model to predict the complete generated answer word by word and calculates the generated correlation score, calculates the semantic retrieval correlation score of the Q&A pair through semantic embedding vectors, calculates the fusion sorting score according to the generated correlation score and the retrieval correlation score by weight, generates an ordered optimal answer sequence, and generates different difficulty levels of the Q&A pairs; A multi-level optimization module that, based on the hierarchical Q&A data output by the generation and retrieval module, gradually optimizes the target language model along the path from low to high; A dynamic trigger generation module that uses a hash chain mechanism to generate a sequence of dynamic triggers, splices verification questions with the dynamic triggers through trigger-based question templates to verify the logical continuity of the triggers, randomly selects triggers for verification using the Monte Carlo method, counts the verification match values, and calculates the average result of the match values; A report generation module that collects the trigger match values and average match values in each round of verification and generates a detailed verification report. A blockchain storage module that archives the sequence of dynamic triggers, verification match values, and generated reports into the blockchain.

[0032] This embodiment also provides a computer device applicable to the scenario of the large model traceability method based on Q&A data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the large model traceability method based on Q&A data proposed in the above embodiment.

[0033] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0034] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for tracing the large model based on Q&A data proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0035] In summary, the present invention randomly selects Q&A pairs to attach triggers and constructs a training data set, avoiding the distribution law of triggers being detected by malicious analysts, further enhancing the concealment of triggers. During the attachment process, only the trigger phrase is connected to the end of the answer, without damaging the Q&A semantics and logic, ensuring that the model retains its original performance in normal generation scenarios. Based on the recalculation of the generation probability under context conditions, the trigger can adapt to any complex context conditions, further enhancing the triggering ability of the model in real Q&A scenarios. At the same time, the Logit interference has a limited impact on the distribution of other non-trigger words and does not disrupt the normal Q&A behavior of the model. Through the verification of the continuity of the trigger sequence, it indirectly checks whether the model has the correct watermark behavior from the generation logic of the model. The good continuity of the verification sequence makes the verification process highly reliable. By randomly selecting a partial subset of the trigger sequence for verification using the Monte Carlo method, the randomness and flexibility of the verification process are enhanced. Compared with linear verification, this randomization method improves the verification efficiency and reduces the risk of specific triggers being predicted and exploited. By averaging the matching values of all dynamic trigger sets, the average matching value directly reflects whether the response behavior of the model to triggers in dynamic verification is stable. The introduction of the average value bridges the possible local noise and interference in dynamic verification, ensuring the stability and reliability of the decision-making mechanism. Combined with dynamic triggers, logical verification, and Monte Carlo verification, the complete matching degree and identity determination framework improve the credibility of verifying the watermark behavior of the model.

[0036] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A large model traceability method based on Q&A data, comprising, characterized in that: Collect a Q&A data set, perform data mapping through a semantic embedding function, and screen and form a similarity set according to semantic similarity. Statistically screen the frequency of common vocabulary to form a keyword set as trigger vocabulary. Set a watermark trigger based on Q&A pairs, and screen the number of trigger Q&A pairs to distinguish Q&A pair data carrying the watermark trigger. For the target language model, calculate the generation probability of words in the context, adjust and intervene in the trigger words containing triggers in the training data, and recalculate the generation probability of the trigger words. Through model training, output the target language model that has completed the training of the embedded trigger watermark; According to the cumulative probability of word-by-word prediction when the target language model generates a complete answer, calculate the relevance score of the text Q&A pair, determine the retrieval relevance score based on the semantic embedding vector obtained through the embedding function, and determine the fusion sorting score, generate the optimal answer sequence, and classify it according to the Q&A difficulty level. Combine the curriculum training method to optimize the model step by step from low difficulty to high difficulty; Generate a dynamic trigger based on the previous trigger, a fixed key, and a hash function according to the trigger, and form a dynamic trigger sequence to verify the optimized target language model. Statistically calculate the verification matching value of the dynamic trigger and calculate the average generation matching value to analyze the matching degree of the optimized target; Generate a report according to the model verification and store it on the blockchain.

2. The method for tracing the large model based on Q&A data according to claim 1, wherein: For the target language model, calculating the generation probability of words in the context and adjusting and intervening in the trigger words containing triggers in the training data includes: Collect a question-and-answer data set, which is organized in the form of binary question-and-answer pairs and denoted as , where represents the text content of the i-th question, represents the text content of the i-th corresponding answer, and a data set is formed according to all binary question-and-answer pairs; Use the pre-trained BERT model to construct a semantic embedding function Embed, and convert the Q&A pairs of the data set into semantic embedding vectors through the embedding function respectively, and store all semantic embedding vectors as a mapping set; Calculate the semantic similarity of the embedding vectors of each Q&A pair using cosine similarity, perform similarity screening based on historical experience values, form a similarity set with the remaining similarity values, and map them back to the corresponding original text Q&A pairs; Determine the common vocabulary set according to historical experience vocabulary, and statistically calculate the frequency of common vocabulary. Based on the sum of the mean and standard deviation of the vocabulary frequency as the frequency threshold, mark the vocabulary with a frequency lower than the frequency threshold as rare vocabulary, and combine them to form a keyword set as trigger vocabulary; Set a watermark trigger based on the question-and-answer pair, and determine the number of trigger question-and-answer pairs according to the proportion of the trigger question-and-answer pairs based on historical experience to the total number of question-and-answer pairs and the total number of question-and-answer pairs, and randomly select question-and-answer pairs to carry the watermark trigger to form a target set; Load the Q&A pair data based on the target set, and distinguish the Q&A pair data carrying the watermark trigger. Use the Q&A pairs with and without triggers as the training data set; For the target language model, calculate the generation probability of words in the context through softmax normalization. Based on the Logit perturbation mechanism algorithm, adjust and intervene in the trigger words containing triggers in the training data, and recalculate the generation probability of the trigger words; Select the cross-entropy loss function to calculate the calculation loss between the category probability predicted by the target language model and the actual label according to the training data set. Use the Adam optimizer for gradient descent optimization. Stop iterating when the loss of the model no longer decreases significantly during continuous iteration, and output the target language model that has completed the training of the embedded trigger watermark.

3. The method for tracing the origin of a large model based on Q&A data according to claim 2, characterized in that: The determination of the fusion sorting score, generating the optimal answer sequence, and grading according to the Q&A difficulty, combined with the curriculum training method, optimize the model step by step from low difficulty to high difficulty, including: Based on the target language model trained by embedding the trigger watermark, calculate the relevance score of the text Q&A pair according to the cumulative probability of the model predicting word by word when generating the complete answer; Based on the semantic embedding vector obtained through the embedding function, determine the retrieval relevance score by calculating the cosine similarity; Design the fusion sorting score based on the generation relevance score calculated by the generation model and the retrieval relevance score calculated by the retrieval model; Sort the answers of all Q&A pairs according to the final fusion relevance score to generate the optimal answer sequence; According to the optimal answer sequence, calculate the complexity of each Q&A pair through the retrieval relevance score; Evenly divide the Q&A pairs from simple to complex according to the generation complexity. The number of divisions is determined by rounding down the logarithm of the number of samples in the dataset. According to the Q&A pairs of different difficulties divided, optimize the target language model step by step from low difficulty to high difficulty. Use the cross-entropy loss function for step-by-step optimization training based on the Q&A pair data of different difficulties as the current optimization training set. After completing the step-by-step optimization training of the Q&A pairs with the highest difficulty, stop the iterative optimization to obtain the optimized target language model.

4. The method for tracing the large model based on Q&A data according to claim 3, wherein: The composition of the dynamic trigger sequence, the verification of the optimized target language model, and the statistical verification matching value of the dynamic trigger and calculation of the average generation matching value, including: Construct a trigger sequence according to the Q&A pairs carrying the watermark trigger, and perform sequence dynamicization. Generate a dynamic trigger based on the previous trigger, a fixed key, and a hash function according to the trigger, and form a dynamic trigger sequence; Use the dynamic trigger sequence to verify the optimized target language model. Concatenate the content of the trigger question template for each dynamic trigger to generate the corresponding verification question, and input it into the optimized target language model to obtain the corresponding generated answer; Compare the generation logic output by the dynamic trigger to ensure that the subsequent dynamic triggers in the dynamic trigger sequence can be predicted by the current dynamic trigger and the system encryption key; At the same time, for each dynamic trigger, perform randomized trigger generation and verification operations through the Monte Carlo method. Determine the random dynamic trigger based on historical experience data, and calculate the matching degree of the generated watermark according to the verification results.

5. The large model traceability method based on Q&A data according to claim 4, characterized in that: The analysis of the matching degree of the optimized target language; including: Calculate the average generation matching value for all dynamic trigger sets; Based on the sum of the mean and standard deviation of the historical matching values as the matching threshold, if the average generation matching value is greater than or equal to the matching threshold, then judge that the model is the original watermark embedding version.

6. The method for tracing the large model based on Q&A data according to claim 5, characterized in that: The generation of a report according to the model verification, including: Collect all trigger verification matching values and average matching values, use a report generation tool to generate a report for the target language model, and record the total number of verified triggers, the number of verification times, and the verification timestamp.

7. The method for tracing the large model based on Q&A data according to claim 6, characterized in that: The implementation of blockchain storage, including Store the dynamically generated triggers in each round of verification in the blockchain. After the verification is completed, record the trigger verification matching values and average matching values in the answers of the target language model, and record the generated reports.

8. A large model traceability system based on Q&A data, based on the large model traceability method based on Q&A data according to any one of claims 1 to 7, characterized in that, Including, A data embedding module that collects a question-and-answer pair dataset, performs similarity screening and filtering using a semantic embedding function, and generates a set of keywords as trigger vocabulary; A trigger module that generates a set of trigger vocabulary based on low-frequency keywords, generates question-and-answer data with watermark triggers, counts the number of trigger question-and-answer pairs, and distinguishes between question-and-answer data with and without triggers; A model training module that constructs contexts for the training dataset, calculates the probability of words generated in the contexts, and uses a Logit adjustment mechanism to train the target language model; A retrieval fusion module that uses the target language model to predict the complete generated answer word by word and calculates the generated correlation score, calculates the semantic retrieval correlation score of the question-and-answer pair through semantic embedding vectors, calculates the fusion ranking score according to the generated correlation score and the retrieval correlation score by weight, generates an optimal answer sequence after ranking, and generates different difficulty levels of the question-and-answer pair; A multi-level optimization module that gradually optimizes the target language model in a path from low to high based on the graded question-and-answer data output by the generation and retrieval modules; A dynamic trigger generation module that generates a sequence of dynamic triggers using a hash chain mechanism, splices the trigger questions with the dynamic triggers through a trigger question template, verifies the logical continuity of the triggers, randomly selects triggers for verification using the Monte Carlo method, counts the verification matching values, and calculates the average result of the matching values; A report generation module that collects the trigger matching values and average matching values in each round of verification and generates a detailed verification report; A blockchain storage module that archives the dynamic trigger sequence, verification matching values, and generated reports into the blockchain.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the large model traceability method based on question-and-answer data according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the large model traceability method based on question-and-answer data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-round dialogue training method for dialogue large model

    CN117556002A

  • Language processing question answering system and method based on AIGC large model

    CN118093834A

  • Port question and answer method based on large language model and related equipment thereof

    CN118410155A

  • Text robot application system based on large model

    CN119474323A

  • Large language model optimization generation method based on optimal cue word selection

    CN119476209A

Cited By

  • Question and answer data set copyright protection method and device based on backdoor watermark and perception encryption

    CN121765697A