Data processing method applied to model and electronic equipment

By implanting watermark information in the retrieval enhancement database and generating a question-answer hash chain, the problem of unauthorized document detection in the artificial intelligence model is solved, and the traceability of data ownership and the credibility of the model output are achieved.

CN120611017APending Publication Date: 2025-09-09LENOVO (BEIJING) LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510689735.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

How to detect whether the output information of the artificial intelligence model contains unauthorized document information to ensure the traceability and accuracy of the data source.

Method used

Watermark information is embedded in the documents in the retrieval enhanced database, and a question-answer hash chain is generated. The ownership characteristics of the document are verified through the watermark information and hash chain to ensure the document's immutability and semantic-level verification.

Benefits of technology

It improves the verification accuracy and detection convenience of the ownership of documents in the retrieval enhancement database, prevents data tampering, and ensures the credibility of the model output information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611017A_ABST
    Figure CN120611017A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method applied to a model and electronic equipment, and the method comprises the steps: carrying out the retrieval of a document in a retrieval enhancement database based on input information, and obtaining a target document; wherein a text unit in the target document is marked with corresponding watermark information; the target document is used for guiding the target model to output output information corresponding to the input information; based on the semantic information of the target document, generating a question and answer hash chain of the target document; and inputting the target document bound with the question and answer hash chain into the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and more specifically to a data processing method and electronic equipment applied to a model. Background Art

[0002] With the development of artificial intelligence (AI) models, information processing has become possible. These models include a retrieval-augmented generation system (RAG), which enhances the AI ​​model's information processing capabilities by searching for documents in corresponding databases. Detecting unauthorized document information in model output has become a challenge in the application of AI models. Summary of the Invention

[0003] In view of this, this application provides the following technical solutions:

[0004] A data processing method applied to a model, comprising:

[0005] Retrieving documents in a search enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide a target model to output output information corresponding to the input information;

[0006] Based on the semantic information of the target document, generating a question-answer hash chain of the target document;

[0007] The target document bound with the question-answer hash chain is input into the target model.

[0008] Optionally, the process of marking the documents in the retrieval enhancement database with watermark information includes:

[0009] Performing text segmentation on the document in the retrieval enhancement database to obtain at least one text unit;

[0010] Determining a document hash value, associated text unit feature information, and a percentage of text units of a target type for the document based on document features of the document;

[0011] If the current text unit to be annotated is of the target type, determining an associated hash value of an associated text unit corresponding to the text unit to be annotated based on the associated text unit feature information;

[0012] Calculating a watermark probability parameter of the text unit to be marked based on the document hash value, the associated hash value, and the ratio information of the text units having the target type;

[0013] If the watermark marking probability parameter satisfies the watermark marking condition, the text unit to be marked is marked with watermark information of a corresponding type to obtain a document marked with the watermark information.

[0014] Optionally, the calculating the watermark probability parameter of the text unit to be marked based on the document hash value, the associated hash value, and the ratio information of the text units having the target type includes:

[0015] Calculating an initial probability parameter of the text unit to be annotated based on the document hash value, the associated hash value, and the ratio information of the text units having the target type;

[0016] determining an adjustment coefficient based on the occurrence frequency of the text unit to be annotated in the document and / or the position information of the text unit to be annotated in the document;

[0017] Based on the adjustment coefficient and the initial probability parameter, a watermark marking probability parameter of the text unit to be marked is determined.

[0018] Optionally, generating a question-answer hash chain of the target document based on the semantic information of the target document includes:

[0019] extracting a question set from the target document based on semantic information of the target document;

[0020] Determining a first hash value of the question set based on a hash index corresponding to each question in the question set;

[0021] Determining a second hash value of an answer set corresponding to the question set based on a hash index of an answer corresponding to each question in the question set in the target document;

[0022] A question-answer hash chain is determined according to the first hash value and the second hash value.

[0023] Optionally, determining the first hash value of the question set based on the hash index corresponding to each question in the question set includes:

[0024] Convert the text corresponding to each question in the question set into a vector to obtain a question vector;

[0025] Determine a hash index corresponding to each question based on the hash mapping matrix and the question vector;

[0026] Perform an XOR operation on the hash index corresponding to each question to obtain a first hash value of the question set.

[0027] Optionally, it also includes:

[0028] The question set, the first hash index, the answer set, the second hash index, and the question-answer hash chain are stored in a trusted storage area.

[0029] Optionally, it also includes:

[0030] Based on the question-answer hash chain bound to the target document, ownership verification information of the document to which the output information of the target model is applied is determined, where the ownership verification information indicates whether the document to which the output information is applied is derived from the retrieval enhancement database.

[0031] Optionally, determining ownership verification information of a document to which the output information of the target model is applied based on the question-answer hash chain bound to the target document includes:

[0032] determining a first answer hash value of answer information corresponding to the input information in the output information of the target model;

[0033] Determine a second answer hash value of the answer information corresponding to the input information according to the question-answer hash chain bound to the target document;

[0034] The ownership verification information of the document to which the output information of the target model is applied is determined according to the first answer hash value and the second answer hash value.

[0035] Optionally, determining ownership verification information of a document to which the output information of the target model is applied based on the first answer hash value and the second answer hash value includes:

[0036] Calculating a similarity parameter between the first answer Hash value and the second answer Hash value;

[0037] If the similarity parameter is greater than a target threshold, obtaining watermark information corresponding to the output information of the target model;

[0038] If the watermark information matches the watermark information annotated with the text unit in the target document, it is determined that the document to which the target model output information is applied originates from the retrieval enhancement database.

[0039] An electronic device, comprising:

[0040] A storage module, used for storing a retrieval enhancement database;

[0041] The processing module is used to execute the application program to achieve:

[0042] Retrieving documents in the retrieval enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide the target model to output output information corresponding to the input information;

[0043] Based on the semantic information of the target document, generating a question-answer hash chain of the target document;

[0044] The target document bound with the question-answer hash chain is input into the target model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0046] Figure 1 A flow chart of a data processing method applied to a model provided in an embodiment of the present application;

[0047] Figure 2 A schematic diagram of a flow chart of a method for marking watermark information provided in an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a processing flow for generating a question-answer hash chain for a target document provided in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of an application architecture provided in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0052] In this application, the terms "first" and "second" are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.

[0053] The embodiment of the present application provides a data processing method applied to a model, which can be applied to the process in which a target model processes data. The target model can be a large language model (LLM), which is used to process natural language tasks. The target model can apply a retrieval enhancement system to improve the accuracy of the model's data processing. The retrieval-augmented generation (RAG) system corresponds to a retrieval enhancement database. The input information input to the target model can be used to obtain relevant documents in the retrieval enhancement database. The document is used to guide the output of the target model, thereby improving the accuracy of the target model output. In order to accurately identify whether documents or data in the retrieval enhancement database are applied, in the embodiment of the present application, the documents in the retrieval enhancement database can be specifically processed to achieve the addition of tamper-proof ownership information to the documents in the retrieval enhancement database, so that the data can be verified based on the ownership information whether it is included in the retrieval enhancement database, thereby improving the effectiveness and accuracy of the verification.

[0054] See also Figure 1 , is a flow chart of a data processing method applied to a model provided in an embodiment of the present application. The method may include the following steps:

[0055] S101: Search documents in a search enhancement database based on input information to obtain target documents.

[0056] The retrieval enhancement system corresponds to a retrieval enhancement database, which includes multiple documents. These documents can be applied to the model's data processing process. In order to achieve traceability protection of data ownership in the retrieval enhancement database, hidden ownership identification can be embedded in the documents included in the retrieval enhancement database, such as adding watermark information to the document. For example, the text unit in the document is marked with corresponding watermark information, wherein the text unit represents the granularity unit for dividing the text in the document, such as a word, a term, or a corresponding text segment. Different watermark information can be added to text units with different attribute characteristics, such as adding watermark information of different colors, or different types of watermark information. For example, the text unit representing the technical keyword in the document is marked with a green watermark, and the ownership can be determined based on the watermark information in subsequent use.

[0057] In the embodiment of the present application, watermark information can be pre-marked on documents in the retrieval enhancement database according to the corresponding watermark information marking method. Watermark information can be marked on all documents in the retrieval enhancement database, or on documents of a specified type or application range in the retrieval enhancement database. Watermark information marking can also be performed on the document after the corresponding document is retrieved using the retrieval enhancement system. Correspondingly, when marking watermark information in the retrieval enhancement database, a corresponding marking method can be selected for marking, or key parameters for watermark information marking can be set, and watermark information of related documents can be marked based on the key parameters, wherein the key parameters can include the type of watermark information, the type of corresponding text unit, and the distribution characteristics of the watermark information. In this way, ownership verification can be performed later based on the watermark information marked in the document.

[0058] The documents in the retrieval enhancement database in the embodiment of the present application are used to guide the target model to output output information corresponding to the input information. Therefore, in order to improve the accuracy of the target model output, the target document matching the input information can be retrieved in the retrieval enhancement database based on the retrieval enhancement system. The input information can be converted into a coding vector, and then the vector representation of each document in the retrieval enhancement database is obtained. According to the vector similarity retrieval, the target document matching the input information is obtained. The text units in the target document are marked with corresponding watermark information, and the target document is used to guide the target model to output output information corresponding to the input information. In the embodiment of the present application, by adding corresponding watermark information to the documents in the retrieval enhancement database, the subsequent verification of the ownership of the documents in the retrieval enhancement database can be facilitated.

[0059] S202: Generate a question-answer hash chain of the target document based on the semantic information of the target document.

[0060] By marking watermark information in the text unit of the document in the retrieval enhancement database, it is possible to embed a hidden statistical signal to facilitate the verification of the ownership characteristics of the document. The watermark information can prevent the tampering of the document. In order to further improve the anti-attack on the document ownership information and solve the problem of semantic tampering, in the embodiment of the present application, it is also possible to process according to the semantic information of the target document to generate a question-and-answer hash chain of the target document that represents its semantic identification, that is, the semantic-level verification of the document can be achieved through the question-and-answer hash chain. The semantic vectors of the core question-and-answer pairs of the documents in the retrieval enhancement database can be bound into a unique chain hash through the corresponding processing mode to obtain a question-and-answer hash chain. In the question-and-answer hash chain, each hash node corresponds to a semantic fragment of "question and answer". The overall question-and-answer hash chain forms an indivisible unique identifier of the document. Any local modification will destroy the hash consistency. Therefore, the ownership information of the document can be verified through the question-and-answer hash chain.

[0061] S103: Input the target document bound with the question-answer hash chain into the target model.

[0062] Since the text units in the target document are marked with corresponding watermark information, and the question-answer hash chain is bound to the target document and then input into the target model, the target document input into the target model includes not only the watermark information but also the question-answer hash chain. Even if the target model rewrites the relevant information in the process of determining the output information using the target document, it is still possible to use the watermark information and the question-answer hash chain to determine whether the documents in the retrieval enhancement database are applied to the output information.

[0063] In an embodiment of the present application, watermark information can be added to documents in the retrieval-enhanced database and bound to a question-and-answer hash chain. The cryptographic characteristics of the hash chain and the statistical characteristics of the watermark can be combined to prevent attackers from simultaneously deciphering both aspects of information. The watermark information is covertly integrated into the document text without affecting its readability, while the question-and-answer hash chain ensures unalterable ownership. This information can be used to provide an effective detection method to identify whether the data applied by the target model is included in the retrieval-enhanced database, improving the convenience and accuracy of detection.

[0064] The following describes the data processing method applied to the model in the embodiment of the present application in combination with actual application scenarios.

[0065] In the embodiment of the present application, a method for marking watermark information of a document is also provided. The method can be applied to mark watermark information of documents in a search enhancement database. Figure 2 , the method may include the following steps:

[0066] S201: Divide the documents in the retrieval enhancement database to obtain at least one text unit.

[0067] S202: Based on the document features of the document, determine the document hash value, associated text unit feature information, and ratio information of text units of the target type.

[0068] S203: If the current text unit to be annotated is of the target type, determine the associated hash value of the associated text unit corresponding to the text unit to be annotated based on the associated text unit feature information.

[0069] S204: Calculate a watermark probability parameter of the text unit to be marked based on the document hash value, the associated hash value, and the ratio information of the text units of the target type.

[0070] S205: If the watermark marking probability satisfies the watermark marking condition, the text unit to be marked is marked with watermark information of a corresponding type to obtain a document marked with the watermark information.

[0071] In step S201, the documents in the search enhancement database are segmented into text. The segmentation can be performed based on the corresponding text granularity, such as a certain character length, or based on the semantic content of the text, such as keyword text units. In an embodiment of the present application, the text segmentation mode can be determined based on the semantic content of the document. For example, different types of documents can adopt different segmentation modes, thereby obtaining at least one text unit corresponding to the document, such as obtaining multiple text units. The text unit can be a character, a word, or a text segment.

[0072] In step S202, the relevant parameters for adding watermark information can be determined based on the document characteristics of the document. The relevant parameters include the document hash value of the document, the characteristic information of the associated text units, and the proportion of text units of the target type. The document characteristics can be determined based on characteristics such as the document subject content, the document type, and the text length. The document hash value of the document can represent the unique identifier of the document, or the unique identifier of the document can be identified by a salt value to ensure that the watermark modes of different documents are independent. For example, a random number can be added to the algorithm corresponding to the calculation of the document hash value to enhance the security of the hash value.

[0073] The features of the associated text units represent the features of the text units adjacent to the current text unit to be annotated. For example, the features of the associated text units can represent the number of text units that the current text unit to be annotated depends on. For example, the context width h represents the number of text units that the current text unit to be annotated depends on. When h = 2, it means that the current text unit to be annotated depends on the hash values ​​of the previous two text units. Then, based on the associated text unit feature information, the associated hash value of the associated text unit corresponding to the text unit to be annotated can be determined. In this way, the associated hash value can be used to subsequently determine the watermark annotation probability parameter of the text unit to be annotated.

[0074] The proportion information of text units with the target type can represent the proportion information of text units with a specific watermark type, or the proportion information of text units of a specific type in the document. For example, if the watermark information includes a green watermark and a red watermark, the proportion information of the green watermark in the document can be set. For example, the green ratio is represented by the parameter γ, and the text unit is represented by a vocabulary. γ = 0.25 can mean that 25% of the text units in the vocabulary of the hash value are marked as green watermarks, and 75% of the text units are marked as red watermarks. For another example, the text units are divided into ordinary text units and keyword text units. The target type can be a keyword type. Therefore, the proportion information of text units with the target type can represent the proportion of the number of keyword text units in all text units in the document.

[0075] In step S204, if the unit of the text to be annotated is of the target type, the watermark annotation probability parameter of the text unit to be annotated is calculated based on the document hash value, the associated hash value, and the ratio of text units of the target type. The watermark annotation probability parameter can be used to determine whether the current text unit to be annotated can be annotated with the watermark information corresponding to the target type. That is, if the watermark annotation probability parameter meets the watermark annotation conditions, the text unit to be annotated is annotated with the watermark information of the corresponding type to obtain the document annotated with the watermark information. The above steps S201-S205 are performed for each text unit to be annotated in the document until the watermark information annotation of each text unit in the document is completed.

[0076] In the embodiment of the present application, the corresponding parameters can also be adjusted to achieve the concealment and anti-attack performance of the added watermark information. In one embodiment, based on the document hash value, the associated hash value and the proportion information of the text units with the target type, the watermark annotation probability parameter of the text unit to be annotated is calculated, including: calculating the initial probability parameter of the text unit to be annotated based on the document hash value, the associated hash value and the proportion information of the text units with the target type; determining the adjustment coefficient based on the frequency of occurrence of the text unit to be annotated in the document and / or the position information of the text unit to be annotated in the document; and determining the watermark annotation probability parameter of the text unit to be annotated based on the adjustment coefficient and the initial probability parameter.

[0077] In this embodiment, to avoid the problem of excessive watermarking of common words or excessive watermarking of non-important parts of a document, the high-frequency offset can be reduced. The value corresponding to the initial probability parameter can be determined as a fixed offset for the current text unit to be annotated. Then, based on the adjustment coefficient and the initial probability parameter, the final watermark probability parameter corresponding to the text unit to be annotated is determined. The adjustment coefficient is determined based on the frequency of occurrence and / or position of the text unit to be annotated in the document. For example, if the text unit to be annotated appears frequently, the adjustment coefficient can be used to reduce its probability of becoming watermark information.

[0078] The following is an example of the process of watermarking text units in a document.

[0079] When watermarking information, the text units in the document are represented by tokens, and the watermark information includes red watermarks and green watermarks. The relevant parameters for watermark information annotation can be pre-set according to the characteristics of the document, such as the context width h of the associated text unit feature information that characterizes the text unit to be annotated. h = 2, that is, whether the current token is marked with a red watermark or a green watermark depends on the hash value of the first two tokens corresponding to it. Further, the proportion information of the text units that characterize the target type can be set, such as setting the green watermark ratio γ. When γ = 0.25, it means that 25% of the tokens in the vocabulary table according to the hash value are marked green and 75% of the tokens are marked red. The initial probability parameters of the text unit token to be annotated can also be set, that is, the original output score is expressed in Logits. By adjusting these scores, the generation probability of a specific watermark can be controlled. For example, the Logits offset δ: δ t = 2.0·(1+λ·f(t)), where f(t) represents the word frequency of the tth token in the target dataset (reducing the watermark weight of high-frequency words), and λ is the adjustment coefficient (also called the perturbation factor, for example, λ can be set between 0.1 and 0.5). When δ = 2.0, the logits value of the green token increases by 2.0, increasing the probability of being sampled. Document hash values ​​can also be set based on document characteristics to ensure that watermark patterns for different documents are independent.

[0080] The process of watermarking a text unit token is as follows:

[0081] First, the original document is segmented into text units according to semantics to obtain several text units (represented by tokens). These text units are used as the basic information for watermarking. The position of each watermark unit to be marked in the document is represented by t. According to the hash values ​​of the previous h tokens and the document hash value represented by salt, the hash value of the current token to be marked as a green watermark or a red watermark can be calculated, that is, the initial probability parameter, such as Hash t express:

[0082] Hash t =sha256(token t-2 ,token t-1 ,salt)

[0083] Among them, token t-2 , token t-1 They represent the hash values ​​of the first two text units of the current text unit to be annotated, sha256 represents the hash algorithm currently used, and salt represents the document hash value.

[0084] You can set the offset δ of the initial probability parameter Logits of the text unit token marked with green watermark t , the corresponding watermark can be embedded at the level of a single token, or it can be extended to n-gram to enhance stability, where n-gram is a sequence structure used to represent n consecutive text units (such as words, characters or subwords) in natural language processing. For example, 2-gram (big tuple) splits the sentence "quantum chip efficiency" into ["quantum chip", "chip efficiency"]. In an embodiment of the present application, the n-gram-level watermark enhances the anti-rewriting ability by adjusting the overall generation probability of n consecutive words (rather than a single word). Even if the attacker modifies part of the vocabulary (such as rewriting "quantum chip" to "quantum processor"), as long as enough watermark marks are retained in the n-gram (such as m≥n / 2 green words), the statistical features can still be detected. This design significantly improves the robustness of the watermark to local rewriting while maintaining the fluency of the text. Based on these two processing methods, the watermark annotation probability parameter (in p) of the current text unit to be annotated is calculated wn (t) represents) the process is:

[0085] Single token level processing: in is an indicator function, t represents the target token whose probability is currently being calculated, and t′ represents all tokens in the vocabulary, which is used for traversal and summation. It takes 1 when t belongs to the green list and 0 otherwise.

[0086] Since token-level watermarks may be affected by the attacker's token-by-token rewriting or repetition, which reduces the watermark's detection ability, n-gram-level watermarks are introduced to enhance robustness. The watermark affects not only a single token, but the entire n-gram (i.e., n consecutive tokens); by adjusting the logits at the n-gram level, the entire phrase has a stronger watermark signal, rather than a single word. In this way, even if the RAG system performs word-level paraphrasing (such as "100qubits" is rewritten as "one hundred quantum bits"), the statistical features at the n-gram level can still remain consistent, ensuring detection capabilities. At this point, the watermark annotation probability parameter of the current text unit token to be annotated is calculated as:

[0087]

[0088] Where w=(t1,t2,...t n ) represents an n-gram; the dynamic watermark weight is expressed as The dynamic green vocabulary can be expressed as If there are at least m tokens in the n-gram, they are marked green. For example, when m = n / 2, it means that at least half of the tokens in the n-gram need to be marked.

[0089] Through the above process, documents marked with watermark information can be obtained, and these documents can be used as documents in the retrieval enhancement database. In the embodiment of the present application, the marking of the text unit watermark in the document depends on the context and the document hash value, which is unpredictable by attackers. By fine-tuning the corresponding parameters, the generated text can still remain natural and smooth. For example, the statistical deviation of the text unit marked with green watermark can be detected through multiple query accumulation, providing information for ownership verification.

[0090] In the embodiment of the present application, in addition to marking the corresponding watermark information on the text units in the document, the semantic information of the document can also be used to generate the corresponding question-and-answer hash chain to add an unalterable proof of ownership from a semantic perspective.

[0091] See also Figure 3 , which is a schematic diagram of a processing flow for generating a question-answer hash chain for a target document provided in an embodiment of the present application. This method can generate a question-answer hash chain for a target document based on the semantic information of the target document, and may include the following steps:

[0092] S301. Extract a question-answer set from a target document based on the semantic information of the target document.

[0093] S302. Determine a first hash value of the question set based on a hash index corresponding to each question in the question set.

[0094] S303. Determine a second hash value of the answer set corresponding to the question set based on the hash index of the answer corresponding to each question in the question set in the target document.

[0095] S304: Determine a question-answer hash chain based on the first hash value and the second hash value.

[0096] Using a general hashing method can only achieve a complete match. If the data is rewritten (for example, a different expression), the original hash will become invalid. In order to support fuzzy matching, it is necessary to adopt semantic hashing, that is, to calculate the low-dimensional semantic representation of the text based on a deep learning model, so that semantically similar texts can obtain similar hash values. Therefore, in the embodiment of the present application, a semantic hash is bound to the documents in the retrieval enhancement database for subsequent verification of ownership information.

[0097] Taking the question-answer hash chain of the target document as an example, based on the semantic information of the target document, a question set is extracted from the target document. The question set mainly represents the question set generated by the key entities in the document. The question set is represented by Q, Q = (q1, q2, ..., qk ), that is, there are k questions in the problem set. Then, for each of the problem sets, its corresponding hash index can be calculated, and then the first hash value corresponding to the problem set can be obtained.

[0098] In one implementation of an embodiment of the present application, the process of determining the first hash value of the question set based on the hash index corresponding to each question in the question set may include: converting the text corresponding to each question in the question set into a vector to obtain the question vector; determining the hash index corresponding to each question based on the hash mapping matrix and the question vector; performing an XOR operation on the hash index corresponding to each question to obtain the first hash value of the question set.

[0099] For example, for each question q i , calculate its hash index H(q i )=BERT(q i )W h ; Among them, BERT(q i ) is the semantic vector of the question after BERT encoding. BERT encoding is the process of converting text (such as words, sentences, or paragraphs) into high-dimensional semantic vectors through the pre-trained BERT model. The generated vector can capture the deep semantic information related to the context. h is a hash mapping matrix used to project the output vector of BERT into a low-dimensional hash space in order to calculate fuzzy matching. Finally, H(q i ) is a low-dimensional vector used for approximate hash matching, such as H(q i ) is a 1*256 floating point vector, such as [-0.12,0.88,-0.43,...,1.12].

[0100] The hash of the question sequence (as represented by the first hash value) is The hash values ​​of multiple questions are XORed to form an overall bound hash. Even if a question is rewritten, the semantics remain unchanged and the overall hash remains stable.

[0101] After obtaining the question set, the corresponding answer set is bound to it, so that semantically similar questions can still find the same answer. This way, if the search enhancement system generates a tampered answer, it can also be detected. The answer set corresponding to the question set can be determined, and then the second hash value corresponding to the answer set can be determined. Referring to the process of determining the hash value of the question set mentioned above, the process of determining the second hash value corresponding to the answer set can include:

[0102] Define candidate answers and determine the answer set T = (t1, t2, ..., t m ) means that the answer set includes m answers. Similar to the question vector generation method, the hash index H(t i)=BERT(t i )W h , then the second hash value of the answer set can be obtained as:

[0103] Then, the question and answer hash binding is performed based on the first hash value and the second hash value to obtain a question and answer hash chain, as represented by H.

[0104] To facilitate the tracing of relevant information during the verification process and enhance the security of hash binding, in embodiments of the present application, the question set, first hash index, answer set, second hash index, and question-answer hash chain can be stored in a trusted storage area. Alternatively, the storage can be in a blockchain, ensuring that this information is publicly verifiable.

[0105] In the embodiment of the present application, by marking the documents in the search enhancement database with watermark information and binding the question-answer hash chain, it is convenient to subsequently verify whether the relevant documents are from the search enhancement database. Accordingly, in one implementation of the embodiment of the present application, it also includes:

[0106] Based on the question-and-answer hash chain bound to the target document, the ownership verification information of the document used in the output information of the target model is determined. The ownership verification information indicates whether the document used in the output information originates from the retrieval-enhanced database. Hash binding detection can be used to verify whether the document used in the output information of the target model has a binding relationship with the retrieval-enhanced database, thereby identifying whether the data is included in the retrieval-enhanced database.

[0107] Correspondingly, based on the question-answer hash chain bound to the target document, the ownership verification information of the document applied to the output information of the target model is determined, including: determining a first answer hash value of the answer information corresponding to the input information in the output information of the target model; determining a second answer hash value of the answer information corresponding to the input information according to the question-answer hash chain bound to the target document; and determining the ownership verification information of the document applied to the output information of the target model according to the first answer hash value and the second answer hash value.

[0108] The corresponding question in the input information can be used to obtain the question set Q associated with the document, and then the question set is sent to the retrieval enhancement system to determine the relevant answer set, such as R, R = (r1, r2, ..., r n ). Then calculate the first answer hash value H(R) corresponding to the answer set, and perform similarity calculation with the second hash value H(T) of the answer set in the pre-stored bound question and answer hash chain. For example, sim(H(R),H(T)) represents the similarity parameter, that is:

[0109]

[0110] Then, a similarity threshold can be set, such as 0.85. If the value corresponding to the calculated similarity parameter is greater than 0.85, watermark detection will continue. Otherwise, relevant information such as no protected data will be detected will be output.

[0111] Because watermark information is added to the document, its ownership information can be further verified based on the watermark information. Accordingly, based on the first answer hash value and the second answer hash value, the ownership verification information of the document to which the output information of the target model applies is determined, including: calculating a similarity parameter between the first answer hash value and the second answer hash value; if the similarity parameter is greater than a target threshold, obtaining the watermark information corresponding to the output information of the target model; if the watermark information matches the watermark information annotated with the text unit in the target document, determining that the document to which the output information of the target model applies originates from the retrieval enhancement database.

[0112] In the embodiment of the present application, the ownership of the data can be obtained through watermark detection, and the amount of use of the corresponding data can be measured to determine the degree of data leakage.

[0113] In the process of watermark detection, the proportion of the documents marked with the corresponding type of watermark information can be calculated. For example, taking the green watermark as an example, the total number of text units (expressed in tokens) in the current document can be counted (expressed as Γ), and the number of text units corresponding to the green watermark can also be counted, such as |s| g For example, watermark information can be verified by the Z-score test method, which is used to determine the relationship between a data point and the mean of a set of data. It measures the relative position of the data point by calculating the number of standard deviations between the data point and the mean. Z-score test:

[0114]

[0115] p=1—Φ(z)

[0116] Where γ represents the target ratio of green watermarks, γΓ represents the theoretically expected number of text units corresponding to green watermarks, γ(1-γ)Γ represents the variance of the binomial distribution, which characterizes the reasonable fluctuation range of the number of green watermark text units in natural text; z represents the standardized outlier, which measures the degree to which the actual number of green words deviates from the theoretical expectation (unit: standard deviation σ).

[0117] When the result p<10 -5 , indicating that the data leakage is significant; it can also be, for example: 100 queries, total token number Γ = 5000, green token number |s| g =1400, then z = 4.9, p < 10 -6 .

[0118] See also Figure 4 , is a schematic diagram of an application architecture corresponding to an application scenario provided in an embodiment of the present application, wherein the application architecture includes a data annotation module, a retrieval enhancement system (in Figure 4 ), and may also include a hash binding module, a target model (in Figure 4 The application architecture also includes a joint verification module (represented by LLM).

[0119] Among them, the data annotation module can be used to retrieve the watermark information corresponding to the document sky in the enhanced database, such as the watermark annotation model (such as Figure 4 The watermark LLM in the watermark annotation model is then used to input the document to be annotated into the watermark annotation model to obtain the watermarked document. The hash binding module can perform question-answer hash binding on the watermarked document. For example, the module extracts the question sequence and determines the question sequence hash; then determines the answer sequence hash to generate a question-answer hash chain, i.e., the response sequence hash.

[0120] When a user enters a query, the RAG retrieval system retrieves the corresponding document, which is watermarked and bound to a question-answer hash chain. This document is then fed into the LLM, which then outputs the corresponding output information. The joint verification module can further verify document ownership, such as confirming data usage through hash chain detection and further confirming data ownership through watermark detection. The specific implementation of this application architecture is described in the previous examples and will not be detailed here.

[0121] In another embodiment of the present application, an electronic device is provided. Figure 5 , the electronic device comprises:

[0122] Storage module 501, used for storing the search enhancement database;

[0123] The processing module 502 is configured to execute an application program to implement:

[0124] Retrieving documents in the retrieval enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide the target model to output output information corresponding to the input information;

[0125] Based on the semantic information of the target document, generating a question-answer hash chain of the target document;

[0126] The target document bound with the question-answer hash chain is input into the target model.

[0127] Retrieving documents in the retrieval enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide the target model to output output information corresponding to the input information;

[0128] Based on the semantic information of the target document, generating a question-answer hash chain of the target document;

[0129] The target document bound with the question-answer hash chain is input into the target model.

[0130] In another embodiment of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data processing method of the application model as described in any one of the above items is implemented.

[0131] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.

[0132] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0133] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0134] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0135] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method applied to a model, comprising: Retrieving documents in a search enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide a target model to output output information corresponding to the input information; Based on the semantic information of the target document, generating a question-answer hash chain of the target document; The target document bound with the question-answer hash chain is input into the target model.

2. The method according to claim 1, wherein The process of annotating the documents in the retrieval enhancement database with watermark information includes: Performing text segmentation on the document in the retrieval enhancement database to obtain at least one text unit; Determining a document hash value, associated text unit feature information, and a percentage of text units of a target type for the document based on document features of the document; If the current text unit to be annotated is of the target type, determining an associated hash value of an associated text unit corresponding to the text unit to be annotated based on the associated text unit feature information; Calculating a watermark probability parameter of the text unit to be marked based on the document hash value, the associated hash value, and the ratio information of the text units having the target type; If the watermark marking probability parameter satisfies the watermark marking condition, the text unit to be marked is marked with watermark information of a corresponding type to obtain a document marked with the watermark information.

3. The method according to claim 2, wherein the step of calculating the watermark probability parameter of the text unit to be annotated based on the document hash value, the associated hash value, and the percentage information of the text units having the target type comprises: Calculating an initial probability parameter of the text unit to be annotated based on the document hash value, the associated hash value, and the ratio information of the text units having the target type; determining an adjustment coefficient based on the occurrence frequency of the text unit to be annotated in the document and / or the position information of the text unit to be annotated in the document; Based on the adjustment coefficient and the initial probability parameter, a watermark marking probability parameter of the text unit to be marked is determined.

4. The method according to claim 1, wherein generating a question-answer hash chain of the target document based on the semantic information of the target document comprises: extracting a question set from the target document based on semantic information of the target document; Determining a first hash value of the question set based on a hash index corresponding to each question in the question set; Determining a second hash value of an answer set corresponding to the question set based on a hash index of an answer corresponding to each question in the question set in the target document; A question-answer hash chain is determined according to the first hash value and the second hash value.

5. The method according to claim 4, wherein determining the first hash value of the question set based on the hash index corresponding to each question in the question set comprises: Convert the text corresponding to each question in the question set into a vector to obtain a question vector; Determine a hash index corresponding to each question based on the hash mapping matrix and the question vector; Perform an XOR operation on the hash index corresponding to each question to obtain a first hash value of the question set.

6. The method according to claim 4, further comprising: The question set, the first hash index, the answer set, the second hash index, and the question-answer hash chain are stored in a trusted storage area.

7. The method according to claim 1, further comprising: Based on the question-answer hash chain bound to the target document, ownership verification information of the document to which the output information of the target model is applied is determined, where the ownership verification information indicates whether the document to which the output information is applied is derived from the retrieval enhancement database.

8. The method according to claim 7, wherein determining ownership verification information of a document to which the output information of the target model is applied based on the question-answer hash chain bound to the target document comprises: determining a first answer hash value of answer information corresponding to the input information in the output information of the target model; Determine a second answer hash value of the answer information corresponding to the input information according to the question-answer hash chain bound to the target document; The ownership verification information of the document to which the output information of the target model is applied is determined according to the first answer hash value and the second answer hash value.

9. The method according to claim 8, wherein determining ownership verification information of a document to which the output information of the target model is applied based on the first answer hash value and the second answer hash value comprises: Calculating a similarity parameter between the first answer Hash value and the second answer Hash value; If the similarity parameter is greater than a target threshold, obtaining watermark information corresponding to the output information of the target model; If the watermark information matches the watermark information annotated with the text unit in the target document, it is determined that the document to which the target model output information is applied originates from the retrieval enhancement database.

10. An electronic device comprising: A storage module, used for storing a retrieval enhancement database; The processing module is used to execute the application program to achieve: Retrieving documents in the retrieval enhancement database based on input information to obtain a target document; wherein text units in the target document are marked with corresponding watermark information; and the target document is used to guide the target model to output output information corresponding to the input information; Based on the semantic information of the target document, generating a question-answer hash chain of the target document; The target document bound with the question-answer hash chain is input into the target model.

Citation Information

Patent Citations

  • Text generation method and device and electronic equipment

    CN116956906A

  • Text watermark embedding and detecting method based on model context learning

    CN118349970A

  • Question and answer processing method, device and equipment

    CN119961415A

  • Systems and methods for retrieval based question answering using neura network models

    US20240428044A1