Data processing method and device

By determining the watermark descriptor based on signature information in the search knowledge base and building and optimizing the watermark information, the problem of watermarks being easily removed in the prior art is solved, and the protection reliability of the knowledge base is improved.

CN120217331APending Publication Date: 2025-06-27LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286955.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, text knowledge bases that generate RAGs based on retrieval enhancement are vulnerable to attacks, and the attacker can easily remove the watermark to bypass protection.

Method used

By determining the relevant watermark descriptor based on the signature information in the search knowledge base, the watermark information in the retrieved text is concealed and difficult to be removed. The specific methods include constructing multiple watermark information, adjusting the statement description method of the description statement, integrating the watermark information, and optimizing the watermark information using semantic coherent parameters.

Benefits of technology

Improve the reliability of protecting the search knowledge base, making the watermark information more difficult to remove or tamper with by the attacker, thereby enhancing the protection effect of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217331A_ABST
    Figure CN120217331A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device. The method comprises the steps of obtaining a data query request; the data query request comprises a query keyword; according to the query keyword, performing text retrieval in a retrieval knowledge base to obtain a retrieval text; wherein the retrieval text comprises at least one piece of watermark information, the watermark information is obtained based on a watermark description word contained in the retrieval knowledge base, and the watermark description word is related to signature information corresponding to the retrieval knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data processing method and apparatus. Background Art

[0002] Currently, in order to protect a text knowledge base based on Retrieval-augmented Generation (RAG), watermarks can be embedded into the retrieved text. For example, watermark embedding can be achieved by modifying the text format, replacing words, or performing grammatical transformations.

[0003] However, this method is vulnerable to attacks, such as paraphrasing attacks and irrelevant content removal attacks. An attacker can easily remove the watermark and thus bypass the protection. Summary of the Invention

[0004] In view of this, this application provides a data processing method and apparatus as follows:

[0005] A data processing method includes:

[0006] Obtain a data query request; the data query request includes a query keyword;

[0007] According to the query keyword, perform text retrieval in a retrieval knowledge base to obtain a retrieved text;

[0008] Wherein, the retrieved text contains at least one watermark information, and the watermark information is obtained based on watermark description words included in the retrieval knowledge base, and the watermark description words are related to signature information corresponding to the retrieval knowledge base.

[0009] In the above method, preferably, the retrieval knowledge base contains multiple watermark information so that the retrieved text contains at least one of the watermark information.

[0010] In the above method, preferably, the multiple watermark information in the retrieval knowledge base is constructed by the following method:

[0011] In the retrieval knowledge base, obtain multiple first description statements containing the watermark description words;

[0012] Adjust the statement description manner of the first description statement to obtain a second description statement;

[0013] Integrate the second description statement with the first description statement to obtain watermark information.

[0014] In the above method, preferably, after integrating the second description statement with the first description statement, the method further includes:

[0015] Obtain a semantic coherence parameter between the second description statement and the first description statement, where the semantic coherence parameter characterizes the semantic coherence of the text composed of the second description statement and the first description statement;

[0016] Obtain a third description statement according to the semantic coherence parameter;

[0017] Integrate by substituting the third description statement for the second description statement and the first description statement to obtain new watermark information.

[0018] In the above method, preferably, obtaining a third description statement according to the semantic coherence parameter includes:

[0019] Adjust the statement description manner of the first description statement according to the semantic coherence parameter to obtain a third description statement; wherein, the statement description manner of the third description statement is different from that of the second description statement;

[0020] Or, according to the semantic coherence parameter, use the signature information to obtain new watermark descriptors; in the retrieval knowledge base, obtain multiple new first description statements containing the new watermark descriptors; adjust the statement description manner of the new first description statements to obtain a third description statement.

[0021] In the above method, preferably, the watermark descriptors are obtained by the following method:

[0022] Perform text parsing on the first text in the retrieval knowledge base to obtain at least one triple corresponding to the first text, where the triple includes an entity descriptor and a relationship descriptor, and the relationship descriptor characterizes the entity relationship between the entity descriptors in the triple where it is located;

[0023] Wherein, the entity descriptors form an entity list, the relationship descriptors form a relationship list, and both the entity descriptors and the relationship descriptors are initial descriptors included in the first text;

[0024] Use the signature information to screen out multiple watermark descriptors from the initial descriptors.

[0025] In the above method, preferably, using the signature information to screen out multiple watermark descriptors from the initial descriptors includes:

[0026] Use the signature information to screen out multiple first descriptors in the entity list;

[0027] According to the first descriptors and the signature information, screen out multiple second descriptors in the relationship list;

[0028] Among them, the first descriptor and the second descriptor are watermark descriptors.

[0029] For the above method, preferably, using the signature information, multiple first descriptors are screened from the entity list, including:

[0030] Select an entity descriptor from the entity list as the current descriptor;

[0031] Determine the current descriptor as the first descriptor;

[0032] Perform a hash calculation on the current descriptor using the signature information to obtain a first index position, and the first index position points to the target descriptor in the entity list;

[0033] Use the target descriptor as the new current descriptor, and execute the step of determining the current descriptor as the first descriptor until the number of first descriptors meets the first screening condition.

[0034] For the above method, preferably, according to the first descriptor and the signature information, multiple second descriptors are screened from the relationship list, including:

[0035] Select two entity candidate words from the multiple first descriptors;

[0036] Perform a hash calculation on the two entity candidate words using the signature information to obtain a second index position, and the second index position points to the target descriptor in the relationship list;

[0037] Determine the target descriptor as the second descriptor, and the number of second descriptors meets the second screening condition.

[0038] A data processing device includes:

[0039] A request acquisition unit, configured to acquire a data query request; the data query request includes a query keyword;

[0040] A text retrieval unit, configured to perform text retrieval in a retrieval knowledge base according to the query keyword to obtain a retrieval text;

[0041] Among them, the retrieval text includes at least one watermark information, and the watermark information is obtained based on the watermark descriptors included in the retrieval knowledge base, and the watermark descriptors are related to the signature information corresponding to the retrieval knowledge base.

[0042] A computer device / system includes: a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the data processing method described in any one of the above.

[0043] A computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the data processing method described in any one of the above is implemented.

[0044] A computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the data processing method described in any one of the above is implemented.

[0045] As can be seen from the above technical solutions, in a data processing method and apparatus disclosed in the present application, relevant watermark description words are determined based on signature information corresponding to a retrieval knowledge base, so that watermark information in the text retrieved from the retrieval knowledge base is obtained based on these watermark description words. It can be seen that in the present application, watermark description words are determined based on signature information, so that the watermark information in the retrieved text has concealment and is not easily removed, thereby improving the reliability of protecting the retrieval knowledge base. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present application;

[0048] Figure 2 It is a partial flowchart of a data processing method provided by an embodiment of the present application;

[0049] Figure 3 It is another partial flowchart of a data processing method provided by an embodiment of the present application;

[0050] Figure 4 It is yet another partial flowchart of a data processing method provided by an embodiment of the present application;

[0051] Figure 5 It is another partial flowchart of a data processing method provided by an embodiment of the present application;

[0052] Figure 6 It is yet another partial flowchart of a data processing method provided by an embodiment of the present application;

[0053] Figure 7 It is a schematic structural diagram of a data processing apparatus provided by an embodiment of the present application;

[0054] Figure 8Another structural schematic diagram of a data processing device provided by an embodiment of the present application;

[0055] Figure 9 Yet another structural schematic diagram of a data processing device provided by an embodiment of the present application;

[0056] Figure 10 Structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0058] Refer to Figure 1 As shown, it is a flowchart of the implementation of a data processing method provided by an embodiment of the present application. This method can be applied to an electronic device capable of data processing. The technical solution in this embodiment is mainly used to improve the reliability of the protected retrieval knowledge base.

[0059] Specifically, the method in this embodiment may include the following steps:

[0060] Step 101: Obtain a data query request.

[0061] Among them, the data query request includes a query keyword. For example, if the user wants to know the way to Scenic Spot A, then enter "How to get to Scenic Spot A" in the search interface provided by the retrieval tool. Based on this input operation and the input content "How to get to Scenic Spot A", a data query request is generated, and the data query request includes query keywords "Scenic Spot A" and "How to get to".

[0062] Step 102: Perform text retrieval in the retrieval knowledge base according to the query keyword to obtain a retrieval text.

[0063] Among them, the retrieval text contains at least one watermark information, and the watermark information is obtained based on the watermark description words included in the retrieval knowledge base. The watermark description words are related to the signature information corresponding to the retrieval knowledge base.

[0064] It should be noted that the watermark information can be understood as: a description statement containing watermark description words.

[0065] Specifically, the watermark description word can be a description word included in the retrieved knowledge base. In this way, the watermark description word is a description word determined from the retrieved knowledge base according to the signature information, so that the retrieved text contains watermark information obtained based on such a watermark description word. In this way, the retrieved text contains a watermark description word that can be determined based on the signature information. The owner of the retrieved knowledge base can determine whether the retrieved text is obtained through the retrieved knowledge base through the watermark description word, thereby achieving the purpose of the retrieved knowledge base.

[0066] As can be seen from the above technical solution, in a data processing method provided by an embodiment of the present application, relevant watermark description words are determined based on the signature information corresponding to the retrieved knowledge base, so that the watermark information in the text retrieved from the retrieved knowledge base is obtained based on these watermark description words. It can be seen that in this embodiment, the watermark description words are determined based on the signature information, so that the watermark information in the retrieved text is concealed and not easily removed, thereby improving the reliability of protecting the retrieved knowledge base.

[0067] In one implementation, the retrieved knowledge base contains multiple pieces of watermark information, so that the retrieved text contains at least one piece of watermark information.

[0068] That is to say, in this embodiment, the watermark description words are determined from the retrieved knowledge base according to the signature information, and multiple pieces of watermark information are constructed in the retrieved knowledge base based on these watermark description words. Therefore, after the retrieved text is retrieved from the retrieved knowledge base according to the query keyword, the retrieved text contains at least one piece of watermark information containing the watermark description word, and the intellectual property rights of the retrieved knowledge base are represented by these watermark information.

[0069] In one implementation, in this embodiment, after determining the watermark description word according to the signature information, the description statement containing the watermark description word in the retrieved knowledge base can be used as the watermark information.

[0070] In another implementation, multiple pieces of watermark information in the retrieved knowledge base can be constructed in the following manner, as Figure 2 shown in:

[0071] Step 201: In the retrieved knowledge base, obtain multiple first description statements containing the watermark description word.

[0072] For example, taking the watermark description word "student x" as an example, in the retrieved knowledge base, determine the first description statement "student x won the ** award in 2020" and so on that contains this watermark description word.

[0073] Again, taking the watermark description word "teacher y" as an example, in the retrieved knowledge base, determine the first description statement "the supervisor of the paper *** published by student x is teacher y" and so on that contains this watermark description word.

[0074] Step 202: Adjust the statement description manner of the first description statement to obtain a second description statement.

[0075] For example, adjust the statement description manner of the first description statement "Student x won the ** award in 2020" to obtain the second description statement "The ** award in 2020 was won by student x".

[0076] Another example, adjust the statement description manner of the first description statement "The supervisor of the paper *** published by student x is teacher y" to obtain the second description statement "Teacher y supervised student x, and student x published the paper ***".

[0077] Step 203: Integrate the second description statement with the first description statement to obtain watermark information.

[0078] Specifically, in step 203, the second description statement can be spliced after or before the first description statement to obtain multiple pieces of watermark information. In fact, both the first description statement and the second description statement are watermark information. After splicing the statements, the resulting text contains watermark descriptors in the description statements, that is, watermark information.

[0079] Based on the above implementation, after step 203, the following processing can also be performed, such as Figure 3 as shown in

[0080] Step 204: Obtain the semantic coherence parameter of the second description statement and the first description statement.

[0081] Among them, the semantic coherence parameter characterizes the semantic coherence of the text composed of the second description statement and the first description statement. For example, when the semantic coherence parameter is greater than the corresponding threshold, it characterizes the semantic coherence of the text composed of the second description statement and the first description statement; when the semantic coherence parameter is not greater than the corresponding threshold, it characterizes the semantic incoherence of the text composed of the second description statement and the first description statement.

[0082] Specifically, in step 204, the semantic coherence of the text spliced by the second description statement and the first description statement can be parsed through a large language model.

[0083] Step 205: Obtain a third description statement according to the semantic coherence parameter.

[0084] Specifically, in step 205, when the semantic coherence parameter characterizes the semantic incoherence of the text composed of the second description statement and the first description statement, a third description statement containing watermark descriptors is re-obtained.

[0085] In one implementation manner, in step 205, the statement description manner of the first description statement can be adjusted according to the semantic coherence parameter to obtain the third description statement.

[0086] Among them, the statement description method of the third description statement is different from that of the second description statement, and the statement description method of the third description statement is also different from that of the first description statement.

[0087] For example, the first description statement "The supervisor of the paper *** published by student x is teacher y" is re-adjusted in its statement description method to obtain the third description statement "Student x published the paper *** under the guidance of teacher y".

[0088] In another implementation, in step 205, new watermark description words can be obtained according to semantic coherence parameters by using signature information; in the retrieval knowledge base, multiple new first description statements containing the new watermark description words are obtained; the statement description method of the new first description statements is adjusted to obtain the third description statement.

[0089] That is to say, when it is found that the text semantics of the watermark information is incoherent after the construction of the watermark information, the watermark description words can be re-determined, and then new first description statements can be determined. Based on the description statements after the adjustment of the statement description method, the watermark information is re-constructed, so that the semantics of the re-constructed watermark information is coherent.

[0090] Step 206: Replace the second description statement with the third description statement and integrate it with the first description statement to obtain new watermark information.

[0091] For example, in this embodiment, the third description statement can be spliced before or after the first description statement to obtain new watermark information.

[0092] In one implementation, the watermark description words can be obtained in the following ways, as Figure 4 shown in

[0093] Step 401: Parse the first text in the retrieval knowledge base to obtain at least one triple corresponding to the first text.

[0094] Among them, the triple includes entity description words and relationship description words. The relationship description word represents the entity relationship between the entity description words in its triple. The entity description words form an entity list, and the relationship description words form a relationship list. The entity description words and relationship description words are both initial description words included in the first text.

[0095] For example, the entity list can be represented by E, which contains multiple entity description words parsed from the first text. The relationship list can be represented by R, which contains multiple relationship description words parsed from the first text. The relationship description words in the relationship list represent the entity relationship between the corresponding two entity description words in the entity list. For example, teacher y and student x are in a teacher-student relationship. Another example is that there is an inclusion relationship between furniture and a dining table.

[0096] In one implementation, in this embodiment, a large language model can be used to perform triple parsing on the first text to obtain at least one triple.

[0097] It should be noted that the first text can be a selected part of the text in the retrieval knowledge base. For example, the first text can be randomly selected from the retrieval knowledge base.

[0098] Step 402: Use the signature information to screen out multiple watermark descriptors from the initial descriptors.

[0099] Among them, the signature information matches the retrieval knowledge base, and the signature information represents the owner of the retrieval knowledge base, such as the identification ID (Identity document) or the ownership signature of the owner.

[0100] Specifically, in this embodiment, the matching descriptors, that is, the watermark descriptors, can be screened out from the initial descriptors according to the signature information.

[0101] In one implementation, in step 402, multiple first descriptors can be screened out from the entity list by using the signature information first; then, according to the first descriptors and the signature information, multiple second descriptors can be screened out from the relationship list.

[0102] That is to say, in this embodiment, the signature information can be used to screen out the entity descriptors as watermark descriptors first, and then based on the screened entity descriptors and the signature information, the relationship descriptors as watermark descriptors can be screened out.

[0103] In a specific implementation, in step 402, the signature information can be used in the following way to screen out multiple first descriptors from the entity list, as Figure 5 shown in

[0104] Step 501: Select an entity descriptor in the entity list as the current descriptor.

[0105] Among them, in step 501, the entity descriptor ranked first in the entity list can be selected as the current descriptor or a random entity descriptor in the entity list can be selected as the current descriptor.

[0106] Step 502: Determine the current descriptor as the first descriptor, that is, the watermark descriptor.

[0107] Step 503: Monitor whether the number of the first descriptors meets the first screening condition. If not, execute step 504. If so, end the current process.

[0108] Among them, the first screening condition can be: the number of the first descriptors reaches the first threshold, such as 10.

[0109] Step 504: Calculate the hash of the current descriptor using the signature information to obtain a first index position.

[0110] Among them, the first index position points to the target descriptor in the entity list.

[0111] Specifically, in step 504, a cryptographic hash function with a key, such as the key-related Hash-based Message Authentication Code (HMAC), can be used to calculate the hash of the signature information key and the current descriptor to obtain the first index position As shown in formula (1):

[0112]

[0113] Among them, is the i-th entity descriptor in the entity list E, is the (i + 1)-th entity descriptor in the entity list E.

[0114] Step 505: Use the target descriptor as the new current descriptor, and execute step 502, step 503, and step 504.

[0115] That is to say, after calculating the first index position, the target descriptor pointed to by the first index position in the entity list can be determined as the watermark descriptor. In this way, if the number of the first descriptors reaches the first threshold, then no longer continue to screen the watermark descriptors in the entity list. If the number of the first descriptors does not reach the first threshold, then the newly screened watermark descriptor can be used in combination with the signature information to screen the next watermark descriptor until the number of the first descriptors reaches the first threshold.

[0116] In a specific implementation manner, in step 402, the following processing can be executed multiple times according to the second screening condition to screen multiple second descriptors in the relationship list, such as Figure 6 shown in

[0117] Step 601: Select two entity candidate words from multiple first descriptors.

[0118] Step 602: Calculate the hash of the two entity candidate words using the signature information to obtain a second index position, and the second index position points to the target descriptor in the relationship list.

[0119] Specifically, in step 602, HMAC can be used to calculate the hash of the signature information key and the two entity candidate words to obtain the second index position As shown in formula (2):

[0120]

[0121] Wherein, is the i-th entity descriptor selected as the watermark descriptor in the entity list E, is the j-th entity descriptor selected as the watermark descriptor in the entity list E, is the relationship descriptor pointed to by the second index position in the relationship list R.

[0122] Step 603: Determine the target descriptor as the second descriptor, and the number of the second descriptors satisfies the second screening condition.

[0123] Wherein, the second screening condition may be: the number of the second descriptors reaches a second threshold, such as 15. Based on this, according to the Figure 6 process, execute 15 times, so as to screen out 15 relationship descriptors from the relationship list as watermark descriptors.

[0124] Based on the above implementation solution, in this embodiment, after obtaining the entity descriptor and the relationship descriptor, before screening the watermark descriptor, first delete the descriptors with a frequency lower than the corresponding threshold, and retain the descriptors with a frequency higher than the target threshold, so as to avoid outliers by eliminating rare types of entity descriptors and relationship descriptors, thereby enhancing the concealment of the subsequent obtained watermark descriptors.

[0125] In one implementation, the retrieved text is obtained by adding at least one watermark message to the original text retrieved from the retrieval knowledge base. Specifically, step 102 can be implemented in the following manner:

[0126] First, according to the query keyword, retrieve at least one description statement in the retrieval knowledge base that matches the query keyword to obtain a retrieval result, and the retrieval result can be marked as the original text; then, splice at least one watermark message in the original text.

[0127] Based on this, the watermark message spliced for the original text in this embodiment can be obtained in the following manner:

[0128] First, in the original text, obtain at least one first description statement containing a watermark descriptor, and then adjust the statement description method of the first description statement to obtain a second description statement, and the second description statement is the watermark message. Based on this, by splicing the second description statement as the watermark message after the original text, that is, the first description statement, the retrieved text can be obtained.

[0129] Further, after splicing at least one watermark information into the original text, the obtained retrieval text can be semantically coherence parsed to obtain the semantic coherence parameter of the retrieval text, where the semantic coherence parameter characterizes the semantic coherence of the retrieval text; then, according to the semantic coherence parameter, the watermark information in the retrieval text is adjusted, such as readjusting the statement description manner of the first description statement to obtain a new second description statement, and then splicing the second description statement after the first description statement to obtain a new retrieval text.

[0130] Reference Figure 7 , which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. The device can be deployed in an electronic device capable of data processing. The technical solution in this embodiment is mainly used to improve the reliability of protecting the retrieval knowledge base.

[0131] Specifically, the device in this embodiment may include the following units:

[0132] A request acquisition unit 701, configured to acquire a data query request; the data query request includes a query keyword;

[0133] A text retrieval unit 702, configured to perform text retrieval in the retrieval knowledge base according to the query keyword to obtain a retrieval text;

[0134] Wherein, the retrieval text includes at least one watermark information, and the watermark information is obtained based on the watermark description words included in the retrieval knowledge base, and the watermark description words are related to the signature information corresponding to the retrieval knowledge base.

[0135] It can be seen from the above technical solution that in a data processing device provided by an embodiment of the present application, relevant watermark description words are determined based on the signature information corresponding to the retrieval knowledge base, so that the watermark information in the text retrieved from the retrieval knowledge base is obtained based on these watermark description words. It can be seen that in this embodiment, the watermark description words are determined based on the signature information, so that the watermark information in the retrieved text has concealment and is not easily removed, thereby improving the reliability of protecting the retrieval knowledge base.

[0136] In one implementation, the retrieval knowledge base includes multiple watermark information so that the retrieval text includes at least one of the watermark information.

[0137] In one implementation, the device may further include the following units, as Figure 8 shown in

[0138] A watermark construction unit 703 is configured to construct the multiple pieces of watermark information in the retrieval knowledge base in the following manner: in the retrieval knowledge base, obtain multiple first description statements containing the watermark descriptor; adjust the statement description manner of the first description statements to obtain second description statements; and integrate the second description statements with the first description statements to obtain multiple pieces of watermark information.

[0139] In one implementation manner, after integrating the second description statements with the first description statements, the watermark construction unit 703 is further configured to: obtain a semantic coherence parameter between the second description statements and the first description statements, where the semantic coherence parameter characterizes the semantic coherence of the text formed by the second description statements and the first description statements; obtain a third description statement according to the semantic coherence parameter; and replace the second description statements and the first description statements with the third description statement for integration to obtain new watermark information.

[0140] Among them, when the watermark construction unit 703 obtains the third description statement according to the semantic coherence parameter, it is specifically configured to:

[0141] Adjust the statement description manner of the first description statements according to the semantic coherence parameter to obtain a third description statement, where the statement description manner of the third description statement is different from that of the second description statements; or, obtain a new watermark descriptor by using the signature information according to the semantic coherence parameter; in the retrieval knowledge base, obtain multiple new first description statements containing the new watermark descriptor; and adjust the statement description manner of the new first description statements to obtain a third description statement.

[0142] In one implementation manner, the apparatus may further include the following units, as Figure 9 shown in:

[0143] A watermark determination unit 704 is configured to obtain the watermark descriptor in the following manner: perform text parsing on a first text in the retrieval knowledge base to obtain at least one triple corresponding to the first text, where the triple includes an entity descriptor and a relationship descriptor, and the relationship descriptor characterizes the entity relationship between the entity descriptors in the triple where it is located; among them, the entity descriptors form an entity list, the relationship descriptors form a relationship list, and both the entity descriptors and the relationship descriptors are initial descriptors included in the first text; and screen out multiple watermark descriptors from the initial descriptors by using the signature information.

[0144] In one implementation, when the watermark determination unit 704 uses the signature information to screen out multiple watermark descriptors from the initial descriptors, it specifically is configured to: use the signature information to screen out multiple first descriptors in the entity list; and according to the first descriptors and the signature information, screen out multiple second descriptors in the relationship list; wherein, the first descriptors and the second descriptors are watermark descriptors.

[0145] In one implementation, when the watermark determination unit 704 uses the signature information to screen out multiple first descriptors in the entity list, it specifically is configured to: select an entity descriptor in the entity list as the current descriptor; determine the current descriptor as a first descriptor; perform a hash calculation on the current descriptor by using the signature information to obtain a first index position, where the first index position points to a target descriptor in the entity list; use the target descriptor as the new current descriptor, and execute the operation of determining the current descriptor as a first descriptor until the number of first descriptors meets a first screening condition.

[0146] In one implementation, when the watermark determination unit 704 screens out multiple second descriptors in the relationship list according to the first descriptors and the signature information, it specifically is configured to: select two entity candidate words from the multiple first descriptors; perform a hash calculation on the two entity candidate words by using the signature information to obtain a second index position, where the second index position points to a target descriptor in the relationship list; and determine the target descriptor as a second descriptor, where the number of second descriptors meets a second screening condition.

[0147] It should be noted that the specific implementation manners of the units in this embodiment may refer to the corresponding content in the foregoing, and will not be elaborated herein.

[0148] Reference Figure 10 , is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may include the following structures:

[0149] A memory 1001, configured to store a computer program and data generated by running the computer program;

[0150] A processor 1002, configured to execute the computer program to implement: obtaining a data query request; the data query request includes a query keyword; performing a text retrieval in a retrieval knowledge base according to the query keyword to obtain a retrieval text; wherein, the retrieval text includes at least one piece of watermark information, and the watermark information is obtained based on watermark descriptors included in the retrieval knowledge base, and the watermark descriptors are related to signature information corresponding to the retrieval knowledge base.

[0151] As can be seen from the above technical solution, in an electronic device provided by an embodiment of the present application, relevant watermark description words are determined based on the signature information corresponding to the retrieval knowledge base, so that the watermark information in the text retrieved from the retrieval knowledge base is obtained based on these watermark description words. It can be seen that in this embodiment, the watermark description words are determined based on the signature information, so that the watermark information in the retrieved text is concealed and not easily removed, thereby improving the reliability of protecting the retrieval knowledge base.

[0152] Taking the knowledge base with RAG as an example of the retrieval knowledge base, the technical solution of the present application will be illustrated by way of example as follows:

[0153] The present application proposes a brand-new black-box watermark method for the RAG system, embedding watermark information in the form of text into the knowledge base, rather than directly embedding it into the retriever or the large language model LLM (Large Language Model), thereby improving the concealment and security of the watermark. In the present application, three components, namely a watermark generation module, a simulation attack system, and a watermark retrieval module, interact to generate high-quality watermark text and ensure that the watermark knowledge information can still be retrieved after being processed by the LLM. Connect the watermark text with the most relevant text in the knowledge base to improve the retrieval success rate of the watermark and enhance its robustness. The technical solution of the present application provides a new solution for the intellectual property protection of the RAG system and has important application value.

[0154] In the specific solution, the present application proposes a black-box watermark method, injecting watermark information into the knowledge base in text form, which can be abstracted into entities and relationships. In the present application, the watermark is first represented as a set of triples in the form of where is an entity descriptor (hereinafter referred to as an entity or entity word), and represents the relationship between them. This structured form simplifies watermark generation and intellectual property infringement detection. Importantly, the watermark injection process can maintain the availability of the RAG and the effectiveness of the watermark.

[0155] Moreover, due to the attacker's lack of expertise in the RAG knowledge base, watermark information containing entities or relationships in the original knowledge base can be injected in the present application, thereby enhancing the concealment of the watermark.

[0156] In addition, in order to verify the watermark, the entities and their relationships injected in the present application are real, but contain watermark information known only to the owner. These watermark information can then be extracted from the output of the LLM to detect intellectual property infringement.

[0157] The watermark injection process proposed by the present application is divided into two main steps:

[0158] 1. Watermark knowledge generation, that is, determining watermark description words:

[0159] (1) Entity and relationship extraction:

[0160] In this application, entities and relationships are first extracted from the original knowledge base, and these entities and relationships will be used as candidates for constructing watermark entities and relationships. Specifically, this application can use an LLM (e.g., the Transformer of the LLM) to parse and classify the entities and their relationships in the text documents of the knowledge base. However, extracting all entities and relationships from a large number of text documents in the knowledge base will be very expensive because of their huge volume. Therefore, in this application, a subset of the text documents in the knowledge base can be randomly selected to create an entity list E and a relationship list R.

[0161] (2) High-frequency screening:

[0162] Since the original entity list E and relationship list R may contain rare types of entities and relationships, this application reduces the size of the lists by focusing on high-frequency entities and relationships, and can also avoid outliers and enhance the concealment of the watermark. Therefore, in this application, the final entity list E and relationship list R are generated, and their sizes are |E| and |R| respectively.

[0163] (3) Generation of watermark entities and relationships:

[0164] To construct the watermark text, this application generates a set of triples in the format of according to the entity list E and the relationship list R. For intellectual property infringement detection, the owner's signature or identifier ID must be embedded in these watermark tuples. However, directly embedding the signature or ID into the entity or relationship will affect the concealment and reduce the watermark quality.

[0165] To solve this problem, this application uses the signature as the key in HMAC to generate the watermark tuples

[0166] Specifically, this application generates a watermark entity list Ewm for embedding watermark information. The first entity ew0 is initialized to Null, and the subsequent generation method uses formulas (1) and (2) to filter out watermark words (i.e., watermark description words) in the entity list E and the relationship list R.

[0167] 2. Generation and embedding of watermark text:

[0168] (1) Interaction framework: The solution proposed in this application uses three components, namely the watermark generation module, the simulated attacker system, and the watermark detection module, to interact and generate high-quality watermark text.

[0169] The watermark generation module refers to the one introduced above that uses an LLM to generate watermark text containing watermark entities, relationships, and relationship types.

[0170] The main function of the simulation attack system is to simulate the way the LLM deployed by the attacker processes the watermarked text, so as to judge whether the watermark is successfully embedded. It is not directly used to generate the final answer, but to analyze whether the retrieved text contains watermark information.

[0171] The watermark detection module refers to checking whether the correct watermark relationship is contained in the output of the simulation attack system. If it is contained, it means that the watermark embedding is successful; otherwise, feedback information is returned to the watermark generation module for adjustment.

[0172] (2) Connection of relevant texts, that is, splicing in the previous text:

[0173] Considering that RAG can retrieve the top k texts highly relevant to the question from the knowledge base, in order to improve the successful retrieval rate of the watermarked text, this application can generate and inject multiple different watermarked texts WT (water text) according to the same This application proposes a watermark injection method based on concatenation of relevant texts. Use the watermark query to retrieve the most relevant text TEXT, that is, the original text, from the original RAG, and then connect the watermarked text WT with the original text TEXT to generate the final watermarked text, that is, TEXT⊕WT.

[0174] To ensure the quality of the concatenated text TEXT⊕WT, this application can further use the LLM to evaluate the semantic coherence of the result. This step is crucial for maintaining the fluency and logical structure of the combined text. If the semantic coherence check passes, the concatenated text is considered valid. Otherwise, the watermark generation module needs to adjust the generated text to improve the quality.

[0175] (3) Watermark embedding:

[0176] Insert the generated watermarked text TEXT⊕WT into the original RAG to generate the final watermarked RAG. The watermarked RAG can be provided to the retriever for text retrieval.

[0177] Based on the above scheme, the watermark detection steps proposed in this application are as follows:

[0178] 1) Generation of watermark queries:

[0179] Generate watermark queries based on the embedded watermark entities and relationship triples. These queries usually appear in the form of questions, such as: "What is the relationship between entity A and entity B?" or "Please introduce the relevant knowledge between entity A and entity B?" The purpose of these queries is to trigger the model to return relevant content containing watermark information.

[0180] 2) Text retrieval:

[0181] Retrieve the most relevant text content from the suspected RAG system using the generated watermark query. Through this step, the text related to the watermark query returned by the RAG system can be obtained, and further analyze whether it contains watermark information.

[0182] 3) Text processing:

[0183] Once the relevant text is retrieved, it is passed to the simulation attack system. The simulation attack system conducts in-depth analysis and processing on the retrieved text to generate an answer. By simulating and analyzing the processing method of the RAG model, potential watermark information can be better extracted.

[0184] 4) Relationship verification:

[0185] After text processing, verify the generated answer to check whether it contains the specific relationship in the watermark query. For example, whether the relationship between the entities involved in the query is mentioned. If the relationship exists, it indicates that the output of the RAG system contains the preset watermark information.

[0186] 5) Result analysis:

[0187] Analyze according to the verification result. If the correct watermark relationship is detected, it indicates that the watermark embedding is successful, and the suspected RAG system is using a pirated model with a watermark, i.e., the knowledge base. If the correct watermark relationship is not detected, it may mean that the watermark has been removed or tampered with by the attacker, or the suspected RAG system does not use a pirated model with a watermark.

[0188] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0189] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0190] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium well known in the art.

[0191] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, comprising: Get data query request; The data query request includes a query keyword; According to the query keywords, a text search is performed in a search knowledge base to obtain a search text; The search text contains at least one watermark information, the watermark information is obtained based on the watermark description words contained in the search knowledge base, and the watermark description words are related to the signature information corresponding to the search knowledge base.

2. According to the method of claim 1, the retrieval knowledge base contains multiple watermark information, so that the retrieval text contains at least one of the watermark information.

3. According to the method of claim 2, the plurality of watermark information in the retrieval knowledge base is constructed in the following manner: In the search knowledge base, a plurality of first description sentences containing the watermark description word are obtained; adjusting the statement description mode of the first description statement to obtain a second description statement; The second description statement is integrated with the first description statement to obtain watermark information.

4. The method according to claim 3, after integrating the second description statement with the first description statement, the method further comprises: Obtaining a semantic coherence parameter between the second description sentence and the first description sentence, wherein the semantic coherence parameter represents semantic coherence of a text composed of the second description sentence and the first description sentence; Obtaining a third description sentence according to the semantic coherence parameter; The third description statement replaces the second description statement and is integrated with the first description statement to obtain new watermark information.

5. The method according to claim 4, obtaining a third description sentence according to the semantic coherence parameter, comprising: According to the semantic coherence parameter, adjusting the sentence description mode of the first description sentence to obtain a third description sentence; wherein the sentence description mode of the third description sentence is different from the sentence description mode of the second description sentence; Or, according to the semantic coherence parameter, using the signature information, obtain a new watermark descriptor; in the retrieval knowledge base, obtain multiple new first description sentences containing the new watermark descriptor; adjust the sentence description method of the new first description sentence to obtain a third description sentence.

6. According to the method of claim 1 or 2, the watermark description word is obtained by: Performing text parsing on the first text in the search knowledge base to obtain at least one triple corresponding to the first text, the triple comprising entity descriptors and relationship descriptors, the relationship descriptors representing entity relationships between entity descriptors in the triple; in, The entity description words constitute an entity list, the relationship description words constitute a relationship list, and the entity description words and the relationship description words are both initial description words contained in the first text; Using the signature information, multiple watermark description words are screened out from the initial description words.

7. The method according to claim 6, using the signature information to select a plurality of watermark description words from the initial description words, comprising: Using the signature information, screening a plurality of first descriptive words in the entity list; Filtering a plurality of second descriptive words in the relationship list according to the first descriptive word and the signature information; The first description word and the second description word are watermark description words.

8. The method according to claim 7, using the signature information to filter a plurality of first descriptive words in the entity list, comprising: Selecting an entity descriptor in the entity list as a current descriptor; Determining the current description word as a first description word; Performing a hash calculation on the current description word using the signature information to obtain a first index position, where the first index position points to a target description word in the entity list; The target description word is used as a new current description word, and the step of determining the current description word as a first description word is performed until the number of the first description words meets a first screening condition.

9. The method according to claim 7, screening a plurality of second descriptive words in the relationship list according to the first descriptive word and the signature information, comprises: Selecting two entity candidate words from the plurality of first description words; Performing hash calculation on the two entity candidate words using the signature information to obtain a second index position, where the second index position points to the target description word in the relationship list; The target descriptor is determined as a second descriptor, and the number of the second descriptors satisfies a second screening condition.

10. A data processing device, comprising: A request obtaining unit, used for obtaining a data query request; The data query request includes a query keyword; A text retrieval unit, used to perform text retrieval in a retrieval knowledge base according to the query keyword to obtain a retrieval text; The search text includes at least one watermark information, the watermark information is obtained based on a watermark description word included in the search knowledge base, and the watermark description word is related to the signature information corresponding to the search knowledge base.