Cross-domain privacy retrieval method and system based on embedding interaction and re-embedding check
By converting documents into embedded vectors for retrieval and verification, the problems of high data verification complexity and privacy protection in cross-domain data retrieval are solved, and efficient and reliable data consistency verification and privacy protection are achieved.
Patent Information
- Application Number
- CN202510677290.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-16
AI Technical Summary
Existing cross-domain data retrieval methods lack an effective data verification mechanism, resulting in a low degree of matching between index information and target data. In addition, existing verification methods have high computational complexity and high storage costs, making it difficult to provide efficient and reliable data consistency verification while protecting data privacy.
Using the method of embedding interaction and re-embedding checking, the document is converted into an embedding vector and stored in a vector database. Retrieval and verification are performed through the embedding vector. The requesting party re-embeds and compares the target document to determine its authenticity, ensuring data privacy protection and consistency.
It achieves efficient verification of the authenticity and consistency of target data while protecting data privacy, reduces computational complexity and storage costs, and improves the reliability and efficiency of cross-domain data interaction.
Smart Images

Figure CN120653822A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data security technology, and in particular to a cross-domain privacy retrieval method and system based on embedding interaction and re-embedding checking. Background Art
[0002] Traditional retrieval systems typically process cross-domain data by directly exchanging raw data. This simple and direct approach carries two security risks: first, the transmission of original documents by data providers can lead to the leakage of intellectual property and trade secrets; second, this model lacks effective privacy protection and data verification mechanisms, making it impossible to ensure data integrity during transmission and difficult to verify that the data retrieved by the query is consistent with the indexed content. This severely limits the application of cross-domain data retrieval in highly sensitive scenarios.
[0003] Existing cross-domain retrieval systems generally lack effective data verification mechanisms, making it difficult to ensure the authenticity and integrity of retrieval results. Data providers may employ a "summary inducement-content replacement" strategy, displaying high-quality summaries to attract users during the retrieval phase, but then providing low-quality or irrelevant content when actually delivering it. They may even return conflicting results for the same query at different times, making data interaction unpredictable. Furthermore, due to the lack of effective verification methods, data providers struggle to prove to querying parties that the data they provide is authentic and reliable, further reducing the efficiency and value of cross-domain data interaction. In an era where data is an asset, the lack of a data verification mechanism has become a key obstacle to the secure circulation and value mining of cross-domain data.
[0004] Currently, the industry has proposed a variety of solutions to address cross-domain data verification, but these solutions all have significant limitations. Hash value comparison-based verification is the most common data verification technique. It verifies data integrity by calculating the hash value of the original data and comparing it with the declared hash value. However, this method is susceptible to hash collisions, where different data may generate the same hash value, rendering verification invalid. Furthermore, it only performs byte-by-byte comparisons, cannot support semantic verification, and cannot tolerate even minor data changes. Blockchain-based data ownership confirmation mechanisms leverage the blockchain's immutable nature to record data summaries or fingerprints on-chain, achieving trusted data authentication. While this method theoretically provides highly reliable verification, its high computational complexity, high storage costs, long transaction confirmation times, and difficulty in handling large-scale data verification scenarios severely limit its practicality. Verification methods based on third-party notarization rely on a trusted third-party organization to verify and notarize the data, simplifying technical implementation. However, this approach introduces additional trust, increasing system complexity and operational costs while also introducing new security risks. In practical applications, it is difficult for the above methods to provide efficient and reliable data consistency verification while protecting data privacy.
[0005] Therefore, existing cross-domain data retrieval methods have the following technical problems: 1. The index information of the target data does not match the target data well, which results in the target data not matching the query party's retrieval purpose.
[0006] 2. In order to ensure that the target data returned is the retrieved data, data verification mechanisms such as hash value comparison, electronic notarization, and blockchain notarization are required to confirm that the target data has not been tampered with. The data verification mechanism will increase the complexity of the communication network and the computing power of the computer hardware.
[0007] In order to solve the above-mentioned defects in the prior art, the present application provides a cross-domain privacy retrieval method and system based on embedding interaction and re-embedding checking. Summary of the Invention
[0008] To overcome the problems existing in the related art, the first aspect of this application provides a cross-domain privacy retrieval method based on embedding interaction and re-embedding check, comprising: The responder converts the first document into a first embedding vector and stores the first embedding vector in a vector database. The requesting party converts the query request into a second embedding vector, and sends the second embedding vector to the responding party; The responder queries the matching first embedding vector according to the second embedding vector, and sends the matching first embedding vector to the requester; The requester sends a target document request to the responder, and obtains the target first document returned by the responder; The requesting party determines the authenticity of the target first document based on a comparison result between the target first document and the first embedding vector.
[0009] In one embodiment, the responding party converts the first document into a first embedding vector and stores the first embedding vector in a vector database, specifically including: obtaining the first document; Parsing the data type of the first document; the data type includes a character data type, a numeric data type, a table data type, and a JSON data type; A corresponding encoder is called according to the data type to encode the first document to obtain a second document; the second document is a text sequence.
[0010] In one embodiment, after calling a corresponding encoder according to the data type to encode the first document to obtain a second document, the method further includes: Segmenting the second document using a sliding window to obtain N consecutive text blocks, where N is an integer greater than or equal to 1; The N text blocks of the second document are converted into first embedding vectors and stored in the vector database.
[0011] In one implementation, the sliding step length of the sliding window is smaller than the left and right boundary lengths of the sliding window.
[0012] In one embodiment, converting the N text blocks of the second document into first embedding vectors and storing the first embedding vectors in the vector database further includes: Generate globally unique hash values as identifiers for the N text blocks.
[0013] In one embodiment, the responding party searches for a matching first embedding vector based on the second embedding vector and sends the matching first embedding vector to the requesting party, specifically including: Obtaining the second embedding vector in the query request; Performing a search based on the semantic similarity between the first embedding vector and the second embedding vector to obtain a matching target first document; The index text of the target first document is sent to the requesting party.
[0014] In one embodiment, the requesting party determines the authenticity of the target first document based on a comparison result between the target first document and the first embedding vector, specifically including: The requesting party converts the target first document into a target embedding vector; The authenticity of the target first document is determined according to a similarity between the target embedding vector and the first embedding vector.
[0015] In one embodiment, before the requesting party converts the target first document into a target embedding vector, the method further includes: The requester calls the key interface of the target first document of the responder to obtain the key corresponding to the target first document; The requesting party calls the key corresponding to the target first document to request the return of the target first document.
[0016] In one embodiment, determining the authenticity of the target first document based on the similarity between the target embedding vector and the first embedding vector specifically includes: The authenticity of the target first document is determined based on the cosine similarity between the target embedding vector and the first embedding vector; the cosine similarity calculation formula is:
[0017] in, represents the first embedding vector, represents the target embedding vector.
[0018] The second aspect of this application provides a cross-domain privacy retrieval system based on embedding interaction and re-embedding checking, which implements cross-domain privacy retrieval based on the steps in the cross-domain privacy retrieval method described in the first aspect of this application.
[0019] The technical solution provided by this application may have the following beneficial effects: To protect the responder's data privacy, this application embeds the requester's query and the first document, then compares the resulting first and second embedding vectors for retrieval. The requester then checks the index information returned by the responder after the retrieval to determine the target document they are requesting. The requester can only view the index information and cannot access the original first document, effectively protecting the responder's data privacy.
[0020] To prevent the responder from returning documents unrelated to the requested first document, the requester in this application requests the return of the retrieved target first document. The responder then sends the first embedding vector of the target first document and the original document to the requester. The requester then re-embeds the original document, comparing the local target embedding vector with the returned first embedding vector to confirm that the original document sent by the responder is the retrieved target first document.
[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other objects, features and advantages of the present application will become more apparent through a more detailed description of exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0023] Figure 1 This is a flowchart of a cross-domain privacy retrieval method based on embedding interaction and re-embedding checking shown in an embodiment of the present application; Figure 2 This is a first network interaction diagram of the cross-domain privacy retrieval method shown in an embodiment of the present application; Figure 3 A second network interaction diagram of the cross-domain privacy retrieval method shown in an embodiment of the present application; Figure 4 This is a third network interaction diagram of the cross-domain privacy retrieval method shown in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The preferred embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0025] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0026] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0027] Example 1 In cross-domain data retrieval, data holders only return index information of the original data to users to protect privacy. Users search based on the index information, select the target data, and send a return request. Due to data privacy concerns, it is difficult to establish a reliable and simple data verification mechanism between data holders and users.
[0028] In order to solve the privacy data protection and data verification problems in the above-mentioned cross-domain retrieval field, an embodiment of the present application provides a cross-domain privacy retrieval method based on embedding interaction and re-embedding check, which can perform target data verification between the data party and the query party, and can ensure the consistency between the index data and the target data.
[0029] Figure 1 A flowchart of the cross-domain privacy retrieval method provided in an embodiment of the present application.
[0030] The following combination Figure 1 Specifically explain the data interaction principle of this cross-domain privacy retrieval method.
[0031] like Figure 1 As shown, the cross-domain privacy retrieval method provided in the embodiment of the present application includes the following steps: S1. The responder converts the first document into a first embedding vector and stores it in a vector database. S2. The requester converts the query request into a second embedding vector and sends the second embedding vector to the responder; S3. The responder queries for the matching first embedding vector based on the second embedding vector, and sends the matching first embedding vector to the requester; S4. The requester sends a target document request to the responder, and obtains the target first document returned by the responder; S5. The requesting party determines the authenticity of the target first document based on a comparison result between the target first document and the first embedding vector.
[0032] In step S1, the text embedding model performs EMBEDDING on all first documents in the document database to convert them into embedding vectors, and stores all first embedding vectors in the vector database.
[0033] It will be appreciated that the network topology of the embodiment of the present application consists of a responder and a requester. The requester sends a query request to the responder, which searches the database for matching data and returns the data to the requester. The responder is provided with a vector database and a document database. The vector database stores the first embedding vector, and the document database stores the first document.
[0034] Exemplarily, the text embedding model can be a pre-trained embedding model such as bge-m3 or jina-embeddings-v3.
[0035] In step S2, the user terminal converts the query request into an embedding vector to obtain a second embedding vector. The second embedding vector is embedded using the text embedding model used by the responder in step S1. The text embedding models of S1 and S2 have the same structure, ensuring the reliability of the search results when comparing the first embedding vector and the second embedding vector.
[0036] In step S3, a mathematical comparison is performed within the vector database to find the first embedding vector that has the highest similarity to the second embedding vector. The responder returns the first embedding vector, along with the document ID and description, to the requester. During this process, only the vector data is transmitted over the network, protecting data privacy.
[0037] In step S4, after the user confirms the retrieval result, the corresponding target first document is determined by the selected first embedding vector, and the responder is again requested to return the original file of the target first document.
[0038] After the requesting party receives the retrieved target first document, it performs embedding on the target first document again, and then compares the mathematical similarity between the target embedding vector of the target first document and the first embedding vector sent by the responding party to determine that the target first document is the user's retrieval target file.
[0039] To protect the responder's data privacy, this embodiment embeds the requester's query request and the first document, then compares the resulting first and second embedding vectors for retrieval. The requester then checks the index information returned by the responder after the retrieval to determine the target document it needs. The requester can only view the index information and cannot access the original first document, effectively protecting the responder's data privacy.
[0040] To prevent the responder from returning documents unrelated to the requested first document, in this embodiment of the present application, the requester requests the return of the retrieved target first document. The responder then sends the first embedding vector of the target first document and the original document to the requester. The requester then re-embeds the original document, comparing the local target embedding vector with the returned first embedding vector to confirm that the original document sent by the responder is the retrieved target first document.
[0041] Example 2 Step S1 in the cross-domain privacy retrieval method shown in Example 1 is the embedding step for the first document. The first document may include multiple types of data, such as tables, JSON data, and mathematical formulas. In the embedding process, the embedding model needs to extract the features of each modality separately, fuse them, and map them into a unified vector space.
[0042] Existing embedding models consume a lot of time and computing power to parse long multimodal documents, such as academic papers and legal contracts.
[0043] In order to speed up the embedding efficiency of long multimodal documents, the embodiment of the present application also provides a cross-domain privacy retrieval method based on embedding interaction and re-embedding check, including steps S1-S5 of the method shown in Example 1.
[0044] Furthermore, S1 in the embodiment of the present application specifically includes: S101, obtaining the first document; S102: Parse the data type of the first document; the data type includes a character data type, a numeric data type, a table data type, and a JSON data type; S103: Call a corresponding encoder according to the data type to encode the first document to obtain a second document; the second document is a text sequence; S104: Segment the second document using a sliding window to obtain N consecutive text blocks; S105: Convert the N text blocks of the second document into first embedding vectors and store them in the vector database.
[0045] Wherein, N is an integer greater than or equal to 1; Specifically, the sliding step of the sliding window is smaller than the distance between the left and right boundaries of the sliding window. It can be understood that the left boundary is the starting position of the sliding window and the right boundary is the ending position of the sliding window.
[0046] In step S104, a sliding window of a fixed size (eg, 800 characters) is set, and each time the window is moved, only a portion of the window size (eg, 30% of the window size) is advanced, so that there is a certain percentage of content overlap between adjacent text blocks.
[0047] In the embodiment of the present application, the sliding strategy of the above sliding window can ensure that key semantic information is not truncated.
[0048] For extremely long multimodal documents, if fixed length segmentation is simply adopted, key semantic information will be truncated or split. Therefore, the sliding window technology is used to achieve overlapping block segmentation. Specifically, a fixed-size window (such as 800 words) is set, and each time the window is moved, only a part of the window size (such as 30% of the window size) is advanced, so that there is a certain proportion of content overlap between adjacent text blocks. This method can ensure that key semantic information is not truncated. In order to solve the embedding problem of multimodal long documents, the embodiment of the present application needs to preprocess the multimodal data before embedding to ensure that the embedding model can embed normally.
[0049] In step S103, different parsers are called to preprocess the data of different data types. For example, in the case of tabular data, the Python pandas library can be used to read the table content and convert each row or column into a text sequence. For JSON data, the JSON parsing library is used to parse it into text content in the form of key-value pairs.
[0050] After the conversion of the structured data into a text sequence is completed, the text block segmentation in step S104 is performed.
[0051] Furthermore, step S105 also includes: generating a globally unique hash value as an identifier for the N text blocks.
[0052] This embodiment of the present application generates a globally unique hash value as an identifier (ID) for each text block, ensuring that text block IDs across different documents are unique. The ID generation process utilizes a cryptographically secure hash algorithm (such as SHA-256) and establishes a strong binding between the ID and the text block content to prevent malicious tampering. Furthermore, a bidirectional mapping index table is created to record the original document metadata (including document name, creation time, domain, etc.) corresponding to each text block ID, as well as its specific location within the document.
[0053] Example 3 Step S5 in the cross-domain privacy retrieval method shown in the first embodiment is a data verification method of the requesting party, that is, checking whether the obtained data and the retrieved data are consistent.
[0054] In order to perform data verification, the embodiment of the present application also provides a cross-domain privacy retrieval method based on embodiment one, including steps S1-S5 of the method shown in embodiment one.
[0055] Wherein, step S5 is: the requesting party determines the authenticity of the target first document according to the comparison result between the target first document and the first embedding vector.
[0056] Further, such as Figure 4 As shown, step S5 specifically includes: S501: The requester calls the key interface of the target first document of the responder to obtain the key corresponding to the target first document; S502: The requesting party calls the key corresponding to the target first document to request the return of the target first document; S503: The requesting party converts the target first document into a target embedding vector; S504: Determine the authenticity of the target first document according to the similarity between the target embedding vector and the first embedding vector.
[0057] Specifically, in step S504, the requesting party compares the recalculated embedding vector with the embedding vector obtained from the responding party during the previous search. The comparison is calculated using cosine similarity, which is as follows:
[0058] in, represents the first embedding vector, represents the target embedding vector.
[0059] Exemplarily, step S504 sets a similarity threshold (usually 0.99) for determining data consistency. If the calculated result is greater than or equal to the similarity threshold, it can be confirmed that the target first document obtained by the requesting party is the search result of its search.
[0060] It should be noted that when performing embedding, the requester uses the same version of the embedding model as the responder to recalculate the embedding vector for the obtained original data.
[0061] If the similarity calculated on the requesting end exceeds a preset threshold, verification passes, confirming that the retrieved original data is consistent with the content indexed during retrieval. At this point, the original data can be safely presented to the user or made available to downstream applications. If the similarity falls below the threshold, a warning is issued, indicating that the data may have been tampered with or that the index does not match the original data, and the data is denied use.
[0062] In step S501, when the requester determines that it needs to obtain the original data of a particular record, it calls the responder's original data key interface (data_key) based on the ID of the target first document to obtain the corresponding original data key. It then calls the original data interface (data) to obtain the original content of the target first document. This two-stage design increases access control flexibility and decouples query and storage services.
[0063] In step S502, the requester sends a raw data key interface request, and the responder verifies the request and returns the access key for the data. The requester then uses the obtained key to send a raw data interface request, and the responder verifies the validity of the key and returns the raw data content and related metadata.
[0064] In the query request and retrieval result return steps of step S2 and step S3, as Figure 2 As shown, specifically including: S201. The requester calls the knowledge base information interface of the responder to obtain the basic information of the responder; Specifically, the requester sends a request containing the knowledge base name and API key, and the knowledge base information interface provides metadata of the knowledge base, including name, description, data size, supported query filter fields, and available embedding models.
[0065] S202: The requester converts the query into an embedding vector using the embedding model supported by the responder.
[0066] This process needs to ensure that the embedding model used by the requester is consistent with the embedding model used by the responder, in preparation for subsequent re-embedding checks.
[0067] It is understandable that the information of the embedded model is obtained by calling the knowledge base information interface in S201. The interactive process of the search interface is as follows Figure 2 As shown, the requester sends the query embedding vector to the responder and declares the necessary filtering conditions.
[0068] In step S3, if Figure 3 As shown, specifically including: S301: After receiving the search request, the responder verifies the validity of the search request; S302: The responder performs a similarity search in the vector database to find documents related to the query. S303: The responder returns the search results.
[0069] Specifically, the responder's search interface performs a semantic similarity search based on the second embedding vector and returns a set of documents related to the responder's request. The returned search results include the document ID, the first embedding vector, and description information.
[0070] Example 4 A cross-domain privacy retrieval system based on embedding interaction and re-embedding checking implements cross-domain privacy retrieval based on the steps of the cross-domain privacy retrieval method described in any one of embodiments one to three.
[0071] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0072] The solution of the present application has been described in detail above with reference to the accompanying drawings. In the above embodiments, the description of each embodiment has its own focus. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. Those skilled in the art should also be aware that the actions and modules mentioned in the description are not necessarily required for this application.
[0073] In addition, it can be understood that the steps in the method of the embodiment of the present application can be adjusted in order, merged and deleted according to actual needs, and the modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0074] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0075] Alternatively, the present application can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or electronic device, server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.
[0076] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the application herein may be implemented as electronic hardware, computer software, or combinations of both.
[0077] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems and methods according to multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0078] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A cross-domain privacy retrieval method based on embedding interaction and re-embedding check, characterized by: include: The responder converts the first document into a first embedding vector and stores the first embedding vector in a vector database. The requesting party converts the query request into a second embedding vector, and sends the second embedding vector to the responding party; The responder queries the matching first embedding vector according to the second embedding vector, and sends the matching first embedding vector to the requester; The requester sends a target document request to the responder, and obtains the target first document returned by the responder; The requesting party determines the authenticity of the target first document based on a comparison result between the target first document and the first embedding vector.
2. A cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 1, characterized in that: The responder converts the first document into a first embedding vector and stores the first embedding vector in a vector database, specifically including: obtaining the first document; Parsing the data type of the first document; the data type includes a character data type, a numeric data type, a table data type, and a JSON data type; A corresponding encoder is called according to the data type to encode the first document to obtain a second document; the second document is a text sequence.
3. A cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 2, characterized in that: After encoding the first document by calling a corresponding encoder according to the data type to obtain a second document, the method further includes: Segmenting the second document using a sliding window to obtain N consecutive text blocks, where N is an integer greater than or equal to 1; The N text blocks of the second document are converted into first embedding vectors and stored in the vector database.
4. A cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 3, characterized in that: include: The sliding step size of the sliding window is smaller than the lengths of the left and right boundaries of the sliding window.
5. A cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 4, characterized in that: Converting the N text blocks of the second document into first embedding vectors and storing the first embedding vectors in the vector database further includes: Generate globally unique hash values as identifiers for the N text blocks.
6. The cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 1 is characterized in that: The responder searches for a matching first embedding vector based on the second embedding vector, and sends the matching first embedding vector to the requester, specifically including: Obtaining the second embedding vector in the query request; Performing a search based on the semantic similarity between the first embedding vector and the second embedding vector to obtain a matching target first document; The index text of the target first document is sent to the requesting party.
7. The cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 1 is characterized in that: The requesting party determines the authenticity of the target first document based on a comparison result between the target first document and the first embedding vector, specifically including: The requesting party converts the target first document into a target embedding vector; The authenticity of the target first document is determined according to a similarity between the target embedding vector and the first embedding vector.
8. The cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 7 is characterized in that: Before the requesting party converts the target first document into a target embedding vector, the method further includes: The requester calls the key interface of the target first document of the responder to obtain the key corresponding to the target first document; The requesting party calls the key corresponding to the target first document to request the return of the target first document.
9. The cross-domain privacy retrieval method based on embedding interaction and re-embedding check according to claim 7 is characterized in that: Determining the authenticity of the target first document according to the similarity between the target embedding vector and the first embedding vector specifically includes: The authenticity of the target first document is determined based on the cosine similarity between the target embedding vector and the first embedding vector; the cosine similarity calculation formula is: in, represents the first embedding vector, represents the target embedding vector.
10. A cross-domain privacy retrieval system based on embedding interaction and re-embedding checking, characterized in that: Cross-domain privacy retrieval is implemented based on the steps in the cross-domain privacy retrieval method described in any one of claims 1 to 9.