Sample processing method and device and electronic equipment

By leveraging trusted third-party assistance and covert query technology, the privacy protection issue in aligning intersecting samples from multiple sources is resolved. This enables efficient sample alignment and modeling, ensuring that intersecting sample information is not leaked and improving data flow and modeling efficiency.

CN121598084APending Publication Date: 2026-03-03CHINA MOBILE GROUP JIANGSU +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511773245.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently align samples from multiple sources while protecting data privacy. In particular, in scenarios where the information of intersecting samples is highly sensitive, traditional methods are prone to leaking intersecting sample information or affecting modeling effectiveness and efficiency.

Method used

By introducing trusted third-party assistance, a list of obfuscated sample line numbers is generated through covert query technology and unintentional scrambling operations. The sample splicing and arrangement are then performed using a ciphertext calculation engine to ensure that the information of the intersecting samples is not exposed.

Benefits of technology

It achieves efficient sample alignment without revealing the identifiers of the intersection samples, ensuring the joint modeling effect of the model on the intersection data and improving privacy protection and data flow capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598084A_ABST
    Figure CN121598084A_ABST
Patent Text Reader

Abstract

The invention provides a sample processing method and device and electronic equipment, and belongs to the technical field of data processing. The sample processing method is applied to a trusted third party, and comprises the following steps: acquiring a first line number list of an intersection sample in a first participant and a second line number list of the intersection sample in a second participant; confusion sample line numbers are added into the first line number list and the second line number list respectively, the line numbers in the first line number list added with the confusion sample line numbers are disrupted according to disruption index information, and a third line number list is obtained; disorganizing the line numbers in the second line number list added with the mixed sample line numbers to obtain a fourth line number list; the third line number list is sent to the first participant, and the fourth line number list is sent to the second participant; and generating anti-scrambling index information according to the scrambling index information, and sending the encrypted anti-scrambling index information and the number of the intersection samples to a ciphertext calculation engine. According to the invention, sample alignment can be efficiently realized on the premise of protecting the sample identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a sample processing method, apparatus and electronic device. Background Technology

[0002] In recent years, with the widespread use of mobile internet, massive amounts of user data have been collected. After being processed and analyzed through information technology, this data can be utilized by governments, enterprises, organizations, and institutions to provide users with better services and products, thus realizing the value of data. However, data reflecting different aspects of user characteristics is often collected by different institutions; for example, banks hold users' consumption data, and operators hold users' communication data, meaning the data often exists in silos. In certain business scenarios, it is necessary to introduce external data sources to empower business operations. Traditional multi-source data processing often simply involves transmitting data stored in different parties to a single processing platform for centralized processing. This approach cannot control the flow of data and is prone to privacy leaks. With the frequent occurrence of information leaks, increased user awareness of privacy protection, and the improvement of relevant laws and regulations, higher demands are being placed on the protection of privacy data. Against this backdrop, how to break down data silos and realize the value of data while protecting data privacy has become a research hotspot in the current information science academic and industrial communities. Summary of the Invention

[0003] This invention provides a sample processing method, apparatus, and electronic device that can efficiently achieve sample alignment while protecting all sample identifiers, including intersection sample identifiers, for subsequent model training or prediction.

[0004] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0005] In a first aspect, embodiments of the present invention provide a sample processing method applied to a trusted third party, comprising:

[0006] Obtain the first row number list of the intersection sample of the first participant and the second participant in the first participant and the second row number list in the second participant;

[0007] Add obfuscated sample row numbers to the first row number list and the second row number list respectively. Then, shuffle the row numbers in the first row number list after adding the obfuscated sample row numbers according to the shuffle index information to obtain the third row number list. Finally, shuffle the row numbers in the second row number list after adding the obfuscated sample row numbers according to the shuffle index information to obtain the fourth row number list.

[0008] Send the third row number list to the first participant, and send the fourth row number list to the second participant;

[0009] Based on the scrambled index information, anti-scrambled index information is generated. The anti-scrambled index information is used to restore the scrambled row number list to the row number list before scrambling. The encrypted anti-scrambled index information and the number of intersection samples are sent to the ciphertext calculation engine.

[0010] Secondly, embodiments of the present invention provide a sample processing method applied to a ciphertext calculation engine, comprising:

[0011] Receive encrypted descrambled index information sent by a trusted third party, as well as the number of intersection samples between the first and second participants. ;

[0012] The encrypted first feature data sent by the first participant and the encrypted second feature data sent by the second participant are obtained. The first feature data is the feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party. The second feature data is the feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party.

[0013] The first feature data and the second feature data are concatenated to obtain ciphertext feature data. Based on the anti-scrambling index information, the ciphertext feature data is rearranged by calling the unintentional ciphertext scrambling operation to obtain a ciphertext array.

[0014] Extract the previous text from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

[0015] Thirdly, embodiments of the present invention provide a sample processing apparatus, applied to a trusted third party, comprising:

[0016] The acquisition module is used to acquire the first row number list of the intersection samples of the first participant and the second participant in the first participant and the second row number list in the second participant;

[0017] The scrambling module is used to add obfuscated sample row numbers to the first row number list and the second row number list respectively, and scramble the row numbers in the first row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the third row number list; and scramble the row numbers in the second row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the fourth row number list.

[0018] The first sending module is used to send the third row number list to the first participant and the fourth row number list to the second participant;

[0019] The second sending module is used to generate descrambled index information based on the scrambled index information. The descrambled index information is used to restore the scrambled row number list to the row number list before scrambling. The encrypted descrambled index information and the number of intersection samples are sent to the ciphertext calculation engine.

[0020] Fourthly, embodiments of the present invention provide a sample processing apparatus applied to a ciphertext calculation engine, comprising:

[0021] The receiving module is used to receive encrypted descrambled index information sent by a trusted third party, as well as the number of intersection samples of the first and second participants. ;

[0022] The acquisition module is used to acquire encrypted first feature data sent by the first participant and encrypted second feature data sent by the second participant. The first feature data is feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party, and the second feature data is feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party.

[0023] The splicing module is used to splice the first feature data and the second feature data to obtain ciphertext feature data, and rearrange the ciphertext feature data by calling the unintentional ciphertext scrambling operation according to the anti-scrambling index information to obtain a ciphertext array;

[0024] Extraction module, used to extract the first part of the ciphertext from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

[0025] Fifthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the sample processing method as described in the first or second aspect above.

[0026] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the sample processing method as described in the first or second aspect above.

[0027] In a seventh aspect, embodiments of the present invention provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the sample processing method as described in the first or second aspect above.

[0028] In this embodiment of the invention, a trusted third party obtains the intersection sample of the first participant and the second participant, specifically the first row number list in the first participant and the second row number list in the second participant. Obfuscated sample row numbers are added to both the first and second row number lists. The row numbers in the first row number list after adding the obfuscated sample row numbers are then shuffled according to shuffling index information to obtain a third row number list. Similarly, the row numbers in the second row number list after adding the obfuscated sample row numbers are shuffled according to shuffling index information to obtain a fourth row number list. The third row number list is sent to the first participant, and the fourth row number list is sent to the second participant. De-shuffling index information is generated based on the shuffling index information, and this de-shuffling index information is used to restore the shuffled row number list to its original shuffled state. The first party sends the encrypted descrambling index information and the number of intersection samples to the ciphertext calculation engine, which then receives the encrypted first feature data from the first party and the encrypted second feature data from the second party. The first feature data is extracted from the third row number list received by the first party from the trusted third party, and the second feature data is extracted from the fourth row number list received by the second party from the trusted third party. The first and second feature data are concatenated to obtain the ciphertext feature data. Based on the descrambling index information, the ciphertext feature data is rearranged by calling an unintentional ciphertext scrambling operation to obtain the ciphertext array. The first feature data is then extracted from the ciphertext array. The encrypted feature data of the intersection samples is obtained through this method. The technical solution of this embodiment, with the assistance of a trusted third party, can obtain the encrypted feature data of the intersection samples without exposing the identifiers of the intersection samples to the two participating parties. It can efficiently achieve sample alignment while protecting all sample identifiers, including the intersection sample identifier, for subsequent model training or prediction. This ensures the effectiveness of joint modeling on the intersection data, achieves higher privacy protection, facilitates data flow, and maximizes data value. Attached Figure Description

[0029] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0030] Figure 1 This is a flowchart illustrating a sample processing method applied to a trusted third party according to an embodiment of the present invention.

[0031] Figure 2 This is a flowchart illustrating the sample processing method applied to the encrypted computing engine according to an embodiment of the present invention.

[0032] Figure 3This is a schematic diagram illustrating how participant A and participant B, in an embodiment of the present invention, use PIR and a trusted third party to align their intersecting samples.

[0033] Figure 4 This is a schematic diagram illustrating how a trusted third party obtains a list of row numbers for the intersection of samples between participant A and participant B through query results, adds obfuscated sample row numbers to the list, shuffles the list of row numbers containing obfuscated sample row numbers, and sends it to the two participants respectively.

[0034] Figure 5 In this embodiment of the invention, the encrypted and obfuscated samples from the two participants are concatenated, and then the ciphertext is used for unintentional shuffling and sorting to obtain the sorted array. A schematic diagram of a ciphertext sample;

[0035] Figure 6 This is a structural block diagram of a sample processing device applied to a trusted third party according to an embodiment of the present invention;

[0036] Figure 7 This is a structural block diagram of a sample processing device applied to a ciphertext computing engine according to an embodiment of the present invention;

[0037] Figure 8 This is a schematic diagram of the composition of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] The emergence of technologies such as secure multi-party computation and federated learning provides a paradigm for solving the problem of "how to break down data silos and realize the value of data while protecting data privacy." Secure multi-party computation refers to different participants jointly completing a computational task without disclosing their own private data, using techniques such as secret sharing, unintentional transmission, and obfuscated circuits. Federated learning is a distributed machine learning framework that protects privacy data. Most computations are performed locally in plaintext, with a small amount of encrypted computation used to train the model while keeping privacy data confidential.

[0040] Whether it's multi-party secure computation or federated learning, common multi-source data joint modeling schemes require prior sample alignment (Private Set Interaction, PSI) among all participants. Sample alignment refers to determining the user intersection by encrypting sample identifiers (IDs) without exposing user IDs, so as to jointly model the features of the intersection users. Common sample alignment schemes include those based on symmetric encryption protocols, oblivious transfers (OT), and oblivious pseudorandom functions (OPRF), etc. Ultimately, all participants will obtain the plaintext result of the intersection IDs. It is generally assumed that these intersection user IDs are data shared by all participants, thus not leading to privacy leaks. However, in some scenarios, participants may need to keep the intersection samples confidential. The intersection samples expose some of a company's users, and if the intersection samples are leaked, other companies may poach users, leading to the risk of user churn for the original company. Therefore, the sample alignment scheme that protects the intersection is also a prerequisite for the joint modeling of the intersection, and it has extremely important practical significance for privacy protection, data circulation, and realizing the value of data.

[0041] Existing PSI technology cannot meet the requirements of joint modeling scenarios that protect intersection information. Some alignment sample schemes can protect the intersection information for a single participant while exposing it to the other. However, in scenarios where both parties are highly sensitive to the intersection sample information, this security approach is not suitable.

[0042] To protect the intersection sample information, existing solutions can also use hidden queries, where the active querying party can obtain the intersection index column of the queried party. This solution exposes the intersection information to the active querying party. This solution can only prevent the queried party from knowing the intersection information, and is not suitable for scenarios where both parties are highly sensitive to the intersection information.

[0043] To perform joint modeling without exposing the intersection sample information to all parties, one existing approach is to output the union data and then perform joint modeling on the union data to preserve the intersection. However, this approach of performing joint modeling on the union data will affect the modeling effect to some extent, resulting in a deviation from the modeling effect on the intersection samples. Furthermore, due to the introduction of calculations of sample data outside the intersection, the modeling efficiency is low.

[0044] Another approach to protect the joint modeling of intersecting samples is based on comparison operations in multi-party secure computation to ensure sample alignment. This approach involves both parties encrypting their sample IDs, sample features, and a randomly generated shuffle sequence of the same length as the other party's sample and sending them to the multi-party secure computation platform. The platform then performs an oblivious shuffle, using the encrypted shuffle sequence provided by one party to shuffle the other party's encrypted dataset. The sample order in the generated ciphertext data is obliviously shuffled, and no party can know this. The encrypted ID fields of the two parties' samples are extracted, and a pairwise ciphertext comparison is performed. After restoring the comparison results to plaintext, the row number correspondence of the intersecting sample IDs in each party's data can be obtained. Based on the row numbers, the intersecting sample features in each party's dataset are obtained. These features are then concatenated and used to train the model on the ciphertext. Because this implementation shuffles the dataset, it is impossible to map the plaintext row numbers to specific user IDs, thus protecting the intersecting sample IDs. However, due to the significant communication overhead of the comparison operation in multi-party secure computation, this scheme has very low operating efficiency.

[0045] To address the aforementioned technical issues, this embodiment proposes a privacy-preserving method for joint modeling of multi-source data (features and labels belonging to different data sources). It introduces trusted third-party assisted computation, improves existing hidden query algorithms, exposes only query results to trusted third parties, and designs a novel algorithm flow. This method enables joint modeling without exposing the intersection information of all data participants (including intersection sample IDs and corresponding feature data), and achieves the same joint modeling effect as joint modeling on intersection samples, resulting in higher operational efficiency.

[0046] Please refer to Figure 1 This invention provides a sample processing method applied to a trusted third party, comprising:

[0047] Step S11: Obtain the first row number list of the intersection sample of the first participant and the second participant in the first participant and the second row number list in the second participant;

[0048] In this embodiment, a sequence number is added in advance to the sample data of the second participant as a query value. The sequence number corresponds one-to-one with the sample of the second participant and is the row number of the corresponding sample in the sample dataset of the second participant.

[0049] In some embodiments, before obtaining the first row number list of the intersection sample of the first participant and the second participant in the first row number list in the second participant, the method further includes:

[0050] The system receives a query result returned by the second participant. The query result is sent by the second participant after receiving a query key sent by the first participant. The query key is the sample identifier of the first participant, which is used to indicate the row number of the sample identifier in the sample dataset of the second participant. The number of samples of the second participant is greater than the number of samples of the first participant. The query result includes the row number of the query key of the first participant in the sample dataset of the second participant.

[0051] After receiving the query results, the first row number list of the intersection samples in the first participant and the second row number list in the second participant can be obtained through the query results.

[0052] In some embodiments, if the query key of the first participant is not in the sample dataset of the second participant, the query result corresponding to the query key is a negative value.

[0053] Step S12: Add obfuscated sample row numbers to the first row number list and the second row number list respectively. Scramble the row numbers in the first row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the third row number list. Scramble the row numbers in the second row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the fourth row number list.

[0054] In some embodiments, the number of the confused sample row numbers can be determined based on the number of intersection samples and a set ratio. The value of the confused sample row number is determined by the number of samples of the first participant and the second participant.

[0055] In some embodiments, the number of the obfuscated sample row numbers can be determined according to the following formula:

[0056]

[0057] in, The number of the confused sample row numbers, The sample size of the first participant. For the intersection

[0058] The number of samples in the set.

[0059] Step S13: Send the third row number list to the first participant and send the fourth row number list to the second participant;

[0060] Step S14: Generate descrambling index information based on the scrambling index information. The descrambling index information is used to restore the scrambled row number list to the row number list before scrambling. Send the encrypted descrambling index information and the number of intersection samples to the ciphertext calculation engine.

[0061] The technical solution of this embodiment, with the assistance of a trusted third party, can obtain the encrypted feature data of the intersection sample without exposing the intersection sample identifier to the two participating parties. It can efficiently achieve sample alignment while protecting all sample identifiers, including the intersection sample identifier, for subsequent model training or prediction. This ensures the effect of joint modeling on the intersection data, achieves a higher level of privacy protection, facilitates data circulation, and maximizes the value of data.

[0062] Please refer to Figure 2 This invention provides a sample processing method applied to a ciphertext calculation engine, comprising:

[0063] Step 21: Receive the encrypted descrambling index information sent by a trusted third party, as well as the number of intersection samples of the first and second participants. ;

[0064] Step 22: Obtain the encrypted first feature data sent by the first participant and the encrypted second feature data sent by the second participant. The first feature data is the feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party. The second feature data is the feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party.

[0065] Step 23: Concatenate the first feature data and the second feature data to obtain ciphertext feature data. Based on the anti-scrambling index information, rearrange the ciphertext feature data by calling the unintentional ciphertext scrambling operation to obtain the ciphertext array.

[0066] Step 24: Extract the first part from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

[0067] The technical solution of this embodiment, with the assistance of a trusted third party, can obtain the encrypted feature data of the intersection sample without exposing the intersection sample identifier to the two participating parties. It can efficiently achieve sample alignment while protecting all sample identifiers, including the intersection sample identifier, for subsequent model training or prediction. This ensures the effect of joint modeling on the intersection data, achieves a higher level of privacy protection, facilitates data circulation, and maximizes the value of data.

[0068] The technical solution in this embodiment can address the data privacy and security requirements of bank customers, and is particularly applicable to application scenarios where features and tags belong to different data providers in multi-source data.

[0069] The technical solution of the present invention will be further described below with reference to specific embodiments and accompanying drawings:

[0070] Suppose there are two participants, participant A (equivalent to the first participant mentioned above) and participant B (equivalent to the second participant mentioned above), and a trusted auxiliary third party S. Each participant holds different feature information, and one of them also holds label information. Let's assume participant A holds the label information; in joint modeling, the party holding the labeled data is usually referred to as the active party. Participant A wants to introduce features from participant B, whose features differ from its own, to jointly train a machine learning model, thereby improving the model's predictive performance on new samples and facilitating its business. Due to business competition or concerns about customer privacy, both participants have extremely high privacy requirements and do not want to expose the IDs of their overlapping samples during joint modeling.

[0071] To avoid exposing the intersection sample IDs of all parties, this embodiment employs a private information retrieval (PIR) technique and a trusted third party to implement a sample alignment scheme. This ensures that neither party A nor party B can obtain specific intersection sample information. Simultaneously, even without the intersection sample information, this embodiment aligns the samples at the intersection of parties A and B. Furthermore, through optimization—unintentional shuffling—this aligned intersection data can be used for subsequent joint modeling without affecting the modeling results, unlike using union data. Which intersection data participates in joint modeling is unknown to all parties, thus meeting the application requirements of specific scenarios highly sensitive to intersection information.

[0072] like Figure 3 As shown, suppose participant A's sample IDs are [abc0, abc1, abc2, abc3, abc4, abc5], and the corresponding sample features have two attributes, V1 and V2, denoted as... The sample size is Participant B's sample IDs are [abc4, abc5, abc0, abc6, abc7, abc8, abc9], and the corresponding sample features have two attributes, V3 and V4, denoted as... The sample size is The technical solution of this embodiment specifically includes the following steps:

[0073] Step 1: Participant A and Participant B use PIR and trusted third-party assistance to align their intersecting samples to protect the data.

[0074] The PIR technology used in this embodiment is a technology that can securely transmit query results to a trusted third party. Specifically, participant A uses an ID to query participant B's data, and the final query result is transmitted to a trusted third party for decryption. During this process, neither participant A nor participant B can obtain the query result; only the trusted third party can.

[0075] The party with more data among the two participants, acting as the query target in the hidden query PIR, adds a serial number column to its own sample data, using the sample ID column and the serial number column as the query key and query value, respectively, as the query objects for subsequent hidden queries.

[0076] In this embodiment, participant B has a larger amount of data. Participant B adds a sequence number to its sample data, that is, numbers from 0 to 1 according to the order of the sample rows. -1 The samples are numbered, and the sample ID is used as the query key, while the sequence number is used as the query value for subsequent hidden queries. Specifically, participant B has 7 samples, which are added to a sequence number column [0, 1, 2, 3, 4, 5, 6].

[0077] The party with less data among the two participants initiates a hidden query (PIR) to the other party, using its own sample ID as the query key, to find the row number of the sample ID in the other participant's dataset.

[0078] In this embodiment, participant A uses its own sample ID [abc0, abc1, abc2, abc3, abc4, abc5] as the query key to perform a hidden query to participant B. The hidden query can return the query value corresponding to the query key without the queried party knowing the query key of the querying party. Therefore, participant B cannot know the sample ID of participant A in this process.

[0079] Participant B securely decrypts the query results to a trusted third party S. The trusted third party obtains the correspondence between the intersection samples in the two participants based on the query results, adds the row number of the obfuscated samples to generate a pseudo intersection sample sequence number, and sends it to the two participants respectively.

[0080] The query result is the query value corresponding to the query key of the querying party in the queryed party's query key, i.e., the row number of the intersection sample ID. If the query key of the querying party is not in the query key of the queryed party, the query result returns -1. Therefore, the trusted third party S receives an integer array of the same length as the querying party, i.e., the number of samples in participant A. The row of the integers in the array corresponds to the sample ID of participant A, and the integer value corresponds to the row number of the sample ID in participant B. For example, in this example, the array received by the trusted third party is [2, -1, -1, -1, 0, 1]. The length of the array is 6, which is the same as the number of samples in participant A. The first integer 2 indicates that the first sample ID of participant A is in the row number 2 of participant B, meaning that the third sample ID is the same. The second integer -1 indicates that the second sample ID of participant A does not exist in the sample IDs of participant B. Therefore, the trusted third party can know the correspondence between the intersection samples in the two participants and the number of intersection samples through the received array information, but cannot know the specific intersection sample IDs.

[0081] Step 2: The trusted third party obtains the row number mapping relationship of the intersection sample of participant A and participant B through the query results, adds the obfuscated sample row number to the row number list, shuffles the row number list with the obfuscated sample row number, and sends it to the two participants respectively.

[0082] like Figure 4 As shown, by querying the result [2, -1, -1, -1, 0, 1], a trusted third party can determine the row number of any integer not equal to -1. The sample IDs of participant A corresponding to [0, 4, 5] are the intersection sample IDs, and non--1 integers are used as row numbers, i.e., row numbers. = [2, 0, 1] The sample ID of participant B is the intersection sample ID. In two integer arrays, the row number at the same position corresponds to the same sample ID for each participant. For example, the first element of both arrays, 0, 2, indicates that the sample ID of participant A at row number 0 is equal to the sample ID of participant B at row number 2. Specifically, the sample IDs are both abc0. Additionally, the number of intersection samples can be determined. =3.

[0083] A trusted third party adds a certain number of non-intersecting sample row numbers to two row number arrays as obfuscated sample row numbers, and sends them to participant A and participant B respectively. The number of obfuscated sample row numbers added can be determined based on the number of intersecting samples and a pre-set ratio. The sample size is determined jointly by the sample size of the party with less sample data (Party A). Specifically, a maximum of [number] members can be added. One obfuscated sample. Assume this embodiment sets the ratio of obfuscated samples to intersection samples. So, the number of rows of the added obfuscated samples This involves adding three obfuscated sample row numbers. All possible values ​​for the obfuscated sample row numbers are the row numbers of non-intersecting samples. Assuming the obfuscated samples added by the trusted third party through random, non-replacement sampling have row numbers [1, 2, 3] and [6, 3, 5] at the two participants, this means that the sample ID corresponding to row number 1 in participant A is equal to the sample ID corresponding to row number 6 in participant B, and they belong to the intersection sample. The trusted third party then adds the obfuscated sample row numbers to the intersection sample row numbers of the corresponding participants.

[0084] A trusted third party then follows a randomly generated order. The intersection of the obfuscated sample row numbers from the two participants is shuffled, and the shuffled obfuscated sample row numbers are... and Send them to participant A and participant B respectively. The list of line numbers after adding the obfuscated sample line numbers is as follows: [0, 4, 5, 1, 2, 3], [2, 0, 1, 6, 3, 5], the randomly generated shuffled index table is as follows: = [0,4, 1, 2, 5, 3], corresponding to the unscrambled index information = [0, 2, 3, 5, 1, 4], the shuffled list of row numbers sent to participant A and participant B are respectively =[0, 2, 4, 5, 3, 1]、 =[2, 3, 0, 1, 5, 6]. This includes the de-scrambling index information. The shuffled list of line numbers can be restored to its original state, satisfying the following conditions:

[0085]

[0086]

[0087]

[0088]

[0089] Generate descrambled index information based on the scrambled index table. and encrypt the result. Number of intersection samples Sends encrypted samples to the Multi-party Computation (MPC) encrypted computation engine. Number of intersection samples sent to the MPC encrypted computation engine. =3.

[0090] Step 3: Based on the received list of row numbers containing the confused sample row numbers (i.e., the third and fourth row number lists mentioned above), the two participants extract the feature data corresponding to the row numbers. , ;

[0091] like Figure 5 As shown, and The intersection samples are aligned one-to-one.

[0092] Step 4, as follows Figure 5 As shown, in the ciphertext calculation engine, the two encrypted and obfuscated samples are concatenated and then the ciphertext is used. Perform an unintentional shuffle sort to obtain the first few elements of the sorted array. A set of ciphertext samples are used as the intersection sample ciphertext data for subsequent modeling.

[0093] Specifically, the joint modeling phase runs in a multi-party secure computation ciphertext computation engine, which requires encrypting data from participant A. Data from participant B Data from trusted third parties And plaintext data from trusted third parties—the number of intersection samples .

[0094] The multi-party secure computation ciphertext computation engine received feature data, including obfuscated samples, from the two participating parties. , ,in, For encrypted feature data , For encrypted feature data ,Will and By concatenating the data, we obtain all the encryption features, including the obfuscated sample:

[0095] ;

[0096] Based on the encrypted descrambling index information sent by a trusted third party The ciphertext computation engine rearranges the ciphertext feature data by invoking an oblivious shuffle operation.

[0097] Unintentional shuffling of ciphertext refers to rearranging a ciphertext array according to a ciphertext index, such that the decrypted result is identical to the result of rearranging the plaintext array according to its plaintext index, i.e., satisfying the following condition:

[0098]

[0099] As can be seen from the construction method of the descrambling index, after arranging the scrambled sample IDs according to the descrambling index, the order of the sample IDs can be restored to the order before scrambling. Therefore, after performing an unintentional scrambling operation on the ciphertext features and the ciphertext descrambling index, the ciphertext features can be rearranged, and the features of a certain row can correspond to the sample features of the sample ID at the same position in the list of obfuscated sample IDs before the scrambling operation. For example, the row number lists of participants A and B after adding obfuscated sample IDs before the scrambling operation are respectively... [0, 4, 5, 1, 2, 3], [2, 0, 1, 6, 3, 5].

[0100] After knowing the size of the intersection After that, we can find out the actual intersection sample ID row number. , That is, the first three integers represent the positions of the intersection ID in the original ID sequences of the two participants. Thus, the first three integers are used to extract the values ​​from the rearranged ciphertext array. The row is the ciphertext of the features corresponding to the intersection sample IDs, that is...

[0101]

[0102]

[0103]

[0104] in, The encrypted intersection sample feature data is obtained in this embodiment without exposing the intersection ID to the two participants, and can be used for subsequent model training.

[0105] This embodiment provides a scheme to protect the sample alignment and joint modeling of the intersection samples of all data participants with the assistance of a trusted third party. Compared with the technical solution of this embodiment, traditional sample alignment schemes will expose the results of the intersection sample IDs. Although it is possible to find the intersection without exposing the sample intersection by simply comparing on the ciphertext, the efficiency of comparing on the ciphertext is extremely low. When the amount of data is large, the time cost of achieving sample alignment by comparing on the ciphertext is unacceptable.

[0106] This embodiment, by calling PIR and with the assistance of a trusted third party, can efficiently achieve sample alignment while protecting all sample IDs, including the intersection sample ID, for subsequent model training or prediction. It also ensures the effect of joint modeling on the intersection data, achieving a higher level of privacy protection, which is conducive to data circulation and realizing the value of data.

[0107] This embodiment provides a solution for federated modeling scenarios with high privacy protection requirements, offering a solution that meets privacy needs while fully utilizing high-quality data. The technical solution of this embodiment has significant market application prospects and can create direct commercial value in various data-intensive scenarios. This technical solution can be used for federated modeling optimization, improving model performance during data science and machine learning model training, enabling more accurate prediction and analysis of data, and is suitable for high-precision industries such as finance and healthcare. This technical solution can also be used for federated statistical augmentation, improving the accuracy of statistical analysis by introducing high-quality data, providing more reliable data support for key business decisions. This technical solution can also be applied to the financial industry, providing more accurate data analysis for risk management and credit assessment in institutions such as banks and insurance companies, improving business efficiency and decision accuracy. This technical solution can also be applied to the healthcare industry, providing efficient computing solutions for medical data analysis and public health policy formulation, ensuring data privacy while improving analytical accuracy and business benefits. This technical solution can also be applied to e-commerce, optimizing user behavior analysis and recommendation systems, improving customer experience and marketing effectiveness, and increasing the market competitiveness of enterprises.

[0108] By attracting high-quality data holders, the technical solution of this embodiment can significantly improve the accuracy and efficiency of data processing, reduce computing resource consumption, and thus lower operating costs. The technical solution of this embodiment utilizes high-precision models for analysis and prediction, enabling enterprises to make more accurate business decisions, reduce risks caused by data errors, and enhance their market competitiveness.

[0109] Please refer to Figure 6 This invention provides a sample processing device 100, applied to a trusted third party, comprising:

[0110] The acquisition module 101 is used to acquire the first row number list of the intersection sample of the first participant and the second participant in the first participant and the second row number list in the second participant;

[0111] The scrambling module 102 is used to add obfuscated sample row numbers to the first row number list and the second row number list respectively, and scramble the row numbers in the first row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain a third row number list; and scramble the row numbers in the second row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain a fourth row number list.

[0112] The first sending module 103 is used to send the third row number list to the first participant and the fourth row number list to the second participant;

[0113] The second sending module 104 is used to generate descrambling index information based on the scrambling index information. The descrambling index information is used to restore the scrambled row number list to the row number list before scrambling. The encrypted descrambling index information and the number of intersection samples are sent to the ciphertext calculation engine.

[0114] In some embodiments, the acquisition module 101 is specifically used to receive the query result returned by the second participant. The query result is sent by the second participant after receiving the query key sent by the first participant. The query key is the sample identifier of the first participant, used to indicate the row number of the sample identifier in the sample dataset of the second participant. The number of samples of the second participant is greater than the number of samples of the first participant. The query result includes the row number corresponding to the query key of the first participant in the sample dataset of the second participant. The first row number list of the intersection samples in the first participant and the second row number list in the second participant are obtained through the query result.

[0115] In this embodiment, a sequence number is added in advance to the sample data of the second participant as a query value. The sequence number corresponds one-to-one with the sample of the second participant and is the row number of the corresponding sample in the sample dataset of the second participant.

[0116] In some embodiments, if the query key of the first participant is not in the sample dataset of the second participant, the query result corresponding to the query key is a negative value.

[0117] In some embodiments, the number of the confused sample row numbers is determined based on the number of intersection samples and a set ratio. The value of the confused sample row number is determined by the number of samples of the first participant and the second participant.

[0118] In some embodiments, the number of the obfuscated sample row numbers is determined according to the following formula:

[0119]

[0120] in, The number of the confused sample row numbers, The sample size of the first participant. For the intersection

[0121] The number of samples in the set.

[0122] Please refer to Figure 7 This invention provides a sample processing device 200, applied to a ciphertext calculation engine, comprising:

[0123] The receiving module 201 is used to receive encrypted descrambling index information sent by a trusted third party and the number of intersection samples of the first participant and the second participant. ;

[0124] The acquisition module 202 is used to acquire encrypted first feature data sent by the first participant and encrypted second feature data sent by the second participant. The first feature data is feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party, and the second feature data is feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party.

[0125] The splicing module 203 is used to splice the first feature data and the second feature data to obtain ciphertext feature data, and rearrange the ciphertext feature data by calling the unintentional ciphertext scrambling operation according to the anti-scrambling index information to obtain a ciphertext array;

[0126] Extraction module 204 is used to extract the preceding text from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

[0127] Please refer to Figure 8 The present invention also provides an electronic device 300, including a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the computer program is executed by the processor 301, it implements the various processes of the above-described sample processing method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0128] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described sample processing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0129] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 and Figure 2 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0132] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A sample processing method, characterized in that, Applied to trusted third parties, including: Obtain the first row number list of the intersection sample of the first participant and the second participant in the first participant and the second row number list in the second participant; Add obfuscated sample row numbers to the first row number list and the second row number list respectively. Then, shuffle the row numbers in the first row number list after adding the obfuscated sample row numbers according to the shuffle index information to obtain the third row number list. Finally, shuffle the row numbers in the second row number list after adding the obfuscated sample row numbers according to the shuffle index information to obtain the fourth row number list. Send the third row number list to the first participant, and send the fourth row number list to the second participant; Based on the scrambled index information, anti-scrambled index information is generated. The anti-scrambled index information is used to restore the scrambled row number list to the row number list before scrambling. The encrypted anti-scrambled index information and the number of intersection samples are sent to the ciphertext calculation engine.

2. The sample processing method according to claim 1, characterized in that, Before obtaining the first row number list of the intersection sample of the first participant and the second participant in the first participant's and second row number list of the second participant's, the method further includes: The system receives a query result returned by the second participant. The query result is sent by the second participant after receiving a query key sent by the first participant. The query key is the sample identifier of the first participant, which is used to indicate the row number of the sample identifier in the sample dataset of the second participant. The number of samples of the second participant is greater than the number of samples of the first participant. The query result includes the row number of the query key of the first participant in the sample dataset of the second participant. The step of obtaining the first row number list of the intersection sample of the first participant and the second participant in the first participant and the second row number list in the second participant includes: The query results yield a list of the first row numbers of the intersection samples in the first participant and a list of the second row numbers in the second participant.

3. The sample processing method according to claim 2, characterized in that, If the query key of the first participant is not in the sample dataset of the second participant, the query result corresponding to the query key is a negative value.

4. The sample processing method according to claim 2, characterized in that, The number of the obfuscated sample row numbers is determined by the number of intersection samples and a set ratio. The value of the confused sample row number is determined by the number of samples of the first participant and the second participant.

5. The sample processing method according to claim 4, characterized in that, The number of the confused sample row numbers is determined according to the following formula: in, The number of the confused sample row numbers, The sample size of the first participant. The number of samples in the intersection.

6. A sample processing method, characterized in that, Applications to the ciphertext computation engine include: Receive encrypted descrambled index information sent by a trusted third party, as well as the number of intersection samples between the first and second participants. ; The encrypted first feature data sent by the first participant and the encrypted second feature data sent by the second participant are obtained. The first feature data is the feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party. The second feature data is the feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party. The first feature data and the second feature data are concatenated to obtain ciphertext feature data. Based on the anti-scrambling index information, the ciphertext feature data is rearranged by calling the unintentional ciphertext scrambling operation to obtain a ciphertext array. Extract the previous text from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

7. A sample processing device, characterized in that, Applied to trusted third parties, including: The acquisition module is used to acquire the first row number list of the intersection samples of the first participant and the second participant in the first participant and the second row number list in the second participant; The scrambling module is used to add obfuscated sample row numbers to the first row number list and the second row number list respectively, and scramble the row numbers in the first row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the third row number list; and scramble the row numbers in the second row number list after adding the obfuscated sample row numbers according to the scrambling index information to obtain the fourth row number list. The first sending module is used to send the third row number list to the first participant and the fourth row number list to the second participant; The second sending module is used to generate descrambled index information based on the scrambled index information. The descrambled index information is used to restore the scrambled row number list to the row number list before scrambling. The encrypted descrambled index information and the number of intersection samples are sent to the ciphertext calculation engine.

8. A sample processing device, characterized in that, Applications to the ciphertext computation engine include: The receiving module is used to receive encrypted descrambled index information sent by a trusted third party, as well as the number of intersection samples of the first and second participants. ; The acquisition module is used to acquire encrypted first feature data sent by the first participant and encrypted second feature data sent by the second participant. The first feature data is feature data extracted by the first participant from the third line number list after receiving the third line number list from the trusted third party, and the second feature data is feature data extracted by the second participant from the fourth line number list after receiving the fourth line number list from the trusted third party. The splicing module is used to splice the first feature data and the second feature data to obtain ciphertext feature data, and rearrange the ciphertext feature data by calling the unintentional ciphertext scrambling operation according to the anti-scrambling index information to obtain a ciphertext array; Extraction module, used to extract the first part of the ciphertext from the ciphertext array. The ciphertext feature data of the intersection samples is obtained by the following steps.

9. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the sample processing method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the sample processing method as described in any one of claims 1 to 6.

11. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the sample processing method as described in any one of claims 1 to 6.