Joint processing method and device, electronic equipment, storage medium and program product

By preprocessing and privacy-enhancing the data of the participating parties, constructing desensitized feature data and identification ciphertext, and passing them into a secure sandbox for model training, the problems of high model training cost and poor accuracy in two-party data modeling are solved, achieving the dual goals of data privacy protection and efficient model training.

CN120671180APending Publication Date: 2025-09-19JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510749368.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In two-party data modeling scenarios, existing technologies have problems with high model training costs and poor model training accuracy. In particular, in federated learning, communication overhead is large, TEE hardware costs are high, and efficiency is low.

Method used

By preprocessing and privacy-enhancing the original data of the participants, desensitized feature data and identification ciphertext are constructed and passed into a secure sandbox for model training and prediction. Differential privacy desensitization, feature name de-identification and OPRF encryption technology are used to ensure data privacy and optimize computing efficiency.

Benefits of technology

It reduces model training costs, improves model training accuracy and computing efficiency while ensuring data privacy. It is suitable for scenarios with limited network bandwidth or high real-time requirements and reduces hardware requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671180A_ABST
    Figure CN120671180A_ABST
Patent Text Reader

Abstract

The invention provides a joint processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of computers. The model joint processing method comprises the following steps: preprocessing original data of a second participant to obtain second sample feature data and a corresponding second sample identifier; respectively executing privacy enhancement operation on the second sample feature data and the corresponding second sample identifier to obtain second desensitization feature data and a corresponding second identifier ciphertext, and constructing a second two-tuple data set by the second desensitization feature data and the corresponding second identifier ciphertext; and receiving the first two-tuple data sent by the first participant, and executing model training and / or model prediction operation based on the first two-tuple data set and the second two-tuple data set in the security sandbox. According to the technical scheme, the isolation calculation characteristic of the security sandbox can optimize the calculation efficiency and reduce the training cost, so that data privacy protection and efficient and accurate model training are realized at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a model joint processing method, a model joint processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In a two-party data modeling scenario, Party A has sample data, while Party B is the executor of the modeling operation. To enhance Party B's modeling effect, Party A needs to provide more sample data. However, due to data privacy protection, Party A cannot directly provide sample data to Party B. Although federated learning solutions or trusted execution environments can protect data privacy to a certain extent, they have the defects of high model training costs and poor model training accuracy.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a model joint processing method, a model joint processing device, an electronic device, a computer-readable storage medium and a computer program product, which can at least to some extent improve the problems of high model training cost and poor model training accuracy in related technologies.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a model joint processing method is provided, which is applied to a first participant, including: preprocessing the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier; performing privacy enhancement operations on the first sample feature data and the corresponding first sample identifier respectively to obtain first desensitized feature data and a corresponding first identifier ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identifier ciphertext; sending the first two-tuple data set to a second participant, so that the second participant transfers the first two-tuple data set and a local second two-tuple data set into a security sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

[0007] In one embodiment of the present disclosure, a privacy enhancement operation is performed on the first sample feature data and the corresponding first sample identifier, respectively, to obtain first desensitized feature data and the corresponding first identifier ciphertext, including: the privacy enhancement operation includes differential privacy desensitization processing and encryption operation, differential privacy desensitization processing is performed on the first sample feature data to obtain first desensitized data; a normalization operation is performed on the first desensitized data to obtain normalized desensitized data; a feature name de-identification operation is performed on the normalized desensitized data to obtain the first desensitized feature data; the encryption operation is performed on the first sample identifier to obtain the first identifier ciphertext.

[0008] In one embodiment of the present disclosure, differential privacy desensitization processing is performed on the first sample feature data to obtain first desensitized data, including: adding noise to the first sensitive feature in the first sample feature data to obtain the first desensitized data; or performing a statistical perturbation operation on the second sensitive feature in the data chain in the first sample feature data to obtain the first desensitized data; or combining multiple features in the first sample feature data and performing differential privacy processing on the combined multiple features to obtain the first desensitized data.

[0009] In one embodiment of the present disclosure, a feature name de-identification operation is performed on the normalized desensitized data to obtain the first desensitized feature data, including: extracting the original feature name in the normalized desensitized data; mapping the original feature name to a string based on a hash mapping, or mapping the original feature name to a random identifier based on a random mapping table to perform the feature name de-identification operation to obtain the first desensitized feature data.

[0010] In one embodiment of the present disclosure, performing the encryption operation on the first sample identifier to obtain the first identifier ciphertext includes: receiving an oblivious pseudo-random function (OPRF) key sent by the second participant, or negotiating with the second participant to generate the OPRF key; and encrypting the first sample identifier based on the OPRF key to obtain the first identifier ciphertext.

[0011] In one embodiment of the present disclosure, the original data of the first participant is preprocessed to obtain first sample feature data and a corresponding first sample identifier, including: identifying abnormal points that deviate from a preset range in the same type of features of the original data of the first participant; deleting the original data including the abnormal points, or performing a repair operation on the abnormal points to obtain cleaned data; determining a sample identifier corresponding to the cleaned data, determining the sample identifier as a primary key, and obtaining the first sample feature data and the corresponding first sample identifier based on the association relationship between the primary key and the foreign key.

[0012] According to another aspect of the present disclosure, a model joint processing method is provided, which is applied to a second participant, including: preprocessing the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier; performing privacy enhancement operations on the second sample feature data and the corresponding second sample identifier respectively to obtain second desensitized feature data and a corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext; receiving first two-tuple data sent by the first participant, so as to transfer the first two-tuple data set and the second two-tuple data set into a security sandbox, and performing model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

[0013] In one embodiment of the present disclosure, model training and / or model prediction operations are performed within the security sandbox based on the first two-tuple data set and the second two-tuple data set, including: within the security sandbox, the first two-tuple data set includes first desensitized feature data and a corresponding first identification ciphertext, calculating the intersection between the first identification ciphertext and the second identification ciphertext to obtain an intersection ciphertext; extracting intersection desensitized feature data corresponding to the intersection ciphertext from the first desensitized feature data and the second desensitized feature data, respectively; and performing the model training and / or model prediction operations based on the intersection ciphertext and the corresponding intersection desensitized feature data.

[0014] In one embodiment of the present disclosure, before transferring the first two-tuple data set and the second two-tuple data set into a security sandbox, the process further includes: initializing the security sandbox and configuring a communication channel; configuring an OPRF algorithm in the security sandbox; generating an OPRF key in the security sandbox based on the OPRF algorithm; and obtaining the OPRF key from the security sandbox based on the communication channel.

[0015] In one embodiment of the present disclosure, a privacy enhancement operation is performed on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and a corresponding second identifier ciphertext, including: the privacy enhancement operation includes differential privacy desensitization processing and encryption operation, differential privacy desensitization processing is performed on the second sample feature data to obtain second desensitized data; a normalization operation is performed on the second desensitized data to obtain normalized desensitized data; a feature name de-identification operation is performed on the normalized desensitized data to obtain the second desensitized feature data; the encryption operation is performed on the second sample identifier to obtain the second identifier ciphertext.

[0016] In one embodiment of the present disclosure, differential privacy desensitization processing is performed on the second sample feature data to obtain second desensitized data, including: adding noise to the first sensitive feature in the second sample feature data to obtain the second desensitized data; or performing a statistical perturbation operation on the second sensitive feature in the data chain in the second sample feature data to obtain the second desensitized data; or combining multiple features in the second sample feature data and performing differential privacy processing on the combined multiple features to obtain the second desensitized data.

[0017] In one embodiment of the present disclosure, a feature name de-identification operation is performed on the normalized desensitized data to obtain the second desensitized feature data, including extracting the original feature name in the normalized desensitized data; mapping the original feature name to a string based on a hash mapping, or mapping the original feature name to a random identifier based on a random mapping table to perform the feature name de-identification operation to obtain the second desensitized feature data.

[0018] In one embodiment of the present disclosure, performing the encryption operation on the second sample identifier to obtain the second identifier ciphertext includes encrypting the second sample identifier based on the acquired OPRF key to obtain the second identifier ciphertext.

[0019] In one embodiment of the present disclosure, the original data of the second participant is preprocessed to obtain second sample feature data and a corresponding second sample identifier, including: identifying abnormal points that deviate from a preset range in the same type of features of the original data of the second participant; deleting the original data including the abnormal points, or performing a repair operation on the abnormal points to obtain cleaned data; determining a sample identifier corresponding to the cleaned data, determining the sample identifier as a primary key, and obtaining the second sample feature data and the corresponding second sample identifier based on the association relationship between the primary key and the foreign key.

[0020] According to another aspect of the present disclosure, a model joint processing device is provided, including: a first preprocessing module, used to preprocess the original data of the first participant to obtain first sample feature data and the corresponding first sample identification; a first privacy enhancement operation module, used to perform privacy enhancement operations on the first sample feature data and the corresponding first sample identification, respectively, to obtain first desensitized feature data and the corresponding first identification ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identification ciphertext; a sending module, used to send the first two-tuple data set to the second participant, so that the second participant transfers the first two-tuple data set and the local second two-tuple data set into a security sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set in the security sandbox.

[0021] According to another aspect of the present disclosure, a model joint processing device is provided, including: a second preprocessing module, used to preprocess the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier; a second privacy enhancement operation module, used to perform privacy enhancement operations on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and a corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext; a processing module, used to receive the first two-tuple data sent by the first participant, so as to transfer the first two-tuple data set and the second two-tuple data set into a security sandbox, and perform model training and / or model prediction operations based on the second two-tuple data set and the second two-tuple data set in the security sandbox.

[0022] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above-mentioned model joint processing methods by executing the executable instructions.

[0023] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned model joint processing methods is implemented.

[0024] The model joint processing solution provided by the embodiments of the present disclosure pre-processes and performs privacy enhancement operations on the original data of at least two participants. Under the premise of ensuring data privacy and security, the processed two-tuple data set is transferred to a secure sandbox to perform model training and prediction operations in the secure sandbox. This not only prevents the direct exposure of the original data and meets the privacy protection requirements, but also expands the feature dimension by fusing multi-party data and enhances the model's ability to capture data features. In addition, the isolated computing characteristics of the secure sandbox can also optimize computing efficiency and reduce training costs, thereby achieving the dual goals of data privacy protection and efficient and accurate model training.

[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0027] Figure 1 A flow chart showing a model joint processing method according to an embodiment of the present disclosure is shown;

[0028] Figure 2 A flowchart showing another model joint processing method according to an embodiment of the present disclosure is shown;

[0029] Figure 3 A schematic diagram showing another model joint processing method in an embodiment of the present disclosure;

[0030] Figure 4 A flow chart showing another model joint processing method according to an embodiment of the present disclosure is shown;

[0031] Figure 5 A flow chart showing another model joint processing method according to an embodiment of the present disclosure is shown;

[0032] Figure 6 A schematic diagram showing a model joint processing device according to an embodiment of the present disclosure is shown;

[0033] Figure 7 A schematic diagram showing another model joint processing device according to an embodiment of the present disclosure;

[0034] Figure 8 A schematic diagram showing an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0036] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0037] In existing two-party data modeling scenarios, Party A possesses the majority of user samples (such as user IDs) and user features, while Party B possesses a smaller portion. To enhance Party B's modeling, Party A needs to provide more samples or user features with more dimensions to help Party B train its model. However, both parties' data are valuable assets and cannot be directly provided to each other.

[0038] Currently, the mainstream solutions are federated learning and trusted execution environment.

[0039] Federated Learning (FL): Federated learning is a distributed machine learning technology that allows parties to collaboratively train models without sharing original data while protecting data privacy.

[0040] Federated learning is divided into horizontal federated learning and vertical federated learning. The application scenario of horizontal federated learning is that all parties have the same user feature space but different user samples, and the application scenario of vertical federated learning is that all parties have the same user samples but different user feature spaces. Horizontal federated learning trains model parameters locally and transmits the model parameters to the server, while vertical federated learning calculates gradients and intermediate results locally and transmits the intermediate results to the server.

[0041] Taking horizontal federated learning as an example, its basic process is as follows: the server initializes a global model and sends the model parameters to the participants; each party trains the model on local data and updates the parameters; each party encrypts the model parameters and shares them; the server aggregates the model parameters of all participants and updates the global model; iterates several rounds until the model converges.

[0042] It can be seen that federated learning requires frequent exchange of model parameters between all parties, resulting in high communication overhead. Since multiple rounds of parameter exchange and updates are required, the training time is also long. In addition, the implementation and maintenance of the federated learning framework is relatively complex and requires a high technical threshold.

[0043] In two-party modeling scenarios, Trusted Execution Environments (TEE) technology can leverage hardware-level security to perform data processing and model training in a hardware-isolated environment, ensuring that data is not leaked during processing. However, the cost of TEE hardware and related technologies is high and requires specialized hardware support. In addition, the computing resources of the TEE environment are limited, which may affect the efficiency of model training and prediction.

[0044] Therefore, a joint processing model is urgently needed to solve the privacy protection, efficiency and cost issues in the two-party data modeling in the above solution.

[0045] The solution provided in this application preprocesses and performs privacy enhancement operations on the original data of at least two participants. Under the premise of ensuring data privacy and security, the processed two-tuple data set is transferred to a secure sandbox to perform model training and prediction operations in the secure sandbox. This not only prevents the direct exposure of the original data and meets the privacy protection requirements, but also expands the feature dimension by fusing multi-party data and enhances the model's ability to capture data features. In addition, the isolated computing characteristics of the secure sandbox can also optimize computing efficiency and reduce training costs, thereby achieving the dual goals of data privacy protection and efficient and accurate model training.

[0046] To facilitate understanding, several terms involved in this application are first explained below.

[0047] OPRF (Oblivious Pseudo-Random Function) is a cryptographic constructor F k (x), which is similar in form to the key hash function Hash k (x), but the difference is that it allows one party (the client) to calculate the pseudo-random function F held by the other party (the server) k (·) without revealing the input x to the server, and the construction process can also be independent of the hash. At the same time, the server cannot know what the client's input x is. In short, OPRF enables one party to encrypt its input and send it to another party for computation without revealing any information about the input.

[0048] Security sandbox: A technical environment that ensures data and computing security through isolation, restriction, and monitoring. Its core goal is to build a trusted execution space in an untrusted external environment to prevent sensitive data leakage or malicious operations.

[0049] The model joint processing method in this example implementation will be described in more detail below with reference to the accompanying drawings and examples.

[0050] like Figure 1 As shown, a model joint processing method according to an embodiment of the present disclosure is applied to a first participant, including:

[0051] Step S102: pre-process the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier.

[0052] In some embodiments, the characteristic fields in the original data include user age, income, behavior records, etc. The first sample identifier is used to uniquely identify the identifier of each sample, which may be the ID of the original data (such as user ID, device IMEI) or a new ID generated after cleaning (such as UUID).

[0053] In some embodiments, the preprocessing may be outlier processing performed locally by each participant on its own feature data.

[0054] Step S104: Perform privacy enhancement operations on the first sample feature data and the corresponding first sample identifier, respectively, to obtain first desensitized feature data and the corresponding first identifier ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identifier ciphertext.

[0055] In some embodiments, the privacy enhancement operation on the first sample feature data refers to performing privacy protection on the feature data, including but not limited to differential privacy and homomorphic encryption.

[0056] In some embodiments, the privacy enhancement operation on the first sample identification refers to applying a cryptographic transformation to the sample identification.

[0057] In some embodiments, the feature data {d j}={data j}, after privacy enhancement operation, the desensitized feature data {m j}.

[0058] In some embodiments, each participant forms a binary data set based on the plaintext ID and its corresponding desensitized feature data. Find new bigram datasets Each of them is passed into the security sandbox, where ID indicates the document ID, c j Refers to the identification ciphertext.

[0059] Step S106: Send the first two-tuple data set to the second participant, so that the second participant transfers the first two-tuple data set and the local second two-tuple data set into the security sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

[0060] In some embodiments, a security sandbox refers to an isolated computing environment. Compared with traditional virtual machines, sandboxes (such as containers and function-level sandboxes) start quickly and have low resource usage, thereby improving processing efficiency while reducing resource usage.

[0061] In some embodiments, when the model predicts, confidential feature data is input, and a confidential prediction result is output, which is then decrypted and returned to the requesting party.

[0062] In this embodiment, by preprocessing and privacy enhancement operations on the original data of at least two participants, the processed two-tuple data set is transferred to a secure sandbox under the premise of ensuring data privacy and security, so as to perform model training and prediction operations in the secure sandbox. This not only prevents the direct exposure of the original data and meets the privacy protection requirements, but also expands the feature dimension by fusing multi-party data and enhances the model's ability to capture data features. In addition, the isolated computing characteristics of the secure sandbox can also optimize computing efficiency and reduce training costs, thereby achieving the dual goals of data privacy protection and efficient and accurate model training.

[0063] In addition, the solution disclosed in the present invention only requires a one-time transmission of the desensitized binary data set (including desensitized feature data and identification ciphertext), without the need for multiple rounds of parameter interaction. The security sandbox directly performs calculations based on the preprocessed static data, greatly reducing the number of communications and data transmission volume. It is especially suitable for scenarios with limited network bandwidth or high real-time requirements. Moreover, the security sandbox is implemented at the software level and does not require additional dedicated hardware. It can be deployed using general computing resources, significantly reducing hardware costs and deployment thresholds, and adapting to more participants.

[0064] like Figure 2 As shown, in one embodiment of the present disclosure, a privacy enhancement operation is performed on the first sample feature data and the corresponding first sample identifier to obtain first desensitized feature data and the corresponding first identifier ciphertext, including:

[0065] In step S202 , the privacy enhancement operation includes differential privacy desensitization processing and encryption operation, and the differential privacy desensitization processing is performed on the first sample feature data to obtain first desensitized data.

[0066] In some embodiments, individual information is protected by adding noise to the data to ensure that the presence or absence of a single data does not significantly affect the data analysis results. Its core is based on mathematical theory, using the randomness of noise to cover up the details of the real data. In specific operations, for numerical features, noise that conforms to the Laplace distribution or Gaussian distribution is usually added. For example, for values ​​such as age and income, a random noise value is added to the original value, so that even if the attacker obtains the data set, he cannot determine the true feature value of a specific individual. For non-numerical features (such as categorical data), the contribution of each data is blurred by adjusting its distribution probability and adding disturbances, thereby achieving protection of individual privacy.

[0067] Step S204: performing a normalization operation on the first desensitized data to obtain normalized desensitized data.

[0068] In some embodiments, the normalization operation is mainly used to eliminate the differences in dimensions and value ranges between different features, making the data comparable and facilitating subsequent model training and analysis.

[0069] Step S206: perform a feature name de-identification operation on the normalized desensitized data to obtain first desensitized feature data.

[0070] In some embodiments, the feature name de-identification operation is to further reduce the risk of data privacy leakage by removing or replacing sensitive information or identifiable information that may be contained in the feature name. For example, "user ID number" is changed to "identity identifier" or replaced with a common code (such as "feature 1", "feature 2", etc.).

[0071] Step S208: Perform an encryption operation on the first sample identifier to obtain a first identifier ciphertext.

[0072] In some embodiments, OPRF encryption is often used as the underlying cryptographic tool within a secure sandbox, which can securely generate shared keys to achieve data alignment between different parties.

[0073] In this embodiment, differential privacy is used to add noise at the data content level to blur individual privacy information. Normalization operations improve data quality to optimize subsequent models without affecting privacy protection. Feature naming and identification operations eliminate sensitive clues from the data identification level, while encryption operations ensure the security of sample identification. Tasks are atomized by configuring a security sandbox to ensure that desensitized data can only be used for model training in a single sandbox environment and cannot be reused. On the one hand, the risk of privacy leakage during data transmission and processing is reduced, meeting the requirements of data security and privacy protection regulations. On the other hand, the processed data retains key feature relationships and distribution characteristics while having better standardization and usability, and can effectively support joint modeling and analysis with other data in environments such as security sandboxes.

[0074] In one embodiment of the present disclosure, performing differential privacy desensitization processing on first sample feature data to obtain first desensitized data includes:

[0075] Noise is added to the first sensitive feature in the first sample feature data to obtain the first desensitized data; or a statistical perturbation operation is performed on the second sensitive feature in the data chain in the first sample feature data to obtain the first desensitized data; or multiple features in the first sample feature data are combined, and differential privacy processing is performed on the combined multiple features to obtain the first desensitized data.

[0076] In some embodiments, for a numerical first sensitive feature (such as income, age), noise that conforms to a specific probability distribution (such as Laplace distribution, Gaussian distribution) is usually added.

[0077] In some embodiments, there are logical associations among the features in the data chain. By perturbing the statistics (such as mean, median, sum, etc.), sensitive correlation information between the data is destroyed. For example, in a set of user consumption data, perturbations are added to the statistics of the total consumption amount, making it impossible for attackers to infer the consumption behavior of specific users by analyzing the changes in the total consumption amount.

[0078] In some embodiments, considering that the combination of multiple features may leak user privacy, after combining multiple features, the combined features are processed as a whole using differential privacy technology. By adding noise to the combined features, the overall distribution of the combined data changes slightly, blurring the individual feature combination information.

[0079] In this embodiment, by performing differential privacy desensitization processing on the first sample feature data, and through differential reconstruction after indirect differential privacy desensitization, it is possible to protect data privacy while effectively filtering out dirty data, retaining the data availability and statistical characteristics as much as possible, ensuring that the data can still be used for subsequent model training, data analysis and other tasks, so as to improve the accuracy of model training and achieve a balance between data privacy protection and data value utilization.

[0080] In one embodiment of the present disclosure, a feature name de-identification operation is performed on the normalized desensitized data to obtain first desensitized feature data, including:

[0081] Extract the original feature name from the normalized desensitized data; map the original feature name to a string based on a hash map, or map the original feature name to a random identifier based on a random mapping table to perform a feature name de-identification operation to obtain the first desensitized feature data.

[0082] In some embodiments, the hash map converts the original feature name into a fixed-length string using a hash function.

[0083] In some embodiments, a random mapping table is pre-created, in which each original feature name is randomly mapped to a unique random identifier.

[0084] In this embodiment, by extracting the original feature name and performing the feature name de-identification operation by means of hash mapping or random mapping table, the sensitive clues and business-related information that may be contained in the original feature name can be eliminated, thereby effectively enhancing the privacy protection capability of the data at the data identification level. Moreover, the de-identified data can still meet the needs of subsequent model training and data analysis while retaining the intrinsic structure and numerical relationship of the data, thus achieving a balance between data privacy protection and data availability, improving the security of data during cross-party interaction and processing, and helping to meet strict data security and privacy protection rules.

[0085] In one embodiment of the present disclosure, performing an encryption operation on the first sample identifier to obtain a first identifier ciphertext includes:

[0086] An oblivious pseudo-random function (OPRF) key is received from the second party, or an OPRF key is generated through negotiation with the second party; and the first sample identifier is encrypting based on the OPRF key to obtain a first identifier ciphertext.

[0087] In some embodiments, after the second party generates or obtains the OPRF key, it sends it to the first party through a secure communication channel (such as a TLS encrypted channel). The transmission process of the key is subject to strict security protection to prevent the key from being stolen or tampered with during transmission, ensuring that only the authorized first party can receive and use the key.

[0088] In some embodiments, the two parties negotiate to determine cryptographic security parameters, including but not limited to the elliptic curve type and key length, as well as the hash function to be used. The first party generates a random number a and calculates the blinding factor x. , , the second party generates a private key b, and receives x , , calculate y , =Bx , And return the result, the first party uses a to calculate y = a -1 y , At this time, y is used as the ciphertext of sample identifier x, and the second participant uses b to encrypt. That is, the complete key is composed of a and b, but each party only holds part (a or b).

[0089] In some embodiments, the oblivious pseudorandom function OPRF is a cryptographic protocol that allows one party (first party) to use a key provided by another party (second party) to perform encryption calculation on a first sample identifier, while neither party can know the other's key information.

[0090] In some embodiments, each participant performs OPRF encryption on his or her sample ID according to the OPRF key k, and obtains the OPRF-encrypted ID ciphertext c j =OPRF(k,ID j ).

[0091] In some embodiments, taking the OPRF algorithm based on elliptic curve as an example, a specific mathematical operation is performed on the sample identifier and the OPRF key on the elliptic curve to generate a pseudo-random output result, which is the first identification ciphertext. The encryption process is deterministic, that is, the same sample identifier and OPRF key will always generate the same identification ciphertext. This feature provides a basis for subsequent ciphertext-based matching operations.

[0092] In this embodiment, in the data interaction scenario, the sample identifier is often associated with specific individual information. The sample identifier is encrypted using the OPRF key so that the sample identification information of the first participant exists only in ciphertext form. The OPRF key bound to a single task ensures that the IDs of all parties involved in the task are in OPRF ciphertext form when introduced into the security sandbox. When the data sets are merged, the intersection of the OPRF ciphertexts of multiple parties needs to be found. Therefore, from the perspective of cryptographic security, it is further ensured that the data of a single task can only be used for the task. Even if the data is intercepted, a third party cannot obtain the original identification content. In addition, the determinism of OPRF encryption is conducive to realizing sample matching and joint analysis based on ciphertext. On the premise of ensuring privacy, it supports the first participant and the second participant to perform data alignment and collaborative calculation in an environment such as a security sandbox.

[0093] In one embodiment of the present disclosure, preprocessing the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier includes:

[0094] Identify abnormal points that deviate from a preset range in the same type of features of the original data of the first participant; delete the original data including the abnormal points, or perform a repair operation on the abnormal points to obtain cleaned data; determine a sample identifier corresponding to the cleaned data, determine the sample identifier as a primary key, and obtain first sample feature data and a corresponding first sample identifier based on the association relationship between the primary key and the foreign key.

[0095] In some embodiments, a field that uniquely identifies each record (such as user ID, order number) is selected from the cleaned data as the primary key. By mapping the primary key with the foreign key (the field in other tables that references the primary key), the cleaned data is associated with other data sets (such as user profiles, transaction records) to construct a complete feature matrix.

[0096] In this embodiment, by deleting or repairing outliers, the interference of noise data on model training can be reduced. By determining the sample identifier as the primary key and based on the association relationship between the primary key and the foreign key, multi-source data (such as user behavior and portrait) can be integrated to enrich the feature dimensions, which is conducive to improving the prediction accuracy of subsequent model predictions.

[0097] like Figure 3 As shown, a model joint processing method according to another embodiment of the present disclosure is applied to a second participant, including:

[0098] Step S302: pre-process the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier.

[0099] Step S304: Perform privacy enhancement operations on the second sample feature data and the corresponding second sample identifier to obtain second desensitized feature data and the corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext.

[0100] Step S306: Receive the first two-tuple data sent by the first participant to transfer the first two-tuple data set and the second two-tuple data set into the security sandbox, and perform model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

[0101] In this embodiment, by preprocessing and privacy enhancement operations on the original data of at least two participants, the processed two-tuple data set is transferred to a secure sandbox under the premise of ensuring data privacy and security, so as to perform model training and prediction operations in the secure sandbox. This not only prevents the direct exposure of the original data and meets the privacy protection requirements, but also expands the feature dimension by fusing multi-party data and enhances the model's ability to capture data features. In addition, the isolated computing characteristics of the secure sandbox can also optimize computing efficiency and reduce training costs, thereby achieving the dual goals of data privacy protection and efficient and accurate model training.

[0102] like Figure 4 As shown, in one embodiment of the present disclosure, performing a model training and / or model prediction operation based on a first two-tuple data set and a second two-tuple data set within a security sandbox includes:

[0103] Step S402: In the security sandbox, the first two-tuple data set includes the first desensitized feature data and the corresponding first identification ciphertext, and the intersection between the first identification ciphertext and the second identification ciphertext is calculated to obtain the intersection ciphertext.

[0104] In some embodiments, based on the determinism of OPRF, that is, the same original ID generates the same ciphertext, hash collision or Bloom filter is used to quickly locate the ciphertext shared by the participants, ensuring that only the intersection ciphertext is identified and the non-intersection ciphertext remains confidential.

[0105] Step S404: extracting intersection desensitized feature data corresponding to the intersection ciphertext from the first desensitized feature data and the second desensitized feature data respectively.

[0106] In some embodiments, the ciphertext intersection is associated with the desensitized feature data through a primary key (such as a ciphertext ID), and a secure JOIN operation is performed within the sandbox to extract only feature rows that match the intersection ciphertext.

[0107] Step S406: Perform model training and / or model prediction operations based on the intersection ciphertext and the corresponding intersection desensitized feature data.

[0108] In some embodiments, the joint model is trained using multi-source feature data of the intersection samples (such as the behavioral features of the first participant + the portrait features of the second participant).

[0109] In this embodiment, by configuring the ciphertext intersection mechanism, it is possible to ensure that the participating parties can only obtain shared sample information while preventing the leakage of non-intersection data. The generated intersection desensitized feature data integrates the complementary features of multiple parties, so that the model training obtains a more comprehensive sample dimension, which is beneficial to improving the model training accuracy and generalization ability.

[0110] In one embodiment of the present disclosure, before the first two-tuple data set and the second two-tuple data set are transferred to the security sandbox, the method further includes:

[0111] Initialize the security sandbox and configure the communication channel.

[0112] In some embodiments, initializing the security sandbox includes: creating an isolated execution environment based on a hardware security module or container technology, and ensuring confidentiality and integrity during data calculation through memory encryption and process isolation.

[0113] In some embodiments, a TLS / SSL encrypted channel is established, and digital certificates are used to verify the identities of the participants, prevent man-in-the-middle attacks, and ensure the security of the OPRF key and data transmission process.

[0114] Configure the OPRF algorithm in the security sandbox.

[0115] In some embodiments, configuring the OPRF algorithm in a security sandbox includes: deploying an OPRF protocol based on elliptic curves or discrete logarithms in the sandbox, and configuring components such as hash functions and random number generators to ensure that the algorithm implementation complies with cryptographic security standards.

[0116] Generate OPRF keys based on the OPRF algorithm in a secure sandbox.

[0117] In some embodiments, a private key k is generated by a random number generator within the sandbox, and a public key K is calculated based on the private key k as the OPRF key.

[0118] Obtain the OPRF key from the secure sandbox based on the communication channel.

[0119] In some embodiments, the sandbox only outputs the public key K or an intermediate calculation result based on K through a restricted interface, and the private key k does not leave the sandbox.

[0120] In this embodiment, by initializing a secure sandbox, configuring an encrypted communication channel and the OPRF algorithm, OPRF keys are generated and securely distributed in a hardware-level isolated environment, building a full-process security environment from key generation to transmission. This strengthens the security and verifiability of key management while reducing the difficulty of multi-party collaboration model processing.

[0121] In one embodiment of the present disclosure, a privacy enhancement operation is performed on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and the corresponding second identifier ciphertext, including: the privacy enhancement operation includes differential privacy desensitization processing and encryption operation, differential privacy desensitization processing is performed on the second sample feature data to obtain second desensitized data; a normalization operation is performed on the second desensitized data to obtain normalized desensitized data; a feature name de-identification operation is performed on the normalized desensitized data to obtain second desensitized feature data; an encryption operation is performed on the second sample identifier to obtain a second identifier ciphertext.

[0122] In one embodiment of the present disclosure, differential privacy desensitization processing is performed on the second sample feature data to obtain second desensitized data, including adding noise to the first sensitive feature in the second sample feature data to obtain the second desensitized data; or performing a statistical perturbation operation on the second sensitive feature in the data chain in the second sample feature data to obtain the second desensitized data; or combining multiple features in the second sample feature data and performing differential privacy processing on the combined multiple features to obtain the second desensitized data.

[0123] In one embodiment of the present disclosure, a feature name de-identification operation is performed on the normalized desensitized data to obtain second desensitized feature data, including: extracting the original feature name in the normalized desensitized data; mapping the original feature name to a string based on a hash map, or mapping the original feature name to a random identifier based on a random mapping table to perform the feature name de-identification operation to obtain the second desensitized feature data.

[0124] In one embodiment of the present disclosure, performing an encryption operation on the second sample identifier to obtain a second identifier ciphertext includes: encrypting the second sample identifier based on the acquired OPRF key to obtain the second identifier ciphertext.

[0125] In one embodiment of the present disclosure, the original data of the second participant is preprocessed to obtain second sample feature data and a corresponding second sample identifier, including: identifying abnormal points that deviate from a preset range in the same type of features of the original data of the second participant; deleting the original data including the abnormal points, or performing a repair operation on the abnormal points to obtain cleaned data; determining the sample identifier corresponding to the cleaned data, determining the sample identifier as the primary key, and obtaining the second sample feature data and the corresponding second sample identifier based on the association relationship between the primary key and the foreign key.

[0126] Figure 5 An architecture involving two-party data processing and model training is shown, including the local environment of participant A (second participant), the local environment of participant B (first participant) and a secure sandbox environment.

[0127] "Party A initializes the sandbox" and generates the "OPRF key k".

[0128] The local environment of participant A includes "participant A feature data original" and "participant A sample ID original", the participant A feature data original corresponds to the second sample feature data, and the participant A sample ID original corresponds to the second sample identifier.

[0129] Among them, the original feature data is the original attributes or variable set that describes the characteristics of the sample. It is a specific characterization of a certain aspect of the sample. For example, when analyzing the customer's credit status, age, income, credit history length, etc. are the original feature data.

[0130] The sample ID is the original number or identifier assigned to uniquely identify each sample in the dataset.

[0131] The original feature data enters the "Data Desensitization Platform", and the "Original Feature Data of Participant A" is processed through "Differential Privacy Desensitization" to obtain the "Desensitized Feature Data of Participant A". The "OPRF(k,ID)" operation is used on the "Original Sample ID of Participant A" to obtain the "OPRF Ciphertext of Participant A Sample ID".

[0132] The local environment of participant B includes "participant B feature data original" and "participant B sample ID original", the participant B feature data original corresponds to the first sample feature data, and the participant B sample ID original corresponds to the first sample identifier.

[0133] The original feature data enters the "Data Desensitization Platform", and the "Original Feature Data of Participant B" is processed through "Differential Privacy Desensitization" to obtain "Desensitized Feature Data of Participant B". The "OPRF(k,ID)" operation is used on the "Original Sample ID of Participant B" to obtain the "OPRF Ciphertext of Participant B Sample ID".

[0134] In a secure sandbox environment, the "Party A desensitized feature data" and "Party A sample ID OPRF ciphertext" from Party A, and the "Party B sample ID OPRF ciphertext" and "Party B desensitized feature data" from Party B are received.

[0135] Based on the received sample ID OPRF ciphertext, the "intersection based on OPRF ciphertext" operation is performed to obtain the "intersection sample ID OPRF ciphertext" and the "intersection sample corresponding desensitized feature data".

[0136] Finally, the "intersection sample ID OPRF ciphertext" and "intersection sample corresponding desensitized feature data" are used for model training and / or model prediction.

[0137] In this embodiment, through data desensitization and OPRF-based encryption processing, the intersection processing of two-party data and subsequent model training or prediction are achieved under the premise of protecting data privacy, emphasizing the key operations in a secure sandbox environment, thereby improving data security and operation efficiency.

[0138] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0139] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0140] Refer to the following Figure 6 The model joint processing device 600 according to this embodiment of the present invention will be described. Figure 6 The model joint processing device 600 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0141] The model joint processing device 600 is implemented in the form of a hardware module. Components of the model joint processing device 600 may include, but are not limited to: a first preprocessing module 602 for preprocessing the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier; a first privacy enhancement operation module 604 for performing privacy enhancement operations on the first sample feature data and the corresponding first sample identifier, respectively, to obtain first desensitized feature data and a corresponding first identifier ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identifier ciphertext; a sending module 606 for sending the first two-tuple data set to the second participant, so that the second participant transfers the first two-tuple data set and the local second two-tuple data set to a secure sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the secure sandbox.

[0142] Refer to the following Figure 7The model joint processing device 700 according to this embodiment of the present invention will be described. Figure 7 The model joint processing device 700 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0143] The model joint processing device 700 is implemented in the form of a hardware module. Components of the model joint processing device 700 may include, but are not limited to: a second preprocessing module 702 for preprocessing the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier; a second privacy enhancement operation module 704 for performing privacy enhancement operations on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and a corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext; and a processing module 706 for receiving the first two-tuple data sent by the first participant, transferring the first two-tuple data set and the second two-tuple data set into a secure sandbox, and performing model training and / or model prediction operations based on the second two-tuple data set and the second two-tuple data set within the secure sandbox.

[0144] Refer to the following Figure 8 An electronic device 800 according to this embodiment of the present invention will be described. Figure 8 The electronic device 800 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0145] like Figure 8 As shown, electronic device 800 is implemented as a general-purpose computing device. Components of electronic device 800 may include, but are not limited to, the aforementioned at least one processing unit 810, the aforementioned at least one storage unit 820, and a bus 830 connecting various system components (including storage unit 820 and processing unit 810).

[0146] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 810 can perform the following steps: Figure 1 Steps S102 and S106 shown in , and other steps defined in the model joint processing method of the present disclosure.

[0147] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .

[0148] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0149] Bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0150] The electronic device 800 can also communicate with one or more external devices 860 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 840. Furthermore, the electronic device 800 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 850. As shown, the network adapter 850 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0151] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0152] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0153] According to an embodiment of the present invention, a program product for implementing the above-mentioned method can be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0154] The program product may be implemented in any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0155] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0156] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0157] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0158] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0159] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0160] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0161] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A model joint processing method, characterized in that: Applicable to the first party, including: Preprocessing the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier; Performing a privacy enhancement operation on the first sample feature data and the corresponding first sample identifier, respectively, to obtain first desensitized feature data and a corresponding first identifier ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identifier ciphertext; The first two-tuple data set is sent to a second participant, so that the second participant transfers the first two-tuple data set and the local second two-tuple data set into a security sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

2. The model joint processing method according to claim 1, characterized in that: Performing a privacy enhancement operation on the first sample feature data and the corresponding first sample identifier to obtain first desensitized feature data and a corresponding first identifier ciphertext, including: The privacy enhancement operation includes differential privacy desensitization processing and encryption operation, performing differential privacy desensitization processing on the first sample feature data to obtain first desensitized data; Performing a normalization operation on the first desensitized data to obtain normalized desensitized data; Performing a feature name de-identification operation on the normalized desensitized data to obtain the first desensitized feature data; The encryption operation is performed on the first sample identifier to obtain the first identifier ciphertext.

3. The model joint processing method according to claim 2, characterized in that: Performing differential privacy desensitization processing on the first sample feature data to obtain first desensitized data includes: adding noise to the first sensitive feature in the first sample feature data to obtain the first desensitized data; or Performing a statistical perturbation operation on a second sensitive feature in the data chain in the first sample feature data to obtain the first desensitized data; or Multiple features in the first sample feature data are combined, and differential privacy processing is performed on the combined multiple features to obtain the first desensitized data.

4. The model joint processing method according to claim 2, characterized in that: Performing a feature name de-identification operation on the normalized desensitized data to obtain the first desensitized feature data includes: Extracting original feature names from the normalized desensitized data; The original feature name is mapped into a character string based on a hash mapping, or the original feature name is mapped into a random identifier based on a random mapping table to perform the feature name de-identification operation to obtain the first desensitized feature data.

5. The model joint processing method according to claim 2, characterized in that: Performing the encryption operation on the first sample identifier to obtain the first identifier ciphertext includes: receiving an oblivious pseudorandom function (OPRF) key sent by the second party, or generating the OPRF key through negotiation with the second party; The first sample identifier is encrypted based on the OPRF key to obtain the first identifier ciphertext.

6. The model joint processing method according to claim 1, characterized in that: Preprocessing the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier includes: Identifying abnormal points that deviate from a preset range among the same type of features of the original data of the first participant; Deleting the original data including the outlier, or performing a repair operation on the outlier to obtain cleaned data; A sample identifier corresponding to the cleaned data is determined, and the sample identifier is determined as a primary key, so as to obtain the first sample feature data and the corresponding first sample identifier based on an association relationship between the primary key and a foreign key.

7. A model joint processing method, characterized in that: Applicable to the second party, including: Preprocessing the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier; Performing a privacy enhancement operation on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and a corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext; Receive first two-tuple data sent by a first participant to transfer the first two-tuple data set and the second two-tuple data set into a security sandbox, and perform model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

8. The model joint processing method according to claim 7, characterized in that: Performing a model training and / or model prediction operation based on the first two-tuple data set and the second two-tuple data set within the security sandbox includes: In the security sandbox, the first two-tuple data set includes first desensitized feature data and a corresponding first identification ciphertext, and an intersection between the first identification ciphertext and the second identification ciphertext is calculated to obtain an intersection ciphertext; Extracting intersection desensitized feature data corresponding to the intersection ciphertext from the first desensitized feature data and the second desensitized feature data respectively; The model training and / or the model prediction operation is performed based on the intersection ciphertext and the corresponding intersection desensitized feature data.

9. The model joint processing method according to claim 7, characterized in that: Before transferring the first two-tuple data set and the second two-tuple data set into the security sandbox, the method further includes: Initializing the security sandbox and configuring a communication channel; Configuring an OPRF algorithm in the security sandbox; generating an OPRF key based on the OPRF algorithm in the security sandbox; The OPRF key is obtained from the security sandbox based on the communication channel.

10. The model joint processing method according to claim 7, characterized in that: Performing a privacy enhancement operation on the second sample feature data and the corresponding second sample identifier to obtain second desensitized feature data and a corresponding second identifier ciphertext, including: The privacy enhancement operation includes differential privacy desensitization processing and encryption operation, performing differential privacy desensitization processing on the second sample feature data to obtain second desensitized data; Performing a normalization operation on the second desensitized data to obtain normalized desensitized data; Performing a feature name de-identification operation on the normalized desensitized data to obtain the second desensitized feature data; The encryption operation is performed on the second sample identifier to obtain the second identifier ciphertext.

11. The model joint processing method according to claim 10, characterized in that: Performing differential privacy desensitization processing on the second sample feature data to obtain second desensitized data includes: adding noise to the first sensitive feature in the second sample feature data to obtain the second desensitized data; or Performing a statistical perturbation operation on a second sensitive feature in the data chain in the second sample feature data to obtain the second desensitized data; or Multiple features in the second sample feature data are combined, and the combined multiple features are subjected to differential privacy processing to obtain the second desensitized data.

12. The model joint processing method according to claim 10, characterized in that: Performing a feature name de-identification operation on the normalized desensitized data to obtain the second desensitized feature data includes: Extracting original feature names from the normalized desensitized data; The original feature name is mapped into a string based on a hash mapping, or the original feature name is mapped into a random identifier based on a random mapping table to perform the feature name de-identification operation to obtain the second desensitized feature data.

13. The model joint processing method according to claim 10, characterized in that: Performing the encryption operation on the second sample identifier to obtain the second identifier ciphertext includes: The second sample identifier is encrypted based on the obtained OPRF key to obtain the second identifier ciphertext.

14. The model joint processing method according to claim 7, characterized in that: Preprocessing the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier includes: Identifying abnormal points that deviate from a preset range among the same characteristics of the original data of the second party; Deleting the original data including the outlier, or performing a repair operation on the outlier to obtain cleaned data; A sample identifier corresponding to the cleaned data is determined, and the sample identifier is determined as a primary key, so as to obtain the second sample feature data and the corresponding second sample identifier based on an association relationship between the primary key and a foreign key.

15. A model joint processing device, characterized in that: Applicable to the first party, including: A first preprocessing module, configured to preprocess the original data of the first participant to obtain first sample feature data and a corresponding first sample identifier; a first privacy enhancement operation module, configured to perform a privacy enhancement operation on the first sample feature data and the corresponding first sample identifier, respectively, to obtain first desensitized feature data and a corresponding first identifier ciphertext, so as to construct a first two-tuple data set from the first desensitized feature data and the corresponding first identifier ciphertext; A sending module is used to send the first two-tuple data set to a second participant, so that the second participant transfers the first two-tuple data set and a local second two-tuple data set into a security sandbox, and performs model training and / or model prediction operations based on the first two-tuple data set and the second two-tuple data set within the security sandbox.

16. A model joint processing device, characterized in that: Applicable to the second party, including: A second preprocessing module, configured to preprocess the original data of the second participant to obtain second sample feature data and a corresponding second sample identifier; a second privacy enhancement operation module, configured to perform a privacy enhancement operation on the second sample feature data and the corresponding second sample identifier, respectively, to obtain second desensitized feature data and a corresponding second identifier ciphertext, so as to construct a second two-tuple data set from the second desensitized feature data and the corresponding second identifier ciphertext; A processing module is used to receive the first two-tuple data sent by the first participant, to transfer the first two-tuple data set and the second two-tuple data set into a security sandbox, and to perform model training and / or model prediction operations based on the second two-tuple data set and the second two-tuple data set within the security sandbox.

17. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the model joint processing method according to any one of claims 1 to 6 or 7 to 14 by executing the executable instructions.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model joint processing method according to any one of claims 1 to 14 is implemented.

19. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model joint processing method according to any one of claims 1 to 14 is implemented.