Data processing method and device and electronic equipment

By performing distributed parallel operations of sensitive information desensitizing data to be processed and using distributed information to be processed, the problem of data privacy protection in the public network AI model is solved, and the dual effects of efficient computing and privacy protection are achieved.

CN120337246APending Publication Date: 2025-07-18LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386864.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When using public network AI models for large-scale knowledge inference and data calculations, how to make full use of its computing power, effectively protect data privacy and avoid privacy data leakage.

Method used

The data to be processed is desensitized, the data is split into encrypted data blocks and distributed to multiple large models for distributed parallel operations, and data processing is performed in an encrypted state through fully homomorphic encryption, and the data results are finally restored.

Benefits of technology

It realizes efficient calculations on public network AI models while effectively protecting user privacy, avoiding data leakage, and ensuring data security and accuracy of processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337246A_ABST
    Figure CN120337246A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and electronic equipment, and relates to the technical field of artificial intelligence and data security, and the method comprises the steps: carrying out sensitive information desensitization processing on obtained first to-be-processed data, and obtaining desensitized second to-be-processed data corresponding to the first to-be-processed data; splitting the second to-be-processed data into at least two to-be-processed data blocks, and encrypting each to-be-processed data block to obtain at least two encrypted data blocks; distributing each encrypted data block to at least two large models for processing; the processing at least comprises the steps of maintaining the encryption state of the encrypted data block by each large model and performing data operation on the encrypted data block; and determining a target processing result corresponding to the first to-be-processed data according to a processing result of each large model on the encrypted data block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of artificial intelligence and data security, and particularly to a data processing method, apparatus, and electronic device. Background Art

[0002] With the development of AI (Artificial Intelligence) technology, enterprises widely adopt AI models for large-scale knowledge reasoning and data calculation. However, although public network AI models can reduce costs and shorten the construction cycle, they pose a serious risk of data privacy leakage. To solve the problem of privacy data security when using public network AI models, a solution that can both make full use of the computing power of public network AI models and effectively protect data privacy is needed. Summary of the Invention

[0003] For this reason, the present application discloses the following technical solutions:

[0004] A data processing method includes:

[0005] Performing sensitive information desensitization processing on the obtained first data to be processed to obtain the desensitized second data to be processed corresponding to the first data to be processed;

[0006] Splitting the second data to be processed into at least two data blocks to be processed, and respectively encrypting each data block to be processed to obtain at least two encrypted data blocks;

[0007] Distributing each encrypted data block to at least two large models for processing; the processing at least includes each large model maintaining the encrypted state of the encrypted data block and performing data operations on the encrypted data block;

[0008] Determining the target processing result corresponding to the first data to be processed according to the processing results of each large model on the encrypted data block.

[0009] Optionally, the performing sensitive information desensitization processing on the obtained first data to be processed includes:

[0010] Identifying each sensitive entity information in the first data to be processed;

[0011] Replacing each sensitive entity information in the first data to be processed with assumed entity information that satisfies an association condition with the sensitive entity information; the assumed entity information does not contain sensitive information.

[0012] Optionally, splitting the second data to be processed into at least two data blocks to be processed includes:

[0013] Split the second data to be processed into at least two semantic units with complete semantics; the at least two data blocks to be processed include the at least two semantic units.

[0014] Optionally, the encrypting each of the data blocks to be processed includes:

[0015] Encrypt each of the data blocks to be processed based on a target encryption method; the target encryption method allows data operations to be performed on the encrypted data block while maintaining the encrypted state of the encrypted data block.

[0016] Optionally, before distributing each of the encrypted data blocks to at least two large models for processing, the method further includes:

[0017] Construct an index for each encrypted data block;

[0018] Determine at least two first data units to be distributed according to the index constructed for each encrypted data block; each first data unit includes an encrypted data block and the index corresponding to the encrypted data block;

[0019] The distributing each of the encrypted data blocks to at least two large models for processing includes:

[0020] Distribute each of the first data units to at least two large models, so that each large model performs distributed parallel operation processing on the encrypted data block in the obtained first data unit.

[0021] Optionally, before determining the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models, the method further includes:

[0022] Obtain each second data unit generated by each of the large models based on data operations performed on the obtained encrypted data blocks; each second data unit includes the data operation result of the large model on the corresponding encrypted data block and the index corresponding to the corresponding encrypted data block;

[0023] Wherein, the processing result of the large model on the encrypted data block includes the corresponding second data unit.

[0024] Optionally, the determining the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models includes:

[0025] Decrypt the data operation result in each second data unit to obtain a decrypted operation result;

[0026] Integrate the decryption operation results according to the indexes respectively corresponding to the respective decryption operation results to obtain an integration result, and restore the desensitized data in the integration result to the corresponding sensitive information to obtain the target processing result corresponding to the first data to be processed; or, restore the desensitized data in each decryption operation result to the corresponding sensitive information, and integrate the decryption operation results with the sensitive information restored according to the corresponding indexes to obtain the target processing result corresponding to the first data to be processed.

[0027] Optionally, before integrating the decryption operation results or restoring the desensitized data in each decryption operation result to the corresponding sensitive information, the method further includes:

[0028] Perform integrity verification on each of the decryption operation results, and trigger the step of integrating the decryption operation results or trigger the step of restoring the desensitized data in each decryption operation result to the corresponding sensitive information when the verification is passed.

[0029] A data processing device includes:

[0030] A desensitization module, configured to perform sensitive information desensitization processing on the obtained first data to be processed to obtain the desensitized second data to be processed corresponding to the first data to be processed;

[0031] A splitting and encryption module, configured to split the second data to be processed into at least two data blocks to be processed, and encrypt each of the data blocks to be processed to obtain at least two encrypted data blocks;

[0032] A distribution module, configured to distribute each of the encrypted data blocks to at least two large models for processing; the processing at least includes that each of the large models maintains the encrypted state of the encrypted data block and performs data operations on the encrypted data block;

[0033] A determination module, configured to determine the target processing result corresponding to the first data to be processed according to the processing results of each of the large models on the encrypted data block.

[0034] An electronic device includes:

[0035] A memory, configured to store at least a set of computer instruction sets;

[0036] A processor, configured to implement any of the data processing methods as described above by executing the instruction sets stored in the memory.

[0037] A storage medium, the storage medium carries one or more computer instruction sets, and when the one or more computer instruction sets are executed by an electronic device, the electronic device can be enabled to implement any of the data processing methods as described above. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0039] Figure 1 is a schematic flowchart of a data processing method provided by the present application;

[0040] Figure 2 is another schematic flowchart of a data processing method provided by the present application;

[0041] Figure 3 is a schematic diagram of data processing between the enterprise side and the public network large model side in an application example provided by the present application;

[0042] Figure 4 is a composition structure diagram of a data processing device provided by the present application;

[0043] Figure 5 is a composition structure diagram of an electronic device provided by the present application. Detailed implementation manners

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0045] The embodiments of the present application provide a data processing method, device and electronic device, which are used to protect privacy data through a distributed encryption calculation method, ensuring that both the distributed computing power of the public network large model can be effectively utilized and the user's privacy data can be fully secured.

[0046] The data processing method can be applied to electronic devices in many general or special computing device environments or configurations, such as: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor devices, and so on.

[0047] See Figure 1 the method flowchart shown. The data processing method provided by the embodiments of the present application may include the following steps 101 to 104, and the following will describe these steps in detail respectively.

[0048] Step 101: Perform sensitive information desensitization processing on the obtained first data to be processed, and obtain the desensitized second data to be processed corresponding to the first data to be processed.

[0049] Optionally, the first data to be processed is text data, which can specifically be, but is not limited to, text data for which a user desires text analysis, data classification, or model prediction (such as probabilistic or numerical speculation on future events, trends, or results based on historical data or input information), etc.

[0050] In practical applications, the first data to be processed can be text data to be processed in the original data format of text, or it can also be text data obtained by converting non - text format data to be processed, such as images, audio, and / or video.

[0051] Exemplarily, for instance, the first data to be processed is a contract to be signed by an enterprise. More specifically, for example, an enterprise hopes to use a public - network AI large - model to analyze adverse factors in the contract, especially to check the reasonableness of the transaction amount, the fairness of the terms, and potential default risks, etc. In this application scenario, the contract document can be used as the first data to be processed.

[0052] With the development of generative AI technology, users such as enterprises widely adopt large - models such as large - language models for large - scale knowledge reasoning and data calculation. In order to fully utilize the computing power of public - network AI models while effectively protecting the data privacy of users, after obtaining the first data to be processed, the embodiments of this application first perform sensitive information desensitization processing on it to ensure that the data sent to the public - network large - model does not contain any sensitive information and to ensure the security of user privacy.

[0053] Among them, the sensitive information desensitization processing performed on the first data to be processed is used to replace the sensitive information in the first data to be processed with non - sensitive information, while ensuring the structural and semantic consistency between the replaced non - sensitive information and the replaced sensitive information, so as to avoid affecting the accuracy of subsequent data processing results.

[0054] When performing sensitive information desensitization processing on the first data to be processed, optionally, each sensitive entity information in the first data to be processed can be first identified, and then each sensitive entity information in the first data to be processed is replaced with assumed entity information that satisfies the association condition with the sensitive entity information; the assumed entity information does not contain sensitive information.

[0055] In implementation, each sensitive entity information in the first data to be processed can be identified through manual annotation and / or automatic recognition methods.

[0056] In the automatic recognition method, optionally, a sensitive entity information template can be predefined in advance, and various types of sensitive entity information can be recognized based on the corresponding template. For example, templates for various types of sensitive entity information such as names, mobile phone numbers, and company addresses can be formulated for entity recognition of names, mobile phone numbers, company addresses, etc. Taking mobile phone numbers and company addresses as specific examples, a mobile phone number is usually a combination of a predetermined number of digits, and even optional combinations of the first few digits can be specified. Therefore, in the template of the mobile phone number, the character type can be required to be digits, and the length of the digit sequence (i.e., the number of digits included) and the optional digit combinations of the first few digits can be restricted; address information is usually divided according to a certain hierarchical structure, specifically including multi-level geographical location information divided according to a certain hierarchical structure. Therefore, the template of the company address can include multiple hierarchical structures of geographical locations, such as provinces, cities, districts / counties, roads, etc.

[0057] In the automatic recognition method, it is also possible to recognize each sensitive entity information in the first data to be processed based on a pre-trained small model or by using a large model. There is no restriction on this, and it can be determined according to the actual application situation.

[0058] Optionally, the association condition may mean that the assumed entity information after replacement has the same type as the sensitive entity information before replacement, and the composition structure is the same or similar, so as to ensure the consistency of the sensitive entity information and the assumed entity information after replacement in terms of structure and semantics. For example, replacing the user name in the contract with a pseudonym of the same type and structure (including "surname" and "name"), replacing the mobile phone number with a fake mobile phone number with the same number of digits, and replacing the company address with a fake address with the same geographical location hierarchical structure, etc., so as to make the executed desensitization operation not damage the structure and semantic type of the original sensitive information as much as possible and avoid affecting the accuracy of the processing result of the first data to be processed.

[0059] In implementation, pseudonymization processing can be introduced. Through pseudonymization processing, assumed entity information associated with the original sensitive entity information in the first data to be processed (i.e., satisfying the above association condition) is generated to ensure the logical consistency of the data information before and after replacement, and at the same time enhance the privacy protection of the first data to be processed.

[0060] Step 102: Split the second data to be processed into at least two data blocks to be processed, and encrypt each data block to be processed to obtain at least two encrypted data blocks.

[0061] After performing sensitive information desensitization processing on the first data to be processed to obtain the desensitized second data to be processed, enter the data splitting and encryption stage.

[0062] In data splitting processing, optionally, natural language processing (NLP) algorithms or context-based word segmentation algorithms can be specifically used to split the second data to be processed into at least two semantic units with complete semantics.

[0063] The at least two data blocks to be processed include the at least two semantic units. That is, each semantic unit obtained by splitting can be regarded as a data block to be processed, and each semantic unit has complete semantics and is used as the smallest processing unit in subsequent data processing.

[0064] The semantic unit can be, but is not limited to, an independent paragraph, sentence or key phrase. For example, a contract document can be split into multiple independent paragraphs, sentences or key phrases, etc.

[0065] After splitting the second data to be processed into at least two data blocks to be processed, each data block to be processed can be further encrypted based on the target encryption method. For example, each semantic unit obtained by splitting is encrypted based on the target encryption method. The target encryption method allows data operations to be performed on the encrypted data block while maintaining the encrypted state of the encrypted data block.

[0066] Optionally, the target encryption method can be a fully homomorphic encryption (Homomorphic Encryption) method.

[0067] Fully homomorphic encryption allows data processing in the data encryption state. Based on the fully homomorphic encryption method in the embodiments of the present application, the public network large model can directly process the encrypted data block without decrypting the encrypted data block, and enterprise users and other clients can decrypt the processing result returned by the large model to restore the final processing result corresponding to the data.

[0068] The sensitive information desensitization processing in this step can be regarded as a preprocessing step in the subsequent data processing process.

[0069] Step 103: Distribute each of the encrypted data blocks to at least two large models for processing; the processing at least includes each of the large models maintaining the encrypted state of the encrypted data block and performing data operations on the encrypted data block.

[0070] After obtaining at least two encrypted data blocks corresponding to the second data to be processed, enter the data distribution stage of the encrypted data block.

[0071] In the data distribution stage, each encrypted data block corresponding to the second data to be processed can be distributed to at least two large models for processing based on a certain distribution strategy.

[0072] Specifically, but not limited to, based on a random distribution strategy, distribute each encrypted data block to at least two large model instances deployed on a public network such as a public cloud, and perform distributed parallel computing processing on the encrypted data blocks based on the at least two large model instances, so as to make full use of the distributed parallel computing power of the public network large model. At the same time, since the large model processes the encrypted data blocks while maintaining the encrypted state, it also protects the user privacy, and can achieve the dual effects of privacy protection and efficient data computing.

[0073] In implementation, it is preferable to adopt a random distribution strategy to distribute the encrypted data blocks by using a distributed routing algorithm. Exemplarily, use the Consistent Hashing algorithm or a similar distributed algorithm to randomly distribute each encrypted data block to multiple large model instances on the public network for distributed parallel computing. This distribution method can effectively balance the load, prevent the encrypted data blocks from concentrating on a single large model instance on the public network, further improve the data processing efficiency and reduce the risk of data leakage.

[0074] After each large model (such as a large model instance) obtains the corresponding distributed encrypted data block, it maintains the encrypted state of the encrypted data block and performs corresponding data operations on the encrypted data block. The data operations performed by the large model on the encrypted data block can be, but not limited to, data operations related to data analysis, classification, or model prediction, etc. There is no limitation on this, depending on the actual data processing requirements.

[0075] Step 104: Determine the target processing result corresponding to the first data to be processed according to the processing results of each large model on the encrypted data block.

[0076] After each large model completes the processing of the encrypted data block and obtains the processing result of the encrypted data block, an enterprise or other user side can obtain the processing results of each large model on the encrypted data block respectively. Optionally, the user side can directly receive the processing results of each large model on the encrypted data block, or, under the condition of meeting the request conditions, send a processing result acquisition request to each large model to obtain the processing results of the encrypted data block from each large model through the initiated processing result acquisition request.

[0077] The request conditions can be, but not limited to, reaching a pre-set processing result acquisition time, or receiving a use / view request from the service side for the target processing result corresponding to the first data to be processed.

[0078] After obtaining the processing results of each large model on the encrypted data block respectively, the target processing result corresponding to the first data to be processed can be determined by performing corresponding processing on the processing results of each encrypted data block.

[0079] Among them, the corresponding processing performed on the processing results of each encrypted data block includes, but is not limited to, decrypting, integrating, and restoring sensitive information from the processing results of each encrypted data block. This part of the content will be described in detail in the following embodiments.

[0080] In summary, the data processing method provided by the embodiments of the present application performs sensitive information desensitization processing on the obtained first data to be processed, obtains the second data to be processed after desensitization, splits the second data to be processed into at least two data blocks to be processed, encrypts each data block to be processed separately, and distributes each encrypted data block obtained by encryption to at least two large models for processing. Finally, according to the processing results of each encrypted data block by each large model, the target processing result corresponding to the first data to be processed is determined. Among them, the large model maintains the encrypted state of the encrypted data block and performs data operations on the encrypted data block. Therefore, the present application can not only make full use of the computing power of the public network AI model for fast and efficient data processing (for example, making full use of the parallel computing power of the public network large model for distributed parallel operations), but also avoid the leakage of private data on the public network, effectively protect the user privacy, and ensure the security of user data.

[0081] In an alternative embodiment, referring to Figure 2 the flowchart of the data processing method shown, before step 103, the data processing method provided by the present application may further include the following processing:

[0082] Step 201: Construct an index for each encrypted data block.

[0083] Specifically, but not limited to, respectively constructing a number or serial number for each encrypted data block as the index of each encrypted data block. Different encrypted data blocks correspond to different unique indexes, which are used to uniquely identify each encrypted data block based on the corresponding index.

[0084] After constructing an index for each encrypted data block, optionally, each encrypted data block and its corresponding index can be stored at the user side. For example, each encrypted data block and its corresponding number / serial number are stored in a secure database at the enterprise side for convenient subsequent further processing.

[0085] Step 202: Determine at least two first data units to be distributed according to the index constructed for each encrypted data block.

[0086] Based on step 201, each encrypted data block and its corresponding index such as a number or serial number can be further determined as a first data unit. That is, each first data unit includes an encrypted data block and the index corresponding to the encrypted data block.

[0087] Correspondingly, referring to Figure 2, in this embodiment, step 103 in the method provided by this application, that is, distributing each of the encrypted data blocks to at least two large models for processing, can be implemented as:

[0088] Step 203, distribute each of the first data units to at least two large models, so that each of the large models performs distributed parallel computing processing on the encrypted data blocks in the obtained first data units respectively.

[0089] After obtaining each first data unit, each first data unit is used as an object to be distributed and distributed to at least two large models. Then, each large model extracts the encrypted data blocks from the obtained first data units, and performs data operations on the encrypted data blocks while maintaining the encrypted state of the encrypted data blocks, and obtains the data operation results corresponding to the encrypted data blocks. On this basis, the large model assembles the data operation result and the index corresponding to each encrypted data block to obtain the second data unit corresponding to each encrypted data block, and returns the second data unit corresponding to the encrypted data block to the user side as the processing result of the encrypted data block. Wherein, each second data unit includes the data operation result of the large model on the corresponding encrypted data block and the index corresponding to the corresponding encrypted data block, so as to facilitate the subsequent user side to perform integration processing on the decryption operation results corresponding to the data operation results in each second data unit according to the indexes in each second data unit.

[0090] After receiving the first data units, each large model on the public network can specifically use a parallel computing framework (such as MapReduce or Spark) to independently calculate the encrypted data blocks in each of the first data units. The calculations here can include but are not limited to text analysis, data classification, model prediction, etc. Each calculation task runs on an independent large model instance, and after the calculation is completed, the processing result (the second data unit) is returned to the user side such as an enterprise.

[0091] Corresponding to the above implementation process, refer to Figure 2 , in this embodiment, step 104 in the method provided by this application, that is, determining the target processing result corresponding to the first data to be processed according to the processing results of each of the large models on the encrypted data blocks, can be implemented as:

[0092] Step 204, decrypt the data operation result in each second data unit to obtain the decryption operation result.

[0093] User sides such as enterprises can first obtain each second data unit generated by each large model through data operations on the encrypted data blocks.

[0094] Since the large model performs data operations on encrypted data blocks while maintaining their encrypted state, that is, directly performing operations on encrypted data blocks without decrypting them, and thus matching them, the data operation results of the large model on encrypted data blocks are essentially in ciphertext form. Based on this, after the enterprise and other client ends obtain the second data units returned by each large model, they need to decrypt the data operation results in each second data unit, for example, by performing the reverse process of data block encryption to achieve decryption, etc., so as to obtain the corresponding decrypted operation results.

[0095] Step 205: Integrate the decrypted operation results according to the indexes respectively corresponding to the decrypted operation results to obtain an integration result, and restore the desensitized data in the integration result to the corresponding sensitive information to obtain the target processing result corresponding to the first data to be processed; or, restore the desensitized data in each decrypted operation result to the corresponding sensitive information, and integrate the decrypted operation results with the sensitive information restored according to the corresponding indexes to obtain the target processing result corresponding to the first data to be processed.

[0096] After decrypting the data operation results in each second data unit to obtain each decrypted operation result, the enterprise and other client ends can further perform integration and sensitive information restoration processing on each decrypted operation result to achieve data reconstruction, so as to restore the original order and correlation relationship between each data segment (each decrypted operation result), and restore the corresponding sensitive information, thereby obtaining the target processing result corresponding to the first data to be processed.

[0097] Among them, the execution order of the integration and the sensitive information restoration processing is not limited, and either of the two can be processed first and the other later.

[0098] Among them, if the integration processing is performed first, the decrypted operation results can be directly integrated according to the indexes respectively corresponding to the decrypted operation results, that is, the decrypted operation results are reorganized into a whole according to the order between the indexes to obtain an integration result. On this basis, the desensitized data in the integration result can be further replaced with the corresponding sensitive information, and this processing process is the reverse process of desensitization processing. Optionally, specifically, the assumed entity information in the integration result can be replaced with the corresponding sensitive entity information according to the entity mapping table between the sensitive entity information and the assumed entity information, for example, replacing the false mobile phone number with the corresponding real mobile phone number, replacing the false enterprise address with the corresponding real enterprise address, etc., so as to restore the desensitized data in the integration result to the corresponding sensitive information, and the integration result after completing the sensitive information restoration can be used as the final target processing result corresponding to the first data to be processed.

[0099] The entity mapping table may be an information table of the mapping relationship between sensitive entity information and assumed entity information pre-established for desensitization processing and sensitive information restoration processing, or it may also be formed by recording the mapping relationship between the generated assumed entity information and its corresponding sensitive entity information during desensitization processing. There is no limitation on this, and it can be determined according to the actual application situation.

[0100] If the sensitive information restoration processing is performed first, each decryption operation result needs to be used as the restoration object, and the sensitive information restoration processing is performed on each decryption operation result one by one. For example, based on the entity mapping table, the assumed entity information in each decryption operation result is replaced with the corresponding sensitive entity information, etc. Then, according to the corresponding index, the decryption operation results that have completed sensitive information restoration are integrated, and they are recombined into a whole according to the index, so as to obtain the target processing result corresponding to the first data to be processed.

[0101] In this embodiment, by using the distributed parallel processing ability of the public network large model for data processing, the data processing efficiency is improved, and at the same time, the risk of a single large model leaking complete data is reduced. Moreover, by introducing encryption and decryption processing of data in the distributed parallel processing of data based on the public network large model, the AI large model can directly perform calculations without decrypting the data, which can avoid the leakage of user privacy on the public network, thus effectively protecting user privacy and ensuring the security of user data, and having the dual effects of efficient data calculation and privacy protection.

[0102] In an optional embodiment, for the data processing method provided in this application, before integrating each decryption operation result or restoring the desensitized data in each decryption operation result to the corresponding sensitive information, the following processing may also be included:

[0103] Perform integrity verification on each decryption operation result, and trigger the step of integrating each decryption operation result or the step of restoring the desensitized data in each decryption operation result to the corresponding sensitive information when the verification is passed.

[0104] Among them, integrity verification of each decryption operation result can be performed by, but not limited to, using hash verification or an integrity verification algorithm based on a Merkle tree to ensure the data integrity and consistency of each decryption operation result. If the verification is passed, the step of integrating each decryption operation result or the step of restoring the desensitized data in each decryption operation result to the corresponding sensitive information can be triggered to determine the target processing result corresponding to the first data to be processed. On the contrary, if the verification fails, for example, if it is found that the data is missing or tampered with, an alarm mechanism can be triggered to prevent the spread of errors.

[0105] In this embodiment, before integrating the results of each decryption operation or restoring the desensitized data in the results of each decryption operation to the corresponding sensitive information, integrity verification is performed on the results of each decryption operation, so as to timely detect data loss or tampering, and trigger an alarm when data loss or tampering is found, so as to prevent the spread of errors.

[0106] See Figure 3 , an application example of the method of this application is provided. Taking the user as an enterprise as an example, this example provides a schematic diagram of data processing between the enterprise side and a public network large model such as a public cloud large model side. Among them, after the enterprise side obtains the data to be processed in the request, the data is desensitized in the preprocessing stage. For example, the sensitive entity information in the data is replaced with assumed entity information through pseudonymization processing. After that, the desensitized data is successively processed such as data splitting, encryption, index construction, and data distribution based on a distributed routing algorithm, so as to distribute each encrypted data block corresponding to the data that has been desensitized and encrypted to multiple large models in the public network for distributed parallel processing. On this basis, the enterprise side receives the data block processing results of each large model in the public network, and performs a series of processing such as data decryption, verification, integration, and desensitization on them, so as to realize data reconstruction and obtain the target processing result corresponding to the data to be processed in the request.

[0107] For a more detailed data processing process between the enterprise side and the large model side, reference can be made to the descriptions of the above method embodiments, which will not be elaborated here.

[0108] This example improves the data operation efficiency by using multiple large models in the public network to perform distributed parallel processing on data, and at the same time reduces the risk of a single large model leaking complete data. And by introducing the data encryption and decryption process in data processing, the privacy and security of the data are further strengthened, which can avoid the leakage of user privacy in the public network, effectively protect user privacy, and ensure the security of user data.

[0109] Corresponding to the above data processing method, an embodiment of this application also provides a data processing device, the composition structure of which is as Figure 4 shown, including:

[0110] A desensitization module 401, configured to perform sensitive information desensitization processing on the obtained first data to be processed, and obtain the second data to be processed that has been desensitized corresponding to the first data to be processed;

[0111] A splitting and encryption module 402, configured to split the second data to be processed into at least two data blocks to be processed, and encrypt each of the data blocks to be processed to obtain at least two encrypted data blocks;

[0112] A distribution module 403, configured to distribute each of the encrypted data blocks to at least two large models for processing; the processing at least includes each of the large models maintaining the encrypted state of the encrypted data block and performing data operations on the encrypted data block;

[0113] A determination module 404, configured to determine a target processing result corresponding to the first data to be processed according to the processing results of each of the large models on the encrypted data block.

[0114] In an optional implementation manner, the desensitization module 401 is specifically configured to:

[0115] Identify each sensitive entity information in the first data to be processed;

[0116] Replace each sensitive entity information in the first data to be processed with a hypothetical entity information that satisfies an association condition with the sensitive entity information; the hypothetical entity information does not contain sensitive information.

[0117] In an optional implementation manner, when splitting the second data to be processed into at least two data blocks to be processed, the splitting and encryption module 402 is specifically configured to:

[0118] Split the second data to be processed into at least two semantic units with complete semantics; the at least two data blocks to be processed include the at least two semantic units.

[0119] In an optional implementation manner, when encrypting each of the data blocks to be processed respectively, the splitting and encryption module 402 is specifically configured to:

[0120] Encrypt each of the data blocks to be processed respectively based on a target encryption method; the target encryption method allows data operations to be performed on the encrypted data block while maintaining the encrypted state of the encrypted data block.

[0121] In an optional implementation manner, the above device further includes an index construction module, configured to construct an index for each encrypted data block;

[0122] The determination module 404 is further configured to: determine at least two first data units to be distributed according to the index constructed for each encrypted data block; each first data unit includes an encrypted data block and the index corresponding to the encrypted data block;

[0123] The distribution module 403 is specifically configured to: distribute each of the first data units to at least two large models, so that each of the large models performs distributed parallel computing processing on the encrypted data blocks in the obtained first data units respectively.

[0124] In an alternative embodiment, the above device further includes an acquisition module, configured to obtain, before determining the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models, each second data unit generated by each of the large models based on data operations performed on the obtained encrypted data blocks; each second data unit includes the data operation result of the large model on the corresponding encrypted data block and the index corresponding to the corresponding encrypted data block.

[0125] Wherein, the processing result of the large model on the encrypted data block includes the corresponding second data unit.

[0126] In an alternative embodiment, when determining module 404 determines the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models, it is specifically configured to:

[0127] Decrypt the data operation result in each second data unit to obtain a decrypted operation result;

[0128] Integrate the decrypted operation results according to the indexes respectively corresponding to the decrypted operation results to obtain an integration result, and restore the desensitized data in the integration result to the corresponding sensitive information to obtain the target processing result corresponding to the first data to be processed; or, restore the desensitized data in each decrypted operation result to the corresponding sensitive information, and integrate the decrypted operation results with the sensitive information restored according to the corresponding indexes to obtain the target processing result corresponding to the first data to be processed.

[0129] In an alternative embodiment, the above device further includes a verification module, configured to:

[0130] Before integrating the decrypted operation results or restoring the desensitized data in each decrypted operation result to the corresponding sensitive information, perform integrity verification on each decrypted operation result, and trigger the step of integrating the decrypted operation results or trigger the step of restoring the desensitized data in each decrypted operation result to the corresponding sensitive information when the verification is passed.

[0131] An embodiment of the present application also discloses an electronic device, the composition structure of the electronic device is as Figure 5 shown, and at least includes:

[0132] A memory 10, configured to store a computer instruction set;

[0133] The computer instruction set can be implemented in the form of a computer program.

[0134] A processor 20, configured to implement the data processing method provided in any of the above method embodiments by executing the computer instruction set in the memory.

[0135] The processor 20 may be a Central Processing Unit (CPU), an application-specific integrated circuit (ASIC), a Digital Signal Processor (DSP), an ASIC, a Field Programmable Gate Array (FPGA), a Neural Network Processor (NPU), a Deep Learning Processor (DPU), or other programmable logic devices, etc.

[0136] Optionally, the electronic device further includes storage resources such as memory and cache.

[0137] Optionally, the electronic device further includes a camera component and / or is connected to an external camera component.

[0138] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0139] The communication interface is used for communication between the electronic device and other devices. The communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.

[0140] The embodiment of the present application also discloses a storage medium. The storage medium carries one or more computer instruction sets. When the one or more computer instruction sets are executed by the electronic device, the electronic device can implement the data processing method provided in any of the above method embodiments.

[0141] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0142] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for description. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0143] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a creative contribution, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0144] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0145] The above are only the preferred embodiments of this application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A data processing method, comprising: Performing sensitive information desensitization processing on the obtained first data to be processed to obtain the desensitized second data to be processed corresponding to the first data to be processed; Splitting the second data to be processed into at least two data blocks to be processed, and respectively encrypting each of the data blocks to be processed to obtain at least two encrypted data blocks; Distributing each of the encrypted data blocks to at least two large models for processing; the processing at least includes each of the large models maintaining the encrypted state of the encrypted data block and performing data operations on the encrypted data block; Determining the target processing result corresponding to the first data to be processed according to the processing results of each of the large models on the encrypted data block.

2. The data processing method according to claim 1, wherein the performing sensitive information desensitization processing on the obtained first data to be processed includes: Identifying each sensitive entity information in the first data to be processed; Replacing each sensitive entity information in the first data to be processed with assumed entity information that satisfies an association condition with the sensitive entity information; The assumed entity information does not contain sensitive information.

3. The data processing method according to claim 1, wherein splitting the second data to be processed into at least two data blocks to be processed includes: Splitting the second data to be processed into at least two semantic units with complete semantics; The at least two data blocks to be processed include the at least two semantic units.

4. The data processing method according to claim 1, wherein the respectively encrypting each of the data blocks to be processed includes: Respectively encrypting each of the data blocks to be processed based on a target encryption method; The target encryption method allows data operations to be performed on the encrypted data block while maintaining the encrypted state of the encrypted data block.

5. The data processing method according to claim 1, before distributing each of the encrypted data blocks to at least two large models for processing, further comprising: Constructing an index for each encrypted data block; Determining at least two first data units to be distributed according to the index constructed for each encrypted data block; Each first data unit includes an encrypted data block and the index corresponding to the encrypted data block; The distributing each of the encrypted data blocks to at least two large models for processing includes: Distributing each of the first data units to at least two large models, so that each of the large models respectively performs distributed parallel operation processing on the encrypted data blocks in the obtained first data units.

6. The data processing method according to claim 5, before determining the target processing result corresponding to the first data to be processed according to the processing results of each of the large models on the encrypted data block, further comprising: Obtaining each second data unit generated by each of the large models respectively based on performing data operations on the obtained encrypted data blocks; each second data unit includes the data operation result of the large model on the corresponding encrypted data block and the index corresponding to the corresponding encrypted data block; Wherein, the processing result of the large model on the encrypted data block includes the corresponding second data unit.

7. The data processing method according to claim 6, wherein determining the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models comprises: Decrypting the data operation result in each second data unit to obtain a decrypted operation result; Integrating the decrypted operation results according to the indexes respectively corresponding to the decrypted operation results to obtain an integration result, and restoring the desensitized data in the integration result to the corresponding sensitive information to obtain the target processing result corresponding to the first data to be processed; or restoring the desensitized data in each decrypted operation result to the corresponding sensitive information, and integrating the decrypted operation results with the sensitive information restored according to the corresponding indexes to obtain the target processing result corresponding to the first data to be processed.

8. The data processing method according to claim 7, further comprising, before integrating the decrypted operation results or restoring the desensitized data in the decrypted operation results to the corresponding sensitive information: Performing integrity verification on each of the decrypted operation results, and triggering the step of integrating the decrypted operation results or triggering the step of restoring the desensitized data in the decrypted operation results to the corresponding sensitive information when the verification is passed.

9. A data processing device, comprising: A desensitization module, configured to perform sensitive information desensitization processing on the obtained first data to be processed to obtain the desensitized second data to be processed corresponding to the first data to be processed; A splitting and encryption module, configured to split the second data to be processed into at least two data blocks to be processed, and encrypt each of the data blocks to be processed to obtain at least two encrypted data blocks; A distribution module, configured to distribute each of the encrypted data blocks to at least two large models for processing; the processing at least includes that each of the large models maintains the encrypted state of the encrypted data block and performs data operations on the encrypted data block; A determination module, configured to determine the target processing result corresponding to the first data to be processed according to the processing results of the encrypted data blocks by each of the large models.

10. An electronic device, comprising: A memory, configured to store at least one set of computer instruction sets; A processor, configured to implement the data processing method according to any one of claims 1-8 by executing the instruction sets stored in the memory.