Large model training method and device based on trusted data space, equipment and medium

By receiving encrypted data in a trusted data space and dividing the model into plaintext and ciphertext domains, the training method solves the problems of insufficient data sources and security in large model training, and improves data security and training efficiency.

CN120952203AActive Publication Date: 2025-11-14LINGSHU TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511064141.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In large-scale model training, a single data source is insufficient to support model training, requiring the integration of multiple data sources. Furthermore, how to improve training efficiency and security while ensuring data security has become an urgent problem to be solved.

Method used

The encrypted model training data is received using a trusted data space. The large model to be trained is divided into a plaintext domain model and a ciphertext domain model. The plaintext domain model is trained in the plaintext computing environment of the data user, while the ciphertext domain model is trained in the trusted data space. The training process of the ciphertext domain model is then transferred to the plaintext computing environment of the plaintext domain model.

Benefits of technology

This ensures data security while improving the training efficiency of large models, preventing data users from accessing the original training data and thus enhancing model training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952203A_ABST
    Figure CN120952203A_ABST
Patent Text Reader

Abstract

The invention discloses a large model training method, device and equipment based on a trusted data space and a medium, and the large model training method based on the trusted data space comprises the steps: receiving model training data from at least two data providers, and storing the model training data to the trusted data space; dividing a preset to-be-trained large model into a plaintext domain model and a ciphertext domain model; based on a model fine tuning method, training the to-be-trained large model according to the model training data, and transferring a training process of the ciphertext domain model to a plaintext computing power environment to which the plaintext domain model belongs; and taking the trained to-be-trained large model as a target large model. Through the technical scheme, the data security in the training process is ensured, and meanwhile, the training efficiency of the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for training large models based on a trusted data space. Background Technology

[0002] In the training of large models, fine-tuning of vertical-category large models is a core technology for applying the capabilities of large models to industries.

[0003] During the fine-tuning of large-scale vertical models, data from a single data source is insufficient to support the fine-tuning of the large model. It is necessary to combine data from multiple data sources for training the vertical model. For example, when training a large-scale medical consultation model, data from a single hospital is insufficient to support the training of the model. It is necessary to integrate data from hospitals across the province or even the country to achieve the required data scale for model training.

[0004] Therefore, how to train large models while ensuring data security, and how to improve model training efficiency and the security of the training process, have become urgent problems to be solved. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for training large models based on a trusted data space, to ensure data security and model training efficiency during the training process.

[0006] According to one aspect of the present invention, a method for training large models based on a trusted data space is provided, the method comprising:

[0007] Receive model training data from at least two data providers and store the model training data in a trusted data space; wherein the model training data is encrypted data provided by the data providers.

[0008] The pre-defined large model to be trained is divided into a plaintext domain model and a ciphertext domain model; wherein, the plaintext domain model is deployed in the plaintext computing environment of the data user, and the ciphertext domain model is deployed in the trusted data space of the data user; the computing power level of the plaintext computing environment is greater than that of the trusted data space.

[0009] Based on the model fine-tuning method, the large model to be trained is trained according to the model training data, and the training process of the ciphertext domain model is transferred to the plaintext computing environment to which the plaintext domain model belongs.

[0010] The trained large model is used as the target large model.

[0011] According to another aspect of the present invention, a large model training apparatus based on a trusted data space is provided, the apparatus comprising:

[0012] A data receiving module is used to receive model training data from at least two data providers and store the model training data in a trusted data space; wherein the model training data is encrypted data provided by the data providers.

[0013] The model deployment module is used to divide the preset large model to be trained into a plaintext domain model and a ciphertext domain model; wherein, the plaintext domain model is deployed in the plaintext computing environment of the data user, and the ciphertext domain model is deployed in the trusted data space of the data user; the computing power level of the plaintext computing environment is greater than that of the trusted data space.

[0014] The model training module is used to train the large model to be trained based on the model fine-tuning method and the model training data, and to transfer the training process of the ciphertext domain model to the plaintext computing environment to which the plaintext domain model belongs.

[0015] The target model generation module is used to take the trained large model to be trained as the target large model.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor;

[0018] and a memory communicatively connected to the at least one processor;

[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the large model training method based on a trusted data space as described in any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the large model training method based on a trusted data space as described in any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the large model training method based on a trusted data space as described in any embodiment of the present invention.

[0022] The technical solution of this invention involves the data user receiving encrypted model training data through a trusted data space and training the large model to be trained on the basis of the trusted data space. This ensures that the data user cannot access the original model training data during the training process, thus guaranteeing data security. Furthermore, during the training of the large model to be trained, the training process of the ciphertext domain model is transferred to the plaintext computing environment of the plaintext domain model, thereby improving the training efficiency of the large model.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1a This is a flowchart of a large model training method based on a trusted data space provided in Embodiment 1 of the present invention;

[0026] Figure 1b This is a schematic diagram illustrating the scenario between a data user and a data provider as provided in Embodiment 1 of the present invention;

[0027] Figure 2 This is a flowchart of a large model training method based on a trusted data space according to Embodiment 2 of the present invention;

[0028] Figure 3 This is a flowchart of a large model training method based on a trusted data space provided in Embodiment 3 of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of a large model training device based on a trusted data space according to Embodiment 4 of the present invention;

[0030] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the large model training method based on a trusted data space according to an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] Example 1

[0034] Figure 1a This document provides a flowchart of a large model training method based on a trusted data space, as described in Embodiment 1 of the present invention. This embodiment is applicable to training large models. The method can be executed by a large model training device based on a trusted data space, which can be implemented in hardware and / or software and can be configured in various general-purpose computing devices. It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. Figure 1a As shown, the method includes:

[0035] S110. Receive model training data from at least two data providers and store the model training data in a trusted data space.

[0036] The model training data can be encrypted data provided by the data provider.

[0037] Specifically, data users, i.e., those who train large models based on the data provided by data providers, can store the model training data in their own deployed trusted data space after receiving the model training data from the data provider, in order to protect the privacy of the industry data.

[0038] In this embodiment of the invention, the model training data provided by the data provider can refer to historical consultation data from hospitals. The data user can train a large-scale medical consultation model by receiving historical consultation data from multiple hospitals. It should be noted that the historical consultation data carries data tags, which can be used to indicate the consultation results corresponding to the historical consultation data. The consultation results include, but are not limited to, the symptom results obtained by doctors based on the historical consultation data.

[0039] S120. Divide the pre-defined large model to be trained into a plaintext domain model and a ciphertext domain model.

[0040] The ciphertext domain model can be deployed in the trusted data space of the data user, while the plaintext domain model can be deployed in the plaintext computing environment of the data user. The computing power level of the plaintext computing environment is greater than that of the trusted data space. It should be noted that the plaintext computing environment includes, but is not limited to, Graphics Processing Unit (GPU), Tensor Processing Unit (TPU), and Neural Processing Unit (NPU) environments. The trusted execution environment (TEE) is an independent processing environment with computation and storage functions, providing security, integrity, and confidentiality protection. It enables confidential computing capabilities such as encrypted data exiting the domain, non-recoverable data from untrusted environments, and controllable logic for processing data exiting the domain. Optionally, the portion of the large model to be trained that involves encrypted computation on encrypted data provided by the data provider can be designated as the ciphertext domain model, while the remaining model structures in the large model to be trained, excluding the ciphertext domain model, can be designated as the plaintext domain model.

[0041] For example, the scenario diagram between the data user and the data provider disclosed in the embodiments of the present invention can be as follows: Figure 1bAs shown, each data provider encrypts its local data within its own deployed trusted data space, then transmits the encrypted data to the data user through its own connector. The data user receives the data through its own deployed connector, stores the received encrypted data from the data provider in its own deployed trusted data space, and decrypts it. The data user can then train the large model to be trained based on its deployed trusted data space, plaintext computing power environment, and the decrypted encrypted data. It should be noted that the number of plaintext computing power environments deployed by the data user can be adaptively configured according to those skilled in the art.

[0042] S130. Based on the model fine-tuning method, the large model to be trained is trained according to the model training data, and the training process of the ciphertext domain model is transferred to the plaintext computing environment to which the plaintext domain model belongs.

[0043] S140. Use the trained large model to be trained as the target large model.

[0044] Specifically, when the large model to be trained meets the preset model training conditions, it indicates that the large model to be trained has completed training, and the trained large model to be trained is used as the target large model. Optionally, the model training conditions can be adaptively set according to those skilled in the art, for example, the number of iterations of the large model to be trained reaches a preset number.

[0045] The technical solution of this invention involves the data user receiving encrypted model training data through a trusted data space and training the large model to be trained on the basis of the trusted data space. This ensures that the data user cannot access the original model training data during the training process, thus guaranteeing data security. Furthermore, during the training of the large model to be trained, the training process of the ciphertext domain model is transferred to the plaintext computing environment of the plaintext domain model, thereby improving the training efficiency of the large model.

[0046] Example 2

[0047] Figure 2 This is a flowchart of a large model training method based on a trusted data space, provided in Embodiment 2 of the present invention. This embodiment further refines the above embodiments, providing specific steps for transferring the training process of the ciphertext domain model to the plaintext computing environment of the plaintext domain model. It should be noted that for parts not described in detail in this embodiment, please refer to the relevant descriptions in other embodiments, which will not be repeated here. Figure 2 As shown, the method includes:

[0048] S210. Deploy the first model sub-module and the second model sub-module in the ciphertext domain model in the plaintext computing environment, respectively, corresponding to the first backup sub-model and the second backup sub-model.

[0049] The first backup sub-model and the first model sub-module can have the same model structure and model parameters, and the second backup sub-model and the second model sub-module can have the same model structure and model parameters.

[0050] It should be noted that, in this embodiment of the invention, the plaintext domain model may include at least two cascaded model sub-modules. The first model sub-module and the second model sub-module can establish a cascaded relationship through the model sub-modules in the plaintext domain model. There is no direct data flow relationship between the first model sub-module and the second model sub-module deployed in the trusted data space. The output of the first model sub-module can serve as the input to at least two cascaded model sub-modules deployed in the plaintext computing environment, and the output of the last model sub-module among these at least two model sub-modules can serve as the input to the second model sub-module. Therefore, the first model sub-module and the second model sub-module can indirectly establish a connection relationship through the model sub-modules deployed in the plaintext computing environment.

[0051] In another specific embodiment of this invention, the large model to be trained can be divided into model sub-modules numbered 1-n. During the training process of the large model to be trained, the loss functions of the initial model sub-models and the final model sub-modules will use privacy data. Therefore, the initial model sub-models and the final model sub-modules can be designated as the first model sub-module and the second model sub-module, respectively, and defined as the ciphertext domain model. The remaining model sub-modules are defined as the plaintext domain model. It should be noted that the model sub-modules after the large model to be trained have the same original model parameters. After determining the ciphertext domain model and the plaintext domain model of the large model to be trained, the above-mentioned plaintext domain model and ciphertext domain model can be physically split so that the ciphertext domain model can be deployed in the trusted data space of the data user. In this embodiment of the present disclosure, the model sub-modules after the above split will be treated as a whole during training. The split does not change the key structure and integrity of the large model, nor does it change the principle of model training.

[0052] S220. The model input data of the first model submodule or the second model submodule is encrypted according to the random matrix generated in the plaintext computing environment, and the encrypted model input data is sent to the first backup sub-model or the second backup sub-model in the plaintext computing environment for training.

[0053] Optionally, the model input data of the first model submodule or the second model submodule is encrypted using a random matrix generated in the plaintext computing environment, and the encrypted model input data is sent to the first backup submodule or the second backup submodule in the plaintext computing environment for training. This includes: generating a random matrix in the plaintext computing environment and sending the random matrix to a trusted data space; in the trusted data space, encrypting the model input data of the first model submodule or the second model submodule according to the original model parameters of the first model submodule or the second model submodule and the random matrix, sending the encrypted model input data to the plaintext computing environment for training the first backup submodule or the second backup submodule, and sending the model training result to the trusted data space; in the trusted data space, decrypting the model training result according to the original model parameters and the random matrix, and updating the first model submodule or the second model submodule deployed in the trusted data space according to the decrypted model training result.

[0054] It should be noted that the process of encrypting the model input data of the first or second model submodule and the process of decrypting the model training results occur in a trusted data space. The data encryption method is also determined in the trusted data space. The encrypted model input data transmitted to the plaintext computing environment cannot be understood by the plaintext computing environment. Optionally, when encrypting the model input data, the model input data can be masked by randomly selecting matrix parameters with the same data dimension as the model input data from a random matrix. The generation of the random matrix can be adapted according to those skilled in the art.

[0055] By encrypting the model input data of the first or second model submodule and using the encrypted model input data for training the first or second backup submodule, the data security of the model input data is ensured during data use in the plaintext computing environment.

[0056] The technical solution of this invention improves the training efficiency of the model by transferring the training process of the first model sub-module or the second model sub-module to the first backup sub-model or the second backup sub-model deployed in the plaintext computing environment, respectively.

[0057] Example 3

[0058] Figure 3This is a flowchart of a large model training method based on a trusted data space provided in Embodiment 3 of the present invention. This embodiment further refines the above embodiments, providing specific steps for training the large model to be trained based on the model fine-tuning method and the model training data. It should be noted that for parts not described in detail in this embodiment, please refer to the relevant descriptions in other embodiments, which will not be repeated here. Figure 3 As shown, the method includes:

[0059] S310. According to the model fine-tuning method, the model training data is input into the ciphertext domain model, and forward propagation is performed sequentially through the first model submodule in the ciphertext domain model, the plaintext domain model, and the second model submodule in the ciphertext domain model to obtain the prediction output result.

[0060] S320. Based on the prediction output, backpropagation is performed through the first model submodule, the plaintext domain model, and the second model submodule to update the model parameters of the first model submodule, the plaintext domain model, and the second model submodule.

[0061] Optionally, in the process of training the large model to be trained based on the model training data using the model fine-tuning method, the method further includes: in the process of training the plaintext domain model and the ciphertext domain model, the model fine-tuning method is used to train the first self-attention layer, the first feedforward neural network layer and the second feedforward neural network layer in the model sub-modules of the plaintext domain model and the ciphertext domain model.

[0062] In this embodiment of the invention, the model sub-modules of the large model to be trained include a first self-attention layer (QKVLinear layer), a second self-attention layer (QK / √d calculation layer), a third self-attention layer (AV calculation unit layer), a first feedforward neural network layer (MLP Linear1 layer), and a second feedforward neural network layer (MLP Linear2 layer). After the model input data is input to the model sub-module, it can pass through the first self-attention layer, the second self-attention layer, the third self-attention layer, the first feedforward neural network layer, and the second feedforward neural network layer in sequence. In the first self-attention layer, the first feedforward neural network layer, and the second feedforward neural network layer, the model is trained using a model fine-tuning method.

[0063] During the training of the large model to be trained, each iteration includes forward propagation and backward propagation. In forward propagation, the output of the preceding model submodule serves as the input to the following model submodule. The following model submodule receives the input from the preceding submodule and continues propagating backward until the last model submodule of the large model to be trained is reached. During the updating of the model parameters of the large model to be trained, the original model parameters of each model submodule remain unchanged. A model fine-tuning method is used to update only the low-rank parameters. After the gradient parameters are calculated by the following model submodule, the model parameters of the preceding model submodule are updated using a backward gradient update method.

[0064] The technical solution of this invention, by using a model fine-tuning method to train the ciphertext domain model in a trusted data space, enables the secure encryption of private data out of the domain, and makes it impossible to restore in a plaintext computing environment, thus protecting the client's private data.

[0065] Example 4

[0066] Figure 4 This is a schematic diagram of the structure of a large model training device based on a trusted data space, provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes:

[0067] The data receiving module 410 is used to receive model training data from at least two data providers and store the model training data in a trusted data space; wherein the model training data is encrypted data provided by the data providers.

[0068] The model deployment module 420 is used to divide the preset large model to be trained into a plaintext domain model and a ciphertext domain model; wherein, the plaintext domain model is deployed in the plaintext computing environment of the data user, and the ciphertext domain model is deployed in the trusted data space of the data user; the computing power level of the plaintext computing environment is greater than that of the trusted data space.

[0069] The model training module 430 is used to train the large model to be trained based on the model training data according to the model fine-tuning method, and to transfer the training process of the ciphertext domain model to the plaintext computing environment to which the plaintext domain model belongs.

[0070] The target model generation module 440 is used to take the trained large model to be trained as the target large model.

[0071] The technical solution of this invention involves the data user receiving encrypted model training data through a trusted data space and training the large model to be trained on the basis of the trusted data space. This ensures that the data user cannot access the original model training data during the training process, thus guaranteeing data security. Furthermore, during the training of the large model to be trained, the training process of the ciphertext domain model is transferred to the plaintext computing environment of the plaintext domain model, thereby improving the training efficiency of the large model.

[0072] Optionally, the model training module 430 includes:

[0073] A backup sub-model unit is used to deploy the first model sub-module and the second model sub-module in the ciphertext domain model in the plaintext computing environment, respectively corresponding to the first backup sub-model and the second backup sub-model; wherein, the first backup sub-model and the first model sub-module have the same model structure and model parameters, and the second backup sub-model and the second model sub-module have the same model structure and model parameters;

[0074] The training transfer unit is used to encrypt the model input data of the first model submodule or the second model submodule according to the random matrix generated in the plaintext computing environment, and send the encrypted model input data to the first backup submodel or the second backup submodel in the plaintext computing environment for training.

[0075] Optionally, the training transfer unit includes:

[0076] A matrix generation subunit is used to generate a random matrix in a plaintext computing environment and send the random matrix to a trusted data space;

[0077] The model training subunit is used to encrypt the model input data of the first model submodule or the second model submodule in the trusted data space according to the original model parameters of the first model submodule or the second model submodule and the random matrix, send the encrypted model input data to the plaintext computing environment, train the first backup submodel or the second backup submodel, and send the model training results to the trusted data space.

[0078] The model update subunit is used to decrypt the model training results in the trusted data space based on the original model parameters and the random matrix, and update the first model submodule or the second model submodule deployed in the trusted data space based on the decrypted model training results.

[0079] Optionally, the model training module 430 also includes:

[0080] The forward processing unit is used to input the model training data into the ciphertext domain model according to the model fine-tuning method, and then perform forward propagation processing through the first model submodule in the ciphertext domain model, the plaintext domain model, and the second model submodule in the ciphertext domain model in sequence to obtain the prediction output result.

[0081] The back-processing unit is used to update the model parameters of the first model submodule, the plaintext domain model, and the second model submodule by back-propagating through the first model submodule, the plaintext domain model, and the second model submodule according to the prediction output result.

[0082] Optionally, the plaintext domain model includes at least two cascaded model sub-modules, wherein the first model sub-module and the second model sub-module are connected in series through the model sub-modules in the plaintext domain model.

[0083] Optionally, the model training module 430 also includes:

[0084] The model fine-tuning unit is used to train the first self-attention layer, the first feedforward neural network layer, and the second feedforward neural network layer in the model sub-modules of the plaintext domain model and the ciphertext domain model by using the model fine-tuning method during the training process.

[0085] The large model training device based on trusted data space provided in the embodiments of the present invention can execute the large model training method based on trusted data space provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0086] Example 5

[0087] Figure 5 A schematic diagram of an electronic device 510 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0088] like Figure 5As shown, the electronic device 510 includes at least one processor 511 and a memory, such as a read-only memory (ROM) 512 or a random access memory (RAM) 513, communicatively connected to the at least one processor 511. The memory stores computer programs executable by the at least one processor. The processor 511 can perform various appropriate actions and processes based on the computer program stored in the ROM 512 or loaded into the RAM 513 from storage unit 518. The RAM 513 may also store various programs and data required for the operation of the electronic device 510. The processor 511, ROM 512, and RAM 513 are interconnected via a bus 514. An input / output (I / O) interface 515 is also connected to the bus 514.

[0089] Multiple components in electronic device 510 are connected to I / O interface 515, including: input unit 516, such as keyboard, mouse, etc.; output unit 517, such as various types of displays, speakers, etc.; storage unit 518, such as disk, optical disk, etc.; and communication unit 519, such as network card, modem, wireless transceiver, etc. Communication unit 519 allows electronic device 510 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0090] Processor 511 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 511 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 511 performs the various methods and processes described above, such as large model training methods based on a trusted data space.

[0091] In some embodiments, the large model training method based on a trusted data space can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 518. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 510 via ROM 512 and / or communication unit 519. When the computer program is loaded into RAM 513 and executed by processor 511, one or more steps of the large model training method based on a trusted data space described above can be performed. Alternatively, in other embodiments, processor 511 can be configured to perform the large model training method based on a trusted data space by any other suitable means (e.g., by means of firmware).

[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0094] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0096] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0097] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0098] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0099] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for training large models based on a trusted data space, characterized in that, include: Receive model training data from at least two data providers and store the model training data in a trusted data space; wherein the model training data is encrypted data provided by the data providers. The pre-defined large model to be trained is divided into a plaintext domain model and a ciphertext domain model; wherein, the plaintext domain model is deployed in the plaintext computing environment of the data user, and the ciphertext domain model is deployed in the trusted data space of the data user; the computing power level of the plaintext computing environment is greater than that of the trusted data space. Based on the model fine-tuning method, the large model to be trained is trained according to the model training data, and the training process of the ciphertext domain model is transferred to the plaintext computing environment to which the plaintext domain model belongs. The trained large model is used as the target large model.

2. The method according to claim 1, characterized in that, The step of transferring the training process of the ciphertext domain model to the plaintext computing environment to which the plaintext domain model belongs includes: In the plaintext computing environment, a first model submodule and a second model submodule are deployed in the ciphertext domain model, corresponding to a first backup submodule and a second backup submodule, respectively; wherein, the first backup submodule has the same model structure and model parameters as the first model submodule, and the second backup submodule has the same model structure and model parameters as the second model submodule; The model input data of the first model submodule or the second model submodule is encrypted using a random matrix generated in the plaintext computing environment, and the encrypted model input data is sent to the first backup submodule or the second backup submodule in the plaintext computing environment for training.

3. The method according to claim 2, characterized in that, The step of encrypting the model input data of the first model submodule or the second model submodule according to the random matrix generated in the plaintext computing environment, and sending the encrypted model input data to the first backup submodule and the second backup submodule in the plaintext computing environment, and training the first backup submodule and the second backup submodule includes: A random matrix is ​​generated in a plaintext computing environment, and the random matrix is ​​sent to a trusted data space; In the trusted data space, the model input data of the first model sub-module or the second model sub-module is encrypted according to the original model parameters of the first model sub-module or the second model sub-module and the random matrix, and the encrypted model input data is sent to the plaintext computing environment to train the first backup sub-model or the second backup sub-model, and the model training results are sent to the trusted data space. In the trusted data space, the model training results are decrypted based on the original model parameters and the random matrix, and the first model submodule or the second model submodule deployed in the trusted data space is updated based on the decrypted model training results.

4. The method according to claim 1, characterized in that, The model-based fine-tuning method trains the large model to be trained based on the model training data, including: According to the model fine-tuning method, the model training data is input into the ciphertext domain model, and forward propagation is performed sequentially through the first model submodule in the ciphertext domain model, the plaintext domain model, and the second model submodule in the ciphertext domain model to obtain the prediction output result. Based on the predicted output, backpropagation is performed through the first model submodule, the plaintext domain model, and the second model submodule to update the model parameters of the first model submodule, the plaintext domain model, and the second model submodule.

5. The method according to any one of claims 2-4, characterized in that, The plaintext domain model includes at least two cascaded model sub-modules, and the first model sub-module and the second model sub-module are connected in series through the model sub-modules in the plaintext domain model.

6. The method according to claim 1, characterized in that, The model-based fine-tuning method, which trains the large model to be trained based on the model training data, further includes: During the training of the plaintext domain model and the ciphertext domain model, a model fine-tuning method is used to train the first self-attention layer, the first feedforward neural network layer, and the second feedforward neural network layer in the model sub-modules of the plaintext domain model and the ciphertext domain model.

7. A large model training device based on a trusted data space, characterized in that, include: A data receiving module is used to receive model training data from at least two data providers and store the model training data in a trusted data space; wherein the model training data is encrypted data provided by the data providers. The model deployment module is used to divide the preset large model to be trained into a plaintext domain model and a ciphertext domain model; wherein, the plaintext domain model is deployed in the plaintext computing environment of the data user, and the ciphertext domain model is deployed in the trusted data space of the data user; the computing power level of the plaintext computing environment is greater than that of the trusted data space. The model training module is used to train the large model to be trained based on the model fine-tuning method and the model training data, and to transfer the training process of the ciphertext domain model to the plaintext computing environment to which the plaintext domain model belongs. The target model generation module is used to take the trained large model to be trained as the target large model.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the large model training method based on a trusted data space as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the large model training method based on a trusted data space as described in any one of claims 1-6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the large model training method based on a trusted data space as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Large model fine tuning method and device based on trusted execution environment

    CN118296615A

  • Large model training method and device based on privacy protection and storage medium

    CN119382937A

  • Training data protection in artificial intelligence model execution environment

    US20220405383A1

  • System and method for privacy-preserving distributed training of machine learning models on distributed datasets

    US20230188319A1