Large model training method and device based on data security, equipment, medium and product

By dividing the data into training data sets at different security levels and training for different security levels, corresponding sub-models are obtained, which solves the data security guarantee problem in large model training, and achieves both efficient training and data security.

CN120086594APending Publication Date: 2025-06-03PIPECHINA SOUTH CHINA CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510158424.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

During the training of large-scale models, the existing technology is difficult to effectively ensure data security, and there are problems such as data leakage and increased computing complexity.

Method used

By dividing the data into different security levels and assigning the data of each security level to the corresponding security data shard. Each security data shard is used as a training data set to fine-tune the basic big model during the model training process to obtain the corresponding security level segment model.

Benefits of technology

It realizes efficient training of large models while fully ensuring data security, avoiding data loss or distortion caused by data desensitization, as well as increasing computational complexity and storage overhead caused by model encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086594A_ABST
    Figure CN120086594A_ABST
Patent Text Reader

Abstract

The invention discloses a large model training method and device based on data security, equipment, a medium and a product. The large model training method based on data security comprises the steps that data are divided into different security levels, the data of each security level are distributed to a corresponding security data fragment, and each security data fragment serves as a training data set; in the model training process, each training data set is used for conducting fine adjustment on the basic large model, and sub-models of the corresponding safety levels are obtained. According to the technical scheme, the data is divided into the training data sets of the different security levels, training is carried out for the different security levels to obtain the corresponding sub-models, data security can be fully guaranteed, the sub-models suitable for the data of the different security levels are obtained, and the training efficiency of the large model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a large model training method, device, equipment, medium and product based on data security. Background Art

[0002] With the wide deployment and application of generative artificial intelligence technology, especially multi-modal large language models and visual generation models, these models are favored by enterprises in various industries due to their excellent performance and emerging reasoning capabilities. In the process of training and fine-tuning large models, a large amount of enterprise private data is required to provide support. These data often involve enterprise sensitive information and may also include personal privacy, business secrets, etc. When enterprises deeply develop and utilize large models, they generally need to use a large amount of internal enterprise private data to train and fine-tune on the basis of these models. The powerful memory ability of the generated model will remember and generate sensitive, outdated or dangerous knowledge and information derived from the training data. While large models provide convenient services, they may lead to the leakage of enterprise internal data. If these sensitive information is leaked or misused during the large model training process, serious consequences will be caused.

[0003] To ensure the security of sensitive data during the training and fine-tuning of large models, existing solutions mainly achieve this through data desensitization and model encryption. Data desensitization refers to processing data before data transmission or storage so that sensitive information cannot be identified or traced. Model encryption refers to encrypting the model during the model training and inference processes to prevent sensitive information from being leaked or misused.

[0004] However, these methods have some deficiencies. For example, data desensitization may lead to data loss or distortion, thus affecting the training effect of the model; model encryption will increase the computational complexity and storage overhead, thus affecting the performance and scalability of the model. In addition, these methods cannot completely avoid the leakage risk of sensitive information during the model training process. Therefore, how to efficiently train large models on the premise of fully ensuring data security has become an urgent problem to be solved. Summary of the Invention

[0005] The present application provides a large model training method, device, equipment, medium and product based on data security to efficiently train large models on the premise of fully ensuring data security.

[0006] In a first aspect, the embodiments of the present application provide a large model training method based on data security, including:

[0007] Dividing the data into different security levels, and respectively allocating the data of each security level to the corresponding secure data shards, and each secure data shard serves as a training data set;

[0008] During the model training process, each training dataset is used to fine-tune the basic large model respectively to obtain sub-models corresponding to different security levels.

[0009] In a second aspect, an embodiment of the present application further provides a large model training device based on data security, including:

[0010] A partitioning module, configured to partition data into different security levels and allocate the data of each security level to corresponding secure data shards respectively, and each secure data shard serves as a training dataset;

[0011] A training module, configured to, during the model training process, use each training dataset to fine-tune the basic large model respectively to obtain sub-models corresponding to different security levels.

[0012] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0013] One or more processors;

[0014] A storage device, configured to store one or more programs;

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the large model training method based on data security as described in the first aspect.

[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the large model training method based on data security as described in the first aspect.

[0017] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program and / or instructions, and when the computer program and / or instructions are executed by a processor, they implement the large model training method based on data security as described in any of the above embodiments.

[0018] An embodiment of the present application provides a large model training method, device, equipment, medium and product based on data security. The large model training method based on data security includes: partitioning data into different security levels, and allocating the data of each security level to corresponding secure data shards respectively, and each secure data shard serves as a training dataset; during the model training process, using each training dataset to fine-tune the basic large model respectively to obtain sub-models corresponding to different security levels. The above technical solution can fully guarantee data security by partitioning data into training datasets of different security levels and training separately for different security levels to obtain corresponding sub-models, and obtain sub-models applicable to data of different security levels, improving the training efficiency of the large model. Brief Description of the Drawings

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0020] Figure 1 It is a flowchart of a large model training method based on data security provided by an embodiment of the present application;

[0021] Figure 2 It is a schematic diagram of secure data sharding and training sub-models provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of a large model training framework based on data security provided by an embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of the structure of a large model training device based on data security provided by an embodiment of the present application;

[0024] Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0025] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application are shown in the accompanying drawings rather than all the structures.

[0026] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0027] It should be noted that the concepts such as "first" and "second" mentioned in the embodiments of the present application are only used to distinguish different devices, modules, units or other objects, and are not used to limit the order of functions performed by these devices, modules, units or other objects or their interdependent relationships.

[0028] In addition, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other.

[0029] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0030] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used the relevant content of this solution.

[0031] Currently, there is a lack of efficient and flexible security supervision for the internal training and fine-tuning data during the development of enterprise large models. A large amount of data input may cause data leakage problems. This application aims to provide a customized large model corresponding to different security levels on the premise of security control to provide personalized services within the corresponding permission scope.

[0032] Figure 1 FIG. is a flowchart of a large model training method based on data security provided for the embodiments of this application. This embodiment is applicable to the situation of training a large model for secure data. Specifically, the large model training method based on data security can be executed by a large model training device based on data security. The large model training device based on data security can be implemented in software and / or hardware and integrated in an electronic device. The electronic device includes but is not limited to devices such as a computer, a laptop, a smartphone, or a server.

[0033] As Figure 1 shown, the method specifically includes the following steps:

[0034] S110. Divide the data into different security levels, and respectively allocate the data of each security level to the corresponding secure data shards. Each secure data shard serves as a training data set.

[0035] In this embodiment, the data may refer to the data used for training or fine-tuning a large model, such as application data, business data, or enterprise internal data, etc., which may include sensitive data or privacy data. According to the security level, the data can be divided into different security levels, and each security level corresponds to different data access permissions and security protection measures. By dividing the data into different security levels, on the one hand, the privacy and security of sensitive data can be effectively ensured, and on the other hand, personalized inference support can be provided for users with different security permissions.

[0036] Exemplarily, the internal data of an enterprise is divided (classified or graded). For example, it can be successively divided into multiple security levels such as internally public data, department - public data, department - restricted data, and non - trainable data. Among them, internally public data can be understood as data that supports full - access within the enterprise; department - public data can be understood as data that supports public access within the corresponding department but cannot be accessed across departments; department - restricted data can be understood as data for which the access scope and rights need to be approved and confirmed by the department's leading supervisor; non - trainable data can be understood as controlled data that needs to be managed by the enterprise and is not suitable for inclusion in the scope of training data, and this type of data is strictly controlled. During the process of dividing security levels, the security level of the data can be identified, and data cleaning can be performed to ensure that each data item has a clear security level.

[0037] Furthermore, according to the security levels, the data is assigned to different security data shards, such as public data shards, internal data shards, and confidential data shards, etc. Each security data shard corresponds to a different security level. Each security data shard can be used as a training data set for training the corresponding sub - model. Through this hierarchical training method, the risk of unauthorized leakage of sensitive information during the construction and training of the large model can be effectively prevented.

[0038] Exemplarily, the training data sets of each security level can be assigned to different security data shards. For each security level, according to actual needs, the data of the corresponding security level can be assigned to one or more security data shards.

[0039] Optionally, for each security level, the data of the corresponding security level can be further divided according to features such as the source department of the data, data collection time, data collection location, and / or relevant business types, etc. Each part of the divided data is respectively used as a training data set and assigned to the corresponding security data shard. Each security data shard can be used as the input data for subsequent fine - tuning of the large - model scoring model training.

[0040] S120. During the model training process, each training data set is used to fine - tune the basic large model respectively to obtain the sub - model corresponding to the security level.

[0041] In this embodiment, different security data shards are used as training data sets, and training fine - tuning is performed on the basis of the base large model to obtain the large - model sub - models corresponding to different security levels. This method of fine - tuning the large - model sub - models can effectively improve the accuracy and reliability of models at different security levels.

[0042] Exemplarily, for training and fine-tuning, the public data shards can be used to train and fine-tune the base large model to obtain a public data sub-model; the internal data shards can be used to train and fine-tune the base large model to obtain an internal data sub-model; and the confidential data shards can be used to train and fine-tune the base large model to obtain a confidential data sub-model.

[0043] Optionally, the base large models corresponding to different secure data shards can be the same large model, such as a large model trained using external public data or open-source data sets, etc., or can be a known or pre-trained off-the-shelf large model.

[0044] Optionally, the base large models corresponding to different secure data shards can be different large models. For example, the base large model corresponding to one secure data shard can be a sub-model trained based on another secure data shard; also, for example, the base large model corresponding to secure data shard 1 is an off-the-shelf large model, denoted as large model 0. Using secure data shard 1 to train and fine-tune this large model, sub-model 1 is obtained; the base large model corresponding to secure data shard 2 is sub-model 1. Using secure data shard 2 to train and fine-tune sub-model 1, sub-model 2 is obtained; the base large model corresponding to secure data shard 3 is sub-model 2. Using secure data shard 3 to train and fine-tune sub-model 2, sub-model 3 is obtained, and so on.

[0045] Based on the above, a series of sub-models with different security levels can be obtained, and then, on the premise of ensuring data security, personalized inference services can be provided for users with different security permissions.

[0046] In one embodiment, a large model proxy can be set up. According to the department where the user is located and the security permissions, etc., the sub-model within the security level range that the user can access is called or docked to provide further inference services, and finally, a personalized intelligent large model service under the premise of security is realized.

[0047] A large model training method based on data security provided by the embodiments of this application, by dividing data into different security levels, on the one hand, can effectively ensure the privacy and security of sensitive data, and on the other hand, can also provide personalized inference support for users with different security permissions; by performing hierarchical training on different security levels to obtain corresponding sub-models, the risk of unauthorized leakage of sensitive information during the construction and training of the large model can be effectively prevented; by fine-tuning to obtain sub-models of different security levels, the accuracy and reliability of models at different security levels can be effectively improved. This method can achieve high-efficiency large model training, and can fully guarantee the privacy and security of sensitive data. In addition, it can avoid problems such as data loss, distortion, and increased computational complexity and storage overhead caused by data desensitization and model encryption.

[0048] In one embodiment, each training dataset is used to fine-tune the basic large model to obtain sub-models corresponding to the corresponding security levels, including:

[0049] For each of the ordered security levels, the training dataset corresponding to the security level is used to fine-tune the basic large model corresponding to the security level to obtain a sub-model corresponding to the security level;

[0050] Among them, the basic large model corresponding to the first security level is a general large model, and the basic large models corresponding to security levels other than the first security level are the basic large models corresponding to the previous security level. The ordered security levels can be understood as the security levels from high to low or from low to high.

[0051] Figure 2 This is a schematic diagram of a secure data sharding and training sub-model provided by an embodiment of the present application. As Figure 2 shown, the enterprise internal data can be divided into security levels and allocated to obtain secure data shards at different security levels, such as secure data shard 0 - enterprise internal public level, secure data shard 1 - department internal public level, secure data shard 2 - department internal controlled level, etc. Among them, a single secure data shard can be further sharded by department.

[0052] In this embodiment, the basic large model corresponding to a secure data shard can be a sub-model trained based on another secure data shard. The general basic large model can be trained using external public data or open-source datasets, etc. This basic large model generally has a large number of parameters and its reasoning ability is comprehensive. This basic large model serves as the basic general large model for subsequent fine-tuning within the enterprise. According to different secure data shards, sub-models corresponding to the corresponding security levels are trained. Exemplarily, using the internal public data shard (level 0 data) as the training dataset, the basic general large model is trained and fine-tuned to obtain a large model sub-model at the internal public level, which can be called the level 0 large model sub-model; then, using the level 0 large model sub-model as the basic large model, and continuing to use a part of the internal public data shard (level 1 data) as the training dataset, the basic large model is trained and fine-tuned to obtain the corresponding level 1 large model sub-model, and so on. Using different secure data shards as the training dataset, training and fine-tuning are performed on the basis of the sub-model of the previous level to obtain sub-models corresponding to the corresponding security levels. Among them, the sub-model of the previous level can be understood as the sub-model obtained by training and fine-tuning using the enterprise internal data of the previous security level.

[0053] In one embodiment, it further includes:

[0054] S130. Locate the secure data shard where the data to be forgotten is located;

[0055] S140. Delete the data to be forgotten to update the corresponding training data set;

[0056] S150. Update the sub-model of the corresponding security level according to the updated training data set.

[0057] Considering that the forgetting and updating costs of large models are relatively high, and the sensitivity and security levels of data will change over time and with the environment, large models can support the forgetting of sensitive data. The forgetting process of data can effectively prevent the leakage risk of sensitive data during subsequent model inference services. The conventional approach is to analyze and remove the corresponding sensitive data from the training data and then retrain the large model when the sensitivity of the training data changes. This approach requires a large amount of computing resources and time costs.

[0058] In this embodiment, the data to be forgotten can be understood as the data to be forgotten, that is, the data that should not continue to be used as training data during subsequent training. According to actual needs, the data to be forgotten can be accurately identified and marked, and the security level or security data shard where the data to be forgotten is located can be located. Then, the data to be forgotten is deleted from the security data shard of the corresponding security level to obtain the updated training data set, and the sub-model of this security level can be trained using this training data set. On this basis, the forgetting and updating costs can be effectively reduced, and the efficiency of model training and updating can be improved.

[0059] Exemplarily, determining the data to be forgotten according to actual needs may include the following situations:

[0060] The security level of the data changes, especially when the sensitivity of the data increases. For example, due to policy requirements and other reasons, the required scope of knowledge of some department data in security data shard 1 has been upgraded or changed in terms of security level, sensitivity, and / or applicable scope, and it is no longer applicable to security level 1. Therefore, this part of the data in the original security data shard 1 needs to be deleted, and the corresponding knowledge in sub-model 1 of the large model needs to be forgotten and deleted synchronously;

[0061] Due to reasons such as data timeliness, data retirement, and / or data quality errors, the data in the original security data shard has been updated, so the data in the original data shard needs to be forgotten, updated, or corrected.

[0062] Based on locating the secure data shards that need to be replaced or deleted, a new training dataset with the above-mentioned data to be forgotten replaced or deleted can be re-prepared, aiming to exclude all the marked data to be forgotten. By locating the security level where the data to be forgotten is located, based on the sub-model of the large model at the previous security level, according to the scale of the sub-model and the training dataset, a machine forgetting method can be selected to forget or delete the above-mentioned data to be forgotten. The specific methods include, but are not limited to, data deletion and retraining, knowledge distillation, etc. On this basis, by accurately identifying and marking the data to be forgotten and locating the security level area where this part of the data is located, refined management of the data is achieved, and sensitive data can be identified and processed more accurately, thus better protecting data privacy and security.

[0063] In one embodiment, updating the sub-model of the corresponding security level according to the updated training dataset includes one of the following:

[0064] According to the updated training dataset, fine-tune the sub-model of the corresponding security level to update the sub-model of the corresponding security level;

[0065] Take the sub-model of the corresponding security level as the teacher model, and train a student model based on the teacher model and the updated training dataset as the updated sub-model of the corresponding security level.

[0066] Exemplarily, for the method of data deletion and retraining, it mainly achieves the purpose of data forgetting processing by performing data deletion and retraining fine-tuning to regenerate the sub-model of the corresponding security level. Based on the sub-model corresponding to the previous data security level, use the remaining dataset after the update to retrain this sub-model, and finally obtain a new sub-model of this security level. This method is mainly applicable to the situation where the dataset scale is small or the model parameters are not large.

[0067] For the method of knowledge distillation, a teacher model can be prepared, and this teacher model can be the sub-model corresponding to the previous data security level; use this teacher model and combine the training data of this security level after removing the changed security level to train a new student model, so that while learning the behavior of the teacher model, it can ensure that the student model does not contact the sensitive data to be deleted and forgotten during the training process; then take this student model as the new sub-model to replace the original sub-model of this level. This method is mainly applicable to sub-models with a large dataset scale or huge model parameters, which can effectively reduce the computational complexity and storage overhead, save time and computing power resources, and improve the scalability of the model.

[0068] Based on the above, data forgetting is processed through two methods: retraining by data deletion and knowledge distillation. This can not only effectively avoid the risk of sensitive data leakage during model training but also ensure that the performance of the model is not affected. While protecting data privacy and security, it improves the performance and scalability of the model.

[0069] In one embodiment, the method further includes:

[0070] S160. For any of the security levels, determine the model parameters related to the data to be forgotten according to the differences between the sub-models before and after the update of the security level;

[0071] S170. Prune or adjust the model parameters.

[0072] Exemplarily, for any security level, the differences between the two sub-models before and after forgetting at this security level (or the differences between the teacher model and the student model) can be analyzed, the model parameters related to sensitive data in the sub-model can be identified, these model parameters can be pruned or adjusted, and the obtained new model parameters can be re-integrated into the subsequent sub-model training and improvement process to form the final sub-model of the large model that forgets sensitive data. On this basis, it helps to reduce the time and computing power costs of re-training the sub-model of the large model in the subsequent improvement of the training process, accurately remove the dependence of the large model on the sensitive data to be forgotten or updated, and if there are similar forgetting and updating requirements in the future, it can greatly improve the model update efficiency.

[0073] In one embodiment, the method further includes:

[0074] S180. For any of the security levels, evaluate the sub-model after the update of the security level;

[0075] S190. Publish the updated sub-model that passes the evaluation to the model version library.

[0076] In this embodiment, the sub-model after forgetting and updating can be comprehensively evaluated and continuously monitored to ensure that it meets the expected forgetting and performance standards. The models that fail the evaluation test can be directionally improved and intensively trained, thereby effectively ensuring the security and accuracy of the model for machine forgetting learning.

[0077] Exemplarily, build a model version library, deploy the updated sub-model to the experimental and production environments, continuously monitor the performance of the sub-model, an audit log can be established, and wait for the next round of large model updates and iterations. Comprehensively evaluate and continuously improve the sub-model after forgetting and updating to ensure that it meets the expected forgetting and performance standards. The evaluation of the sub-model can include the following aspects:

[0078] Accuracy: Measure the degree to which the sub-model successfully forgets specific unwanted knowledge (including data to be forgotten), ensuring that the sub-model does not produce outputs that should be forgotten. For example, it can be checked whether the sub-model can successfully avoid generating outputs related to sensitive information when processing inputs containing sensitive information. Some test cases can be designed to observe whether the outputs of the sub-model meet the expectations, that is, no longer contain sensitive information.

[0079] Locality: Measure the ability of the sub-model after forgetting learning to retain knowledge related to the non-forgotten content, ensuring that the performance of the sub-model on the retained dataset is maintained. Evaluate the impact of the method on the performance of the sub-model in other non-sensitive tasks. Observe indicators such as accuracy, recall, and F1 value of the sub-model when processing normal tasks to ensure that its performance is not overly affected by forgetting sensitive information.

[0080] Generalizability: Test whether the sub-model can equally effectively forget sensitive information when faced with new, unseen inputs containing sensitive information. This can be verified by using new test data, and the ability of the sub-model to extend forgetting to unseen forgotten datasets can be evaluated to ensure that the sub-model does not produce outputs for new inputs related to the concept of forgetting.

[0081] The machine forgetting process of the large model sub-model may produce adverse results such as incomplete forgetting (the sub-model may still retain the influence of the data to be forgotten to some extent, but only does not show it in the output) and the impact on the performance of the sub-model (the performance of the sub-model in other non-sensitive tasks is affected); in addition, for some sub-models that fail the evaluation, they can be improved by re-executing the forgetting training process, continuous reinforcement learning, etc.

[0082] Furthermore, the sub-model that meets the forgetting requirements after evaluation can be released to the model version library. In this embodiment, it is supported to be first deployed to the experimental environment, and the performance of the sub-model is continuously monitored and accepted by user detection to ensure that it meets the expected forgetting and performance standards. During this process, monitoring tools can be used to monitor the inputs and outputs of the sub-model in real time to ensure that the sub-model does not output sensitive data, that is, verify the updated sub-model in the experimental environment, and then release it to the production environment on the basis of achieving good results.

[0083] On this basis, through steps such as parameter pruning, model fusion, evaluation, and testing, the security of the model is further enhanced, which can more effectively prevent the leakage and abuse of sensitive data during model training, and improve the security and reliability of the model; in the model deployment and monitoring stage, continuously monitor the performance of the model to ensure that it meets the expected forgetting and performance standards, and can promptly detect and solve problems that occur during the operation of the model, improving the reliability and stability of the model.

[0084] Figure 3 It is a schematic diagram of a large model training framework based on data security provided by an embodiment of this application. As Figure 3 shown, the design idea for building and updating a large model based on data security is as follows:

[0085] Data preprocessing module: It can perform high-quality cleaning on the received internal enterprise data, including data quality verification, data format conversion, data labeling, and / or metadata cleaning, etc., to ensure data quality;

[0086] Data security classification and grading module: It can complete the security classification and identification of internal enterprise data, confirm the security level of each data item, and label each data with a security classification label;

[0087] Data sharding, partitioning, and data storage warehouse module: It is responsible for sharding the data after security classification and determining the data storage partition. This module is also responsible for storing all training dataset data, including the latest version of data security data shards, storage of other test training datasets and different version datasets;

[0088] Data change and version management module: It can provide control and registration of different versions of data sharding and partitioning, provide data sharding change processing and record event logs;

[0089] Basic large model call interface: It can provide the basic large model API interface for enterprises to train business large models, and can choose to adapt different basic large models as the base for fine-tuning and customization;

[0090] Model training and fine-tuning module: Based on the computing power resources within the enterprise, at different security level levels, combined with the internal enterprise data security data shards, carry out the training and fine-tuning of the security level large model sub-models, and output different sub-models to the model warehouse;

[0091] Model change and version management module: It can provide version control and registration for large model sub-models of different security levels. Each model version number should be unique, provide management for changes and iterative releases of different large model sub-models, record monitoring logs and model version metadata, and provide processing such as version rollback;

[0092] Model storage repository module: It can store the latest fine-tuned sub-models of the large model and provide model storage for various historical process versions;

[0093] Model configuration module: It can provide comparison between different models, locate the parameters related to the data to be forgotten in the model, and provide configuration management functions such as modification, debugging of model parameters, and fusion between different models;

[0094] Model evaluation and testing module: It can provide evaluation and testing before the release of different versions of the model, conduct trial operation and monitoring of the model using the test data set, and provide test conclusions and problem location of the sub-model;

[0095] User interaction agent module: It can provide release interfaces for different large model security sub-models, and through analyzing the security level and permissions of users, provide an interaction agent for the inference service of the sub-model for users.

[0096] The large model training method based on data security in the embodiments of the present application can, while efficiently performing large model training and inference, ensure the privacy protection and security of sensitive data; it can avoid problems such as data loss, distortion, increased computational complexity, and increased storage overhead brought by existing data desensitization and model encryption technologies; it can effectively prevent the leakage risk of sensitive information during the machine learning and large model training processes; it can effectively identify and forget the data that needs to be forgotten during the training and inference processes of the new model, and locate the data classification and stratification area where this part of the data is located; it can effectively prune and adjust the parameters related to sensitive data without affecting the performance of the model on non-sensitive data, thereby reducing the dependence of the updated sub-model on the sensitive data to be forgotten or corrected; it can comprehensively evaluate and continuously monitor the model after forgetting to ensure that it meets the expected forgetting and performance standards. In addition, this method can effectively protect the privacy and security of sensitive data and also ensure the performance of the model on non-sensitive data.

[0097] Figure 4 It is a structural schematic diagram of a large model training device based on data security provided by the embodiments of the present application. As Figure 4 shown, the large model training device based on data security provided by this embodiment includes:

[0098] Partitioning module 210, which is used to partition the data into different security levels, and respectively allocate the data of each security level to the corresponding secure data shards, and each secure data shard serves as a training data set;

[0099] Training module 220, which is used to fine-tune the basic large model using each training data set respectively during the model training process to obtain sub-models corresponding to the security levels.

[0100] By dividing the data into training data sets of different security levels and training the corresponding sub-models separately for different security levels, the device can fully ensure data security and obtain sub-models applicable to data of different security levels, improving the training efficiency of the large model.

[0101] Optionally, the training module 220 is specifically configured to, during the model training process, for each of the ordered security levels, use the training data set corresponding to the security level to fine-tune the basic large model corresponding to the security level to obtain a sub-model corresponding to the security level;

[0102] Among them, the basic large model corresponding to the first security level is a general large model, and the basic large models corresponding to security levels other than the first security level are the basic large models corresponding to the previous security level.

[0103] Optionally, the device further includes:

[0104] A positioning module, configured to locate the secure data shard where the data to be forgotten is located;

[0105] A deletion module, configured to delete the data to be forgotten to update the corresponding training data set;

[0106] An update module, configured to update the sub-model of the corresponding security level according to the updated training data set.

[0107] Optionally, the update module is specifically configured to perform one of the following:

[0108] According to the updated training data set, fine-tune the sub-model of the corresponding security level to update the sub-model of the corresponding security level;

[0109] Use the sub-model of the corresponding security level as a teacher model, and train a student model based on the teacher model and the updated training data set as the updated sub-model of the corresponding security level.

[0110] Optionally, the device further includes:

[0111] A determination module, configured to, for any of the security levels, determine the model parameters related to the data to be forgotten according to the differences between the sub-models before and after the update of the security level;

[0112] An adjustment module, configured to prune or adjust the model parameters.

[0113] Optionally, the device further includes:

[0114] An evaluation module, configured to, for any of the security levels, evaluate the sub-model after the update of the security level:

[0115] A publishing module for publishing the updated sub-model that has passed the evaluation to the model version library.

[0116] The large model training device based on data security provided by the embodiments of the present application can be used to execute the large model training method based on data security provided by any of the above embodiments, and has the corresponding functions and beneficial effects.

[0117] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present application. The electronic device 10 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 10 can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, user equipment, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0118] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0119] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, wireless networks.

[0120] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above.

[0121] In some embodiments, the methods of the above embodiments may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the methods described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the methods of any of the above embodiments in any other suitable manner (e.g., by means of firmware).

[0122] The various embodiments of the systems and techniques described above may be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0123] The computer programs for implementing the methods of the present application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0124] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device 10 having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device 10. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0126] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0127] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0128] An embodiment of the present application also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implement the large model training method based on data security as described in any of the above embodiments.

[0129] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved, and no limitation is made herein.

[0130] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A large model training method based on data security, characterized in that: include: Divide the data into different security levels, and assign the data of each security level to the corresponding security data shards. Each security data shard is used as a training data set. During the model training process, each training data set is used to fine-tune the basic large model to obtain a sub-model of the corresponding security level.

2. The method according to claim 1, characterized in that Each training data set is used to fine-tune the basic large model to obtain sub-models of corresponding security levels, including: For each of the ordered security levels, using the training data set corresponding to the security level, fine-tune the basic large model corresponding to the security level to obtain a sub-model corresponding to the security level; Among them, the basic large model corresponding to the first security level is the general large model. The basic large model corresponding to the security level other than the first security level is the basic large model corresponding to the previous security level.

3. The method according to claim 1, characterized in that Also includes: Locate the secure data shard where the data to be forgotten is located; Deleting the to-be-forgotten data to update the corresponding training data set; Update the sub-models of the corresponding security levels according to the updated training data set.

4. The method according to claim 3, characterized in that Update the sub-model of the corresponding security level according to the updated training data set, including one of the following: According to the updated training data set, the sub-model of the corresponding security level is fine-tuned to update the sub-model of the corresponding security level; The sub-model of the corresponding security level is used as the teacher model, and the student model is obtained by training based on the teacher model and the updated training data set as the updated sub-model of the corresponding security level.

5. The method according to claim 3, characterized in that: Also includes: For any of the security levels, determining model parameters related to the data to be forgotten according to the difference between the sub-models before and after the security level is updated; The model parameters are pruned or adjusted.

6. The method according to claim 3, characterized in that Also includes: For any of the security levels, evaluating the updated sub-models of the security level; Publish the updated sub-models that have passed the evaluation to the model version library.

7. A large model training device based on data security, characterized in that: include: A partitioning module is used to divide the data into different security levels and allocate the data of each security level to the corresponding security data shards. Each security data shard is used as a training data set. The training module is used to fine-tune the basic large model using each training data set during the model training process to obtain a sub-model of the corresponding security level.

8. An electronic device, characterized in that: include: at least one processor; a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the large model training method based on data security as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the large model training method based on data security as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program and / or instructions, characterized in that: When the computer program and / or instructions are executed by the processor, the large model training method based on data security as described in any one of claims 1-6 is implemented.