Model training method and device based on cloud environment, equipment and medium

By processing sensitive data in a private cloud environment and transmitting public data to the public cloud for preliminary training, the problem of high security and cost of training data in a specific industry model is solved, and the effect of reducing training costs and ensuring data security is achieved.

CN120146156APending Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311717887.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Model training in specific industries requires a large amount of data in sensitive industry fields, but these data are difficult to train externally. At the same time, the model training process consumes time and resources, resulting in high costs.

Method used

Obtain and process industry data in a private cloud environment, divide the data into two parts: public and non-public, and transmit the public data to the public cloud environment for preliminary training, use the elasticity of public cloud resources to reduce hardware costs, and conduct final training in a private cloud environment to ensure data security.

Benefits of technology

On the premise of ensuring industry data security, the cost of model training is reduced, the resource elasticity of public clouds is fully utilized, and the resource burden of private clouds is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146156A_ABST
    Figure CN120146156A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and device based on a cloud environment, equipment and a medium. The method comprises the steps of obtaining industry data of a target industry in a private cloud environment; dividing the industry data according to the security level of the industry data to obtain public data and non-public data, and transmitting the public data from the private cloud environment to the public cloud environment; obtaining a target basic model, and in the public cloud environment, adjusting the target basic model according to the public data to obtain a first-level industry model; and adjusting the first-level industry model according to non-public data in the private cloud environment to obtain a target industry model, wherein the target industry model is used for processing a task of a target industry. According to the technical scheme provided by the embodiment of the invention, the cost of the training model can be further reduced on the premise of ensuring the security of industry data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer and communication technologies. Specifically, it relates to a model training method based on a cloud environment, a model training apparatus based on a cloud environment, an electronic device, and a computer-readable storage medium. Background Art

[0002] The emergence of models has brought great opportunities for the intelligent development of all walks of life. In order to focus more on a specific industry and meet the needs of the corresponding industry, models used in specific industries are developed. Currently, models used in specific industries require a large amount of industry domain data as a basis. However, a lot of industry domain data is very sensitive and cannot be submitted to the outside for training. Moreover, the training of models is very time-consuming and resource-consuming, resulting in a relatively high training cost of the models. Summary of the Invention

[0003] Embodiments of this application provide a model training method based on a cloud environment, a model training apparatus based on a cloud environment, an electronic device, a computer-readable storage medium, and a computer program product, which can further reduce the cost of training a model while ensuring the security of industry data.

[0004] Other features and advantages of this application will become apparent through the following detailed description, or will be partially learned through the practice of this application.

[0005] In a first aspect, embodiments of this application provide a model training method based on a cloud environment, including: obtaining industry data of a target industry in a private cloud environment; dividing the industry data into public data and non-public data according to the security level of the industry data, and transmitting the public data from the private cloud environment to a public cloud environment; obtaining a target basic model, and in the public cloud environment, adjusting the target basic model according to the public data to obtain a first-level industry model; adjusting the first-level industry model according to the non-public data in the private cloud environment to obtain a target industry model, where the target industry model is used to process tasks of the target industry.

[0006] Second aspect, an embodiment of the present application further provides a model training device based on a cloud environment. The device includes: an acquisition module, configured to acquire industry data of a target industry in a private cloud environment; a division module, configured to divide the industry data into public data and non-public data according to the security level of the industry data, and transmit the public data from the private cloud environment to a public cloud environment; an adjustment module, configured to acquire a target basic model, and in the public cloud environment, adjust the target basic model according to the public data to obtain a first-level industry model; the adjustment module is further configured to adjust the first-level industry model according to the non-public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process tasks of the target industry.

[0007] Third aspect, an embodiment of the present application provides an electronic device, including one or more processors; a storage device, configured to store one or more computer programs, and when the one or more computer programs are executed by the one or more processors, the electronic device implements the model training method based on a cloud environment as described above.

[0008] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of an electronic device, the electronic device executes the model training method based on a cloud environment as described above.

[0009] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, the computer program is stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the model training method based on a cloud environment as described above.

[0010] In the technical solution provided by the embodiments of the present application, industry data of a target industry in a private cloud environment is obtained; the industry data is divided into public data and non-public data according to the security level of the industry data, and the public data is transmitted from the private cloud environment to the public cloud environment; a target basic model is obtained, and in the public cloud environment, the target basic model is adjusted according to the public data to obtain a first-level industry model; the first-level industry model is adjusted according to the non-public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process tasks of the target industry; that is, in the technical solution provided by the present application, on the one hand, the private cloud environment is used for the final training of the model, allowing non-public data to be always stored and processed in the private cloud environment without being provided externally, ensuring data security; on the other hand, the public data is transmitted to the public cloud environment, and the public cloud environment is used for the preliminary training and adjustment of the model, making full use of the resource elasticity of the public cloud without bearing the hardware cost required for construction and training, and further reducing the cost of training the model on the premise of ensuring the security of industry data.

[0011] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings

[0012] Figure 1 is a schematic diagram of an implementation environment related to the present application;

[0013] Figure 2 is a flowchart of a model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0014] Figure 3 is a flowchart of a model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0015] Figure 4 is a schematic diagram of a detection result and an object distribution diagram shown in an exemplary embodiment of the present application;

[0016] Figure 5 is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0017] Figure 6 is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0018] Figure 7 is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0019] Figure 8It is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0020] Figure 9 It is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0021] Figure 10 It is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0022] Figure 11 It is a flowchart of another model training method based on a cloud environment shown in an exemplary embodiment of the present application;

[0023] Figure 12 It is a flowchart of a model training method based on a private cloud and a public cloud shown in an exemplary embodiment of the present application;

[0024] Figure 13 It is a flowchart of a model training method shown in another exemplary embodiment of the present application;

[0025] Figure 14 It is a structural block diagram of a model training device based on a cloud environment shown in an exemplary embodiment of the present application;

[0026] Figure 15 It shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0029] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0030] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0031] It should also be noted that: "a plurality of" mentioned in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally means that the associated objects before and after are in an "or" relationship.

[0032] The technical solution of the embodiment of this application relates to the field of Artificial Intelligence (AI) technology. Before introducing the technical solution of the embodiment of this application, AI technology will be briefly introduced. AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, method, technology, and application system. In other words, AI is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.

[0033] Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specializes in studying how a computer simulates or implements human learning behavior to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of AI and the fundamental way to make a computer intelligent. Its applications cover all fields of AI. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0034] The technical solution of the embodiment of the present application specifically relates to the machine learning technology in AI. Specifically, it is based on the machine learning technology to implement the training of industry models. The following is a detailed introduction to the technical solution of the embodiment of the present application:

[0035] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an implementation environment involved in the present application. This implementation environment includes a model training party 10, a private cloud environment 20, and a public cloud environment 30. The private cloud environment belongs to the model training party.

[0036] The model training party obtains the industry data of the target industry in the private cloud environment, divides the industry data into public data and non-public data according to the security of the industry data, and transmits the public data from the private cloud environment to the public cloud environment; then obtains the target basic model, and in the public cloud environment, adjusts the parameters of the target basic model according to the public data to obtain a first-level industry model; adjusts the parameters of the first-level industry model according to the non-public data in the private cloud environment to obtain the target industry model.

[0037] The model training party can apply the target industry model to task processing to process the tasks corresponding to the target industry.

[0038] The aforementioned model training party includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.; the model training party can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and intelligent platforms. This is not limited here.

[0039] The embodiment of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.; for example, when the target industry is the vehicle industry, obtain the assisted driving data of the vehicle industry in the private cloud environment, divide it into public data and non-public data according to the security level, such as the non-public data is the driver's identity information, and the public data is the driving habit, and then train the target industry model in the public cloud environment and the private cloud environment. This target industry model is used to predict the activation method and assistance method of assisted driving.

[0040] It should be noted that in the specific implementation manner of the present application, if the industry data involves objects, when the embodiment of the present application is applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0041] The following elaborates in detail on various implementation details of the technical solution of the embodiments of the present application:

[0042] As Figure 2 shown Figure 2 is a flowchart of a model training method based on a cloud environment shown in an embodiment of the present application. This method can be applied to Figure 1 the implementation environment shown. This method can be executed by the model training party. It should be understood that this method can also be applicable to other exemplary implementation environments and be specifically executed by devices in other implementation environments. The embodiments of the present application do not limit the implementation environment applicable to this method. In the embodiments of the present application, taking the execution of this method by the model training party as an example, the model training method based on the cloud environment may include steps S210 to S240, which are introduced in detail as follows:

[0043] S210. Obtain industry data of the target industry in the private cloud environment.

[0044] In the embodiments of the present application, the data in the private cloud environment is managed by the party to which the private cloud environment belongs. This data is used solely by the party to which the private cloud environment belongs to protect the privacy and security of the data. The target industry in the private environment can be any industry or a specific industry. The target industry can also include one or more industries. The industry data of the target industry includes various types of data related to the target industry, such as text data, voice data, and image data, etc. For example, if the target industry is the financial industry, the industry data can be the financial data, market data, transaction data, etc. of a certain company. If the target industry is the medical industry, the industry data can be medical imaging data and laboratory verification data, etc.

[0045] It should be noted that the acquisition and use of the data in the private cloud environment require authorization from the data owner to ensure the legality and security of the data.

[0046] In one example, the private cloud environment belongs to the model training party. Therefore, the model training party can directly access all the data in the private cloud environment and then query and obtain the industry data of the target industry according to keywords.

[0047] S220. Divide the industry data into public data and non - public data according to the security level of the industry data, and transfer the public data from the private cloud environment to the public cloud environment.

[0048] It is understandable that industry data in many industries may be very sensitive due to data privacy. If it is provided externally, the security of the data cannot be guaranteed. Therefore, in the embodiments of the present application, industry data is divided into public data and non-public data according to the security of the industry data. Among them, public data refers to data that does not require confidentiality processing or has obtained relevant authorization at the time of disclosure. These data can be publicly obtained and used by the public and do not contain any sensitive information or confidential content. For example, public data includes merchant review data, real-time bus data, etc.; non-public data refers to information that cannot be publicly obtained by the general public, such as financial data, market data, etc. It involves the trade secrets and core competitiveness of enterprises, so it needs to be protected and managed.

[0049] The security level of industry data can reflect the degree of security guarantee requirements for the data, that is, the higher the degree of security guarantee requirements for the data, the higher its security level. Therefore, industry data can be divided into public data and non-public data based on the security level of the industry data. For example, industry data with a security level greater than the preset level threshold is regarded as public data, and industry data with a security level less than or equal to the preset level threshold is regarded as non-public data.

[0050] In one example, the security level of industry data can be set by the party to which the industry data belongs, or determined according to the data content of the industry data, which is not limited here.

[0051] Since public data can be publicly obtained and used by the public, the public data can be transferred from the private cloud environment to the public cloud environment, while non-public data is always stored and processed in the private cloud environment; among them, the public cloud environment refers to the one that can be used provided by a third-party provider for the general public or enterprises.

[0052] In one example, the public data can be transferred from the private cloud environment to the public cloud environment through a dedicated line, or exported from the private cloud environment and then uploaded to the public cloud environment.

[0053] S230. Obtain a target basic model, and in the public cloud environment, adjust the target basic model according to the public data to obtain a first-level industry model.

[0054] In the embodiments of the present application, the basic model refers to some general, widely covered, and large-scale deep learning models. They are usually pre-trained on a large amount of data and have strong feature extraction and representation capabilities, and can be used for different tasks and fields; while the target basic model is a model selected from multiple basic models that conforms to the target industry, and the target basic model can also be composed of two basic models combined.

[0055] In one example, the steps of obtaining the target basic model and obtaining the industry data can be carried out simultaneously.

[0056] The public cloud environment can provide various resources such as computing, storage, and networking. Therefore, in the public cloud environment, the target basic model can be adjusted using the public cloud resources and public data provided by it to obtain a first-level industry model. Here, the adjustment of the model refers to, based on a pre-trained model, using some new data to adjust and optimize the model's parameters, so that the model can better adapt to the new data and tasks. The adjustment of the model can improve the model's performance and generalization ability, and can also save the time and cost of training the model from scratch. In one example, adjusting the model is model fine-tuning. When the model training party uses the resources in the public cloud environment for model fine-tuning, it needs to pay according to demand.

[0057] The process of adjusting the model includes: initializing the model weights using the pre-trained weights of the target basic model to retain the knowledge learned by the model in previous tasks. Then, an appropriate loss function is selected, which matches the task type of the target industry. For classification tasks, cross-entropy loss can be selected. By inputting the data into the model, calculating the loss, and adjusting the model parameters through backpropagation. In one example, during the adjustment process, hyperparameters such as the learning rate and batch size may need to be adjusted to achieve better performance.

[0058] S240. Adjust the first-level industry model according to the non-public data in the private cloud environment to obtain a target industry model, which is used to process the tasks of the target industry.

[0059] The first-level industry model already has some industry knowledge understanding ability. To improve the model's accuracy, the first-level industry model needs to be optimized. As previously described, the non-public data is still stored in the private cloud environment. Therefore, the private cloud environment is used for re-training the model, that is, re-adjusting the first-level industry model according to this non-public data. Re-adjustment means, based on the first-level industry large model, using the non-disclosed data on the private cloud environment to further adjust and optimize the model's parameters, so that the model can better adapt to the specific data and tasks of the target industry.

[0060] In one example, after obtaining the first-level industry model, it is necessary to transfer the first-level industry model from the public cloud environment to the private cloud environment.

[0061] In one example, after obtaining the target industry model, the target industry model can be deployed to the corresponding business scenario to process the tasks of the target industry.

[0062] In an embodiment of the present application, industry data of a target industry is obtained in a private cloud environment; industry data is divided into public data and non-public data according to the security level of the industry data, and the public data is transmitted from the private cloud environment to the public cloud environment; a target basic model is obtained, and in a public cloud environment, the target basic model is adjusted according to the public data to obtain a first-level industry model; the first-level industry model is adjusted according to the non-public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process tasks of the target industry; that is, in the technical solution provided by the present application, on the one hand, a private cloud environment is used for final training of the model, allowing non-public data to be always stored and processed in the private cloud environment without being provided to the outside, thereby ensuring data security; on the other hand, public data is transmitted to a public cloud environment, and the public cloud environment is used to perform preliminary training and adjustment of the model, making full use of the resource elasticity of the public cloud without having to bear the hardware cost required for training construction, thereby further reducing the cost of training the model while ensuring the security of industry data.

[0063] In one embodiment of the present application, another cloud-based model training method is provided, which can be applied to Figure 1 Implementation environment, this method can be executed by the model training party, such as Figure 2 As shown in Figure 2, the model training method based on the cloud environment Figure 2 Based on S210 to S240 shown in Figure 2 The step S240 shown in FIG. 1 is expanded into steps S310 to S330. Steps S310 to S330 are described in detail as follows:

[0064] S310. In a public cloud environment, the first-level industry model is compressed to obtain a second-level industry model.

[0065] The first-level industry model still retains parameters of the same scale as the target basic model, and the number of resources in the private cloud environment is effective. In order to reduce the resource burden of the private cloud environment, in the embodiment of the present application, in the public cloud environment, it is necessary to perform model compression processing on the first-level industry model to obtain a second-level industry model. Model compression processing refers to some technologies that can reduce the number of model parameters, reduce the complexity of model calculations, and improve the efficiency of model operation. It can reduce the size and computing requirements of the model while maintaining the accuracy of the model as much as possible.

[0066] S320. Transfer the secondary industry model from the public cloud environment to the private cloud environment, and build a training environment for the secondary industry model in the private cloud environment.

[0067] Since the non-public data is stored in the private cloud environment, it is necessary to transfer the secondary industry model from the public cloud environment to the private cloud environment, which can be transferred to the private cloud environment through a dedicated line, or exported from the public cloud environment and transferred offline to the private cloud environment.

[0068] In the embodiment of the present application, in order to ensure the secure and efficient training of the model, a training environment for the secondary industry model can be built in the private cloud environment. The training environment includes a hardware environment and a software environment. Among them, the hardware environment includes multi-core CPUs, large-capacity memories, high-speed memories, etc., and the software environment includes data processing tools, object storage, etc.

[0069] In one example, in order to ensure that the model is not attacked during transmission, resulting in the leakage of data in the private cloud environment, it is necessary to first encrypt the secondary industry model to obtain an encrypted model, and transfer the encrypted model to the private cloud environment through a dedicated line to ensure the security of model transmission. Decrypt the encrypted model in the private cloud environment to obtain the secondary industry model. It is necessary to verify the transmitted secondary industry model to ensure that the model transmission has not been attacked; after the transmitted secondary industry model passes the verification, in order to ensure the security of the model training process, divide an isolation space for the secondary industry model in the private cloud environment, and build a training environment in the isolation space to avoid the training of this model from affecting other data in the private cloud environment. Among them, only trusted nodes can access the training environment to reduce the risk of network attacks.

[0070] In one example, when the obtained secondary industry model is in the public cloud environment, a hash calculation can be performed on the secondary industry model to obtain a hash value, and the secondary industry model is signed. The signature, hash value, and secondary industry model are encrypted to obtain an encrypted model. Then, in the private cloud environment, after decryption, first verify the signature to ensure that it comes from the public cloud environment. After the signature passes, perform a hash calculation on the decrypted secondary industry model to obtain a hash value, and then match the decrypted hash value to determine that the secondary industry model has not been tampered with during transmission, thereby completing the verification of the secondary industry model.

[0071] S330. In the training environment, adjust the secondary industry model according to the non-public data to obtain the target industry model.

[0072] In the embodiments of the present application, in a training environment, non-public data is input into a secondary industry model, and the target industry model for adjusting the parameters of the model. Among them, when adjusting the secondary industry model, iterative training can be performed according to the actual situation until the model accuracy and performance indicators that meet business requirements are achieved. In one example, when the model training party accesses non-public data in a controlled private cloud environment, verification needs to be performed first to ensure that only authorized parties can access it to protect data security. In one example, after adjusting according to non-public data, it is also possible to detect whether there are modification operations on the non-public data to discover potential security threats in a timely manner.

[0073] It should be noted that Figure 3 For other detailed introductions of steps S210 to S230 shown in Figure 2 Please refer to steps S210 to S230 shown in

[0074] In the embodiments of the present application, a compression operation is performed on the adjusted model on a public cloud, and the compressed model is transmitted to a private cloud environment for adjustment, reducing the number of private cloud resources required for subsequent model adjustment, so as to relieve the pressure and training cost of the private cloud, and at the same time meet the security requirements of non-public data.

[0075] The embodiments of the present application provide another model training method based on a cloud environment. This model training method based on a cloud environment can be applied to Figure 1 the implementation environment of Figure 4 As shown in Figure 3 On the basis shown in

[0076] S410. In a public cloud environment, connections or parameters with a contribution degree less than a preset contribution degree threshold are removed from the primary industry model to obtain a primary industry model.

[0077] In the embodiments of the present application, the model compression process is performed in a public cloud environment. After adjusting the parameters of the target basic model to obtain a primary industry model, the contribution degree of each connection or parameter in the primary industry model to the model performance can be calculated, and then connections or parameters with a contribution degree less than the preset contribution degree threshold are removed. For example, connections or parameters with a contribution degree less than 5% are deleted, which can effectively reduce the overfitting risk of the model, improve the generalization ability of the model, and can greatly reduce the number of model parameters without affecting the model accuracy.

[0078] In one example, the contribution of each parameter to the loss function can be measured by gradient importance. For example, when training the base model, forward propagation is performed to calculate the loss, and backward propagation is performed to calculate the gradients. For each connection or parameter, the absolute value in the gradient vector is calculated. The larger the absolute value, the greater the impact of the connection or parameter on the loss. Then, the absolute values of the gradients of each connection or parameter are normalized to ensure comparison between different parameters, and the absolute value of the gradient is used as the contribution degree.

[0079] In one example, for each connection or parameter, the coefficient of its L1 regularization term can be regarded as its contribution degree; the information entropy can also be used to measure the uncertainty of the connection or parameter to the model. For example, for each connection or parameter, calculate its impact on the model output and represent it in the form of information entropy.

[0080] It can be understood that different types of models may require different contribution degree calculation methods. For example, for deep neural networks, gradient analysis may be more common; for neural networks with non-linear activation functions, the importance of connections or parameters can be evaluated by analyzing the derivatives of the activation functions. For example, for the ReLU activation function, if the activation output of a connection is zero, then the corresponding connection may be less important.

[0081] S420. Quantize the primary industry model to compress the model parameters of the primary industry model from floating-point representation to low-precision integers.

[0082] Quantization processing converts the parameters of the model from high-precision floating-point numbers (such as 32-bit or 64-bit) to low-precision integers (such as 8-bit or 16-bit), reducing the model size.

[0083] The quantization processing in the embodiments of the present application includes weight quantization and activation quantization. Among them, weight quantization mainly quantizes the weights of the neural network to reduce the model size, usually represented by integers with lower bits; activation quantization quantizes the activations of the neural network, usually also represented by low-precision integers. Activation quantization may have a greater impact on the accuracy of the model and requires careful selection. Mixed precision uses different precision bits for quantization. For example, weights are represented by low-precision integers, while activations are represented by higher-precision integers. This method provides a certain balance between accuracy and speed. Therefore, in the embodiments of the present application, an appropriate quantization algorithm can be selected according to the model structure. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) may have different sensitivities to weight quantization and activation quantization; an appropriate quantization algorithm can also be selected according to the business scenario.

[0084] S430. Use the computing resources of the public cloud environment to evaluate the quantized primary industry model, and determine whether to adjust the compressed model parameters according to the evaluation results to obtain a secondary industry model.

[0085] In an embodiment of the present application, after the model is compressed, the computing resources of the public cloud environment can be utilized to evaluate the quantized primary industry model using a test set, ensuring that the model has a certain accuracy and the performance loss is within an acceptable range. The evaluation method can select appropriate evaluation metrics and data sets according to different tasks and fields, such as accuracy and recall rate. For example, if accuracy is selected as the evaluation metric, and the evaluation result indicates that the accuracy of the initial industry model after quantization is lower than the preset threshold compared to the accuracy of the initial industry model before quantization, it means that the accuracy and performance loss of the model exceed the acceptable range. At this time, the model compression process needs to be adjusted, such as restoring the compressed model parameters. For example, restoring some of the quantized model parameters from low-precision integers to high-precision integers to obtain a secondary industry model; or restoring the deleted connections or parameters to obtain a secondary industry model.

[0086] If the evaluation result indicates that the accuracy and performance loss of the model are within the acceptable range, there is no need to adjust the compressed model parameters, and the initial industry model after quantization is directly used as the secondary industry model.

[0087] It should be noted that Figure 4 For the detailed introduction of steps S210 - S230 and S320 - S330 shown in Figure 3 Please refer to steps S210 - S230 and S320 - S330 shown in

[0088] In an embodiment of the present application, by removing connections or parameters with a contribution degree less than the preset contribution degree threshold, the number of model parameters can be significantly reduced without affecting the model accuracy, thereby reducing the complexity and computational amount of the model. On this basis, the model is quantized to further reduce the storage space and computing resources of the model, and it is determined whether to adjust the compressed model parameters according to the evaluation result, ensuring that the secondary industry model has a certain accuracy and the performance loss is within the acceptable range.

[0089] The embodiment of the present application also provides another cloud environment-based model training method. This cloud environment-based model training method can be applied to Figure 1 the implementation environment shown in Figure 5 Taking the example that this method is executed by the model trainer, as Figure 3Based on what is shown, S330 is extended to S510 - S530. Among them, the number of private cloud environments in S210 includes multiple, and the non - public data in each private cloud environment is different. For example, there are private cloud environment 1, private cloud environment 2, and private cloud environment 3. The industry data of the target industry in each private cloud environment is divided respectively to obtain the corresponding non - public data 1 - 3 in private cloud environments 1 - 3, and the non - public data 1 - 3 are different. Steps S510 - S530 are introduced in detail as follows:

[0090] S510. In the training environments deployed in each private cloud environment, adjust the secondary industry model respectively according to the local non - public data in each private cloud environment.

[0091] S520. Aggregate the model update parameters of each private cloud environment for the secondary industry model to obtain the target model update parameters.

[0092] S530. Adjust the secondary industry model according to the target model update parameters to obtain the target industry model.

[0093] In order to enable the target industry model to learn more comprehensive features, in the embodiments of this application, the target industry model is adjusted using decentralized data. In each private cloud environment, each private cloud environment stores part of the non - public data. In each private cloud environment, the secondary industry model is adjusted locally. After local training is completed, each private cloud environment can elect a primary private cloud environment, and other private cloud environments are used as secondary private cloud environments. Then, the secondary private cloud environments transmit the model update parameters obtained after adjusting the secondary industry model to the primary private cloud environment, and the primary private cloud environment aggregates the model update parameters transmitted by each secondary private cloud environment to obtain the target model update parameters. Then, the secondary industry model is adjusted according to the target model parameters, that is, the model parameters of the secondary industry model are adjusted to the target model parameters to obtain the target industry model. Among them, aggregating the model update parameters corresponding to each private cloud environment can be taking the mean value of each model update parameter as the target model update parameter.

[0094] In one example, electing a primary private cloud environment from each private cloud environment can be randomly selecting a private cloud environment as the primary private cloud environment, or determining the security level of the non - public data in each private cloud environment, and taking the private cloud environment to which the non - public data with the highest security level belongs as the primary private cloud environment.

[0095] In one example, steps S510 - S530 are an iterative process, that is, the secondary industry model is iteratively trained in each private cloud environment. After each round of training, the model parameter updates are sent to the primary private cloud environment for aggregation. This process is repeated until the model converges or reaches the predetermined number of training rounds to obtain the target industry model.

[0096] It should be noted that Figure 5 For other detailed descriptions of steps S210 to S230 and S310 to S320 shown in Figure 3 The steps S210 to S230 and S310 to S320 shown in FIG. 1 are not described in detail here.

[0097] In an embodiment of the present application, in cross-cloud collaboration, there may be scattered data resources in different private clouds. The use of federated learning can enable the model to obtain information from various data islands while protecting industry data privacy, and make full use of these scattered data for model training, so that the model can learn more comprehensive features from data in different environments.

[0098] In one embodiment of the present application, another cloud-based model training method is provided. The cloud-based model training method can be applied to Figure 1 The implementation environment shown in the figure takes the method executed by the model training party as an example. Figure 6 As shown in Figure 2, the model training method based on the cloud environment Figure 5 Based on the information shown in , S510 is expanded to S610-S620; steps S610-S620 are described in detail as follows:

[0099] S610: Perform data anonymization processing on designated words in the local non-public data of each private cloud environment, and add preset noise to the non-public data after the data anonymization processing to obtain target non-public data.

[0100] In an embodiment of the present application, in order to ensure that data privacy is not disclosed during the model adjustment process, for the local non-public data of each private cloud environment, the designated words in the non-public data are anonymized. The designated words can be customized by the model training party. The data anonymization processing includes replacement processing, that is, replacing the designated words with other words that do not involve privacy. For example, if the non-public data is financial data, the designated words are specific financial values, which are replaced with "target values". It also includes generalization processing, replacing specific data values ​​with more general data values. For example, the original data is 24 years old, and the data after generalization processing is 20-30 years old; the data anonymization processing also includes data replacement, that is, by replacing the value in one record with the corresponding value of another record.

[0101] In one example, when performing data anonymization processing on designated words in non-public data, it is necessary to first determine the security level of the non-public data, and only when the security level reaches the designated level, perform data anonymization processing on the designated words therein.

[0102] Furthermore, a stronger privacy protection mechanism is introduced. Preset noise is added to the non-public data after data anonymization to obtain target non-public data, so as to prevent the leakage of sensitive information during the model adjustment process. Among them, adding preset noise to the non-public data after data anonymization can be randomly introducing preset noise into the data, or inputting the data into the corresponding model to add noise by the model.

[0103] It should be noted that, in order to ensure that the target non-public data can enable the secondary industry model to learn effective features, when adding noise, the value of this noise should be less than the preset noise threshold, and the preset noise threshold can be flexibly adjusted according to the quantity of the non-public data.

[0104] S620. In the training environment deployed in each private cloud environment, adjust the secondary industry model according to the target non-public data.

[0105] In the embodiment of the present application, in the training environment deployed in each private cloud environment, adjusting the secondary industry model according to the local target non-public data increases the privacy protection during the model training process.

[0106] It can be understood that S610 and S620 can also be executed Figure 3 on the basis of S330.

[0107] It should be noted that Figure 6 For other detailed introductions of the steps S210~S230, S310~S320, S520~S530 shown in Figure 5 please refer to the steps S210~S230, S310~S320, S520~S530 shown in

[0108] In the embodiment of the present application, through data anonymization processing with specified words and adding preset noise to obtain target non-public data, it is ensured to prevent the leakage of sensitive information during the model adjustment process and enhance the privacy protection mechanism.

[0109] In an embodiment of the present application, another model training method based on the cloud environment is also provided. The model training method based on the cloud environment can be applied to Figure 1 the implementation environment shown in Figure 7 Taking the example that this method is executed by the model training party, as Figure 2 shown, on the basis of what is shown in

[0110] S710. Obtain a target basic model, deploy a training environment for the target basic model in a public cloud environment, and in the training environment, adjust the target basic model according to public data to obtain an initial model.

[0111] In an embodiment of the present application, create a virtual machine or container suitable for deep learning tasks in a public cloud environment, ensure that the configuration meets the resource requirements for training, such as including GPUs, etc., and then install and configure the required deep learning frameworks, such as TensorFlow, etc., to obtain a training environment. Finally, upload the target basic model to the public cloud environment to ensure that it can be loaded in the training environment. Furthermore, in the training environment, input the public data into the target basic model and adjust the target basic model so that the target basic model adapts to the tasks corresponding to the public data.

[0112] S720. Test the initial model according to multiple business scenario tasks to determine the coverage of the initial model for business scenarios.

[0113] S730. Iteratively adjust the initial model according to the coverage to obtain a first-level industry model.

[0114] It can be understood that the obtained initial industry model after adjustment already has some industry knowledge understanding capabilities and can be used for some industry-related tasks. Therefore, in a public cloud environment, for specific industry scenario tasks, the model effects and performance of the primary industry model can be preliminarily tested and evaluated, and iterative training of the model can be carried out. Among them, this round of training should focus on the coverage of the model for business scenarios, and let the primary industry model achieve better application effects as much as possible under the training of publicly available datasets to obtain a first-level industry model.

[0115] In an example, obtain test data for multiple business scenario tasks. For example, the test data for normal business scenario tasks includes data on normal business routine operations to reflect the situation of normal model use; the test for abnormal business scenario tasks includes simulating possible abnormal or error operations to test the model's ability to handle abnormal situations; new business scenario tasks can also be set to test the model's adaptability to new scenarios; input the test data for multiple business scenario tasks into the initial model, select corresponding evaluation metrics according to business requirements, such as accuracy, recall rate, precision, F1 score, etc. Assuming that the F1 score is selected as the evaluation metric, obtain the F1 score of the initial model under different business scenario tasks, and use this F1 score as the coverage of the initial model for business scenarios.

[0116] If the coverage of the business scenario is less than the preset degree threshold, the initial model is iteratively adjusted until the coverage is greater than or equal to the preset degree threshold to obtain a first-level industry model; among them, during the iterative adjustment, data augmentation processing can be performed on the public data, and the initial model is iteratively adjusted based on the data after the data augmentation processing.

[0117] It should be noted that Figure 7 For other detailed introductions of steps S210 - S220 and S240 shown in Figure 2 Please refer to steps S210 - S220 and S240 shown in

[0118] In the embodiments of the present application, the preliminarily adjusted model is tested through multiple business scenario tasks, paying attention to the coverage of the business scenario by the model, and enabling the first-level industry large model to achieve a better application effect as much as possible under the training of the publicly available dataset.

[0119] In an embodiment of the present application, another model training method based on the cloud environment is also provided. The model training method based on the cloud environment can be applied to Figure 1 the implementation environment shown in Figure 8 Taking the execution of this method by the model training party as an example, as Figure 7 shown, on the basis of what is shown in

[0120] S810. Obtain a target basic model, deploy a training environment for the target basic model in the public cloud environment, and select a corresponding adjustment strategy according to the annotation situation of the public data and the task type corresponding to the target industry.

[0121] S820. In the training environment, adjust the target basic model according to the adjustment strategy and the public data to obtain an initial model.

[0122] For different task types and different data annotation situations, different adjustment strategies can be selected; in the embodiments of the present application, a corresponding first relationship table between the data annotation situation and the adjustment strategy is preset, and a second relationship table between the task type and the adjustment strategy is also preset. Then, the annotation situation of the public data is matched with the first relationship table to obtain a first adjustment strategy, and the task type corresponding to the target industry is matched with the second relationship table to obtain a second adjustment strategy. The intersection adjustment strategy in the first adjustment strategy and the second adjustment strategy can be used as the final adjustment strategy, or one of the first adjustment strategy and the second adjustment strategy can be selected as the final adjustment strategy according to business requirements.

[0123] Among them, the first relationship table includes, for example, the task is image classification, and the applicable adjustment strategies include standard Stochastic Gradient Descent (SGD), SGD with momentum, Adam, etc., which usually adjust the underlying convolutional layers of the model. For the object detection task, it is necessary to select adjustment strategies that support the object detection task, such as Faster R-CNN, YOLO, etc., and use appropriate loss functions. For text tasks in natural language processing, the adjustment strategy can use optimizers based on the Transformer architecture, such as AdamW.

[0124] The second relationship table includes that for publicly available labeled data, supervised learning algorithms can be used for adjustment, and common choices include SGD, Adam, etc.; for publicly available unlabeled data, some unsupervised fine-tuning algorithms can be selected, such as MLM, ELECTRA, etc.

[0125] After selecting the adjustment strategy, the publicly available data is input into the target base model, and the adjustment strategy is used to adjust the target base model to obtain an initial model.

[0126] It should be noted that Figure 8 For other detailed introductions of steps S210 - S220, S720 - S730, and S240 shown in Figure 7 Please refer to steps S210 - S220, S720 - S730, and S240 shown in

[0127] In the embodiments of the present application, according to the annotation situation of the publicly available data and the task type corresponding to the target industry, the corresponding adjustment strategy is selected, and then the target base model is adjusted based on the adjustment strategy to ensure the accuracy and reliability of model adjustment.

[0128] In an embodiment of the present application, another model training method based on the cloud environment is also provided. The model training method based on the cloud environment can be applied to Figure 1 the implementation environment shown in Figure 9 Taking the example that this method is executed by the model training party, as Figure 2 shown, on the basis of what is shown in

[0129] S910. Determine the security level of each industry data according to the source and content of the industry data.

[0130] In the embodiments of the present application, the sources of industry data include the upload sources of industry data uploaded to the private cloud environment and the sources of the industry data itself; among them, the sources of the industry data itself include data collection methods and generation channels; the upload sources include the uploader information of the uploaded data, including the roles and permissions of the uploader.

[0131] Determine the first security level of the industry data according to the source of the industry data, determine the second security level of the industry data according to the content of the industry data, and then determine the final security level of the industry data according to the first security level and the second security level. For example, set weights for the first security level and the second security level, and then perform weighted summation to obtain the final security level; or use the average value of the first security level and the second security level as the final security level of the industry data.

[0132] Among them, determining the first security level according to the source of the industry data includes: if the permission of the uploader who uploads the industry data is the highest permission, then determine the first security level as the highest level, that is, the permission is positively correlated with the security level; if the data collection method and generation channel of the uploaded industry data, then determine the corresponding security level of the industry data according to the corresponding relationship between the preset collection method and generation channel and the security level.

[0133] Determining the second security level according to the content of the industry data includes: first obtain the content belonging to the preset sensitive information in the industry data, including personal identity information, financial information, intellectual property rights, and trade secrets, etc., and determine the security level corresponding to the content belonging to the preset sensitive information as the high level; for the undetermined content that does not belong to the preset sensitive information, obtain the regulatory requirements related to the target industry, match the undetermined content with the regulatory requirements, if the match is successful, then determine the security level of the undetermined content as the medium level, and if the match fails, then determine the security level of the undetermined content as the low level.

[0134] S920. Use the industry data with a security level higher than the preset security level threshold as non-public data.

[0135] S930. Use the industry data with a security level lower than the preset security level threshold as undetermined industry data, obtain the sensitivity degree of the undetermined industry data, and determine the public data in the undetermined industry data according to the sensitivity degree.

[0136] In the embodiments of the present application, the industry data with a security level higher than the preset security level threshold includes sensitive information and needs to be highly protected, so it is used as non-public data; for the industry data with a security level lower than the preset threshold as undetermined industry data, further processing and classification are required. At this time, obtain the sensitivity degree of the undetermined industry data, and screen out the public data among them according to the sensitivity degree.

[0137] In one example, the sensitivity level of the to-be-determined industry data is determined according to the number of target words included in the to-be-determined industry data. The to-be-determined industry data is divided into data to be desensitized and public data according to the sensitivity level, and the data obtained after desensitizing the data to be desensitized is used as public data.

[0138] Among them, the target word is not a sensitive information word corresponding to the divided security level, and can be flexibly adjusted according to the actual situation; for example, the industry data is medical data, and the target words are "diagnosis", "treatment plan", etc.; the more the number of target words included, the higher the sensitivity level of the to-be-determined industry data. Furthermore, the to-be-determined industry data with a sensitivity level greater than the first preset degree threshold and less than the second preset degree threshold is used as the data to be desensitized, and the public data is obtained by desensitizing the data to be desensitized. The desensitization process includes, but is not limited to, methods such as encryption, replacement, and deletion; the to-be-determined industry data with a sensitivity level less than the first preset degree threshold is also used as public data.

[0139] It should be noted that Figure 9 For other detailed introductions of steps S210, S230 to S240 shown in Figure 2 Please refer to steps S210, S230 to S240 shown in

[0140] In the embodiments of the present application, the security level of the data can be accurately determined according to the content and source of the data, ensuring the accuracy and reliability of the division of industry data.

[0141] The embodiments of the present application provide another model training method based on a cloud environment. This method can be applied to Figure 1 the implementation environment shown in Figure 2 On the basis shown in

[0142] As Figure 10 shown, the details of S1010 to S1040 are as follows:

[0143] S1010. Obtain the data type of the industry data and the resource conditions for model deployment.

[0144] In the embodiments of the present application, it is necessary to determine the data type according to the data characteristics of the industry data. The data type includes different types such as text, image, time series, etc.; the resource conditions for model deployment include hardware resources (CPU, GPU, etc.), software frameworks (TensorFlow, PyTorch, etc.), deployment platforms (cloud, edge devices, etc.).

[0145] S1020. Determine the model structure corresponding to the data type according to the data type, and determine the parameter version corresponding to the resource conditions according to the resource conditions.

[0146] In the embodiments of the present application, different types of data require different model structures for processing. Therefore, it is necessary to select an appropriate model structure according to the data type. For example, for image data, a convolutional neural network (CNN) can be selected; for text data, a recurrent neural network (RNN) or Transformer, etc. can be selected.

[0147] Since models with different parameter scales require different computing capabilities and storage spaces, it is necessary to determine the corresponding parameter version according to the resource conditions for model deployment. If deployed in a resource-constrained environment, a version with smaller parameters may need to be selected. If resources are sufficient, then some large models with larger parameter scales, such as BERT-Large, GPT-3 Davinci, etc., can be selected.

[0148] S1030. Obtain the task type to which the industry data belongs, and determine the model output corresponding to the task type according to the task type.

[0149] In the embodiments of the present application, the task type to which the industry data belongs determines the output layer structure of the model and the selection of the loss function. The task type includes classification, regression, clustering, etc.; for a classification task, the output of the model may be a class probability distribution; for a regression task, the output of the model may be a continuous value.

[0150] S1040. Select a target base model from multiple base models according to the model structure, parameter version, and model output, and in a public cloud environment, adjust the target base model according to the public data to obtain a first-level industry model.

[0151] In the embodiments of the present application, a model structure matching the data type and task type can be selected from the base model library, and an appropriate version can be selected from the defined different parameter versions; considering the resource conditions and performance requirements, the model structure selected from the base model library is modified to configure the output layer structure of the model to ensure compliance with the task type, thereby obtaining the target base model.

[0152] It should be noted that Figure 10 For other detailed introductions of steps S210-S220 and S240 shown in Figure 2 Please refer to steps S210-S220 and S240 shown in

[0153] In the embodiments of the present application, a target basic model is obtained according to the data type of industry data, the resource conditions for model deployment, the task type, and the resource conditions, so as to improve the ability of the target basic model to effectively process corresponding tasks.

[0154] The embodiments of the present application provide another model training method based on a cloud environment. This method can be applied to Figure 1 the implementation environment shown. Taking the execution of this method by the model trainer as an example for illustration, as Figure 11 shown, after S240 in the model training method based on the cloud environment shown in Figure 2 add steps S1110 to S1120. Among them, steps S1110 to S1120 are introduced in detail as follows:

[0155] S1110. Obtain the deployment scenario of the target industry model, and determine whether to perform model compression processing on the target industry model according to the deployment scenario.

[0156] It can be understood that the trained target industry model needs to be deployed to a scenario for application. Therefore, it is necessary to obtain the deployment scenario of the target industry model, which includes device types (such as cloud servers, edge devices, mobile devices, etc.), network connectivity, computing resources, etc. The model trainer can receive the deployment request from the model deployer, and the deployment request carries the specific requirements and restrictions for deployment. Then, based on the specific requirements and restrictions for deployment, obtain the deployment requirements of the deployment scenario.

[0157] Determine whether to perform model compression processing on the target industry model according to the performance and resource limitations of the deployment scenario, the size and computational complexity of the model. If the size and computational complexity of the model exceed the performance and computational resource limitations of the deployment scenario, it is necessary to perform model compression processing on the target industry model.

[0158] S1120. If it is determined to perform model compression processing on the target industry model, deploy the model after model compression processing from the private cloud environment to the deployment scenario.

[0159] If it is determined to perform model compression processing on the target industry model, select a model compression method, such as weight quantization, pruning, distillation, etc., and set corresponding compression parameters according to the scenario requirements, such as the number of quantization bits, the pruning ratio, etc. Then, perform model compression processing according to the model compression method and compression parameters; after that, deploy the model after model compression processing from the private cloud environment to the deployment scenario.

[0160] In one example, when deploying a model to a deployment scenario, model conversion, editing, and loading can be performed. For example, the model can be converted from the original deep learning framework format (such as TensorFlow, PyTorch) to a format supported by the target deployment environment, such as the TensorFlow Lite format, the ONNX format, etc., using a selected tool; the model can be compiled into hardware-executable code using a selected compilation tool, and then loaded into the deployment scenario.

[0161] In other embodiments of the present application, if it is determined not to perform model compression processing on the target industry model, different parts of the target industry model are selected according to the device performance and requirements of the deployment scenario and loaded into the device of the deployment scenario. For example, if the overall model is not compressed, parts of the model are selected according to the scenario requirements. For example, only the first few layers of the model can be loaded, and more layers are gradually loaded according to the device performance to reduce the memory occupancy during the operation of the device in the deployment scenario.

[0162] It should be noted that Figure 11 For the detailed introduction of S210~S240 shown in Figure 2 S210~S240 shown in, which will not be elaborated here.

[0163] In the embodiments of the present application, when training the target industry model, before deployment and inference, a model compression operation can also be performed according to the deployment scenario to further reduce the model parameter scale, so as to facilitate embedding the model into more scenarios and devices to serve more objects that call the model service.

[0164] For the convenience of understanding, the embodiments of the present application also provide a model training method based on a cloud environment, as Figure 12 shown. This training method includes four parts. The first part is to sort out industry data to divide industry data into publicly available data and non-publicly available data; the second part is to select a basic model; the third part is public cloud training; and the fourth part is private cloud training.

[0165] Based on Figure 13 shown, the model training method based on a cloud environment provided by the embodiments of the present application, as Figure 13 shown, includes:

[0166] S1310. Sort out industry data and divide industry data into publicly available data and non-publicly available data.

[0167] In order to classify and manage industry data so as to reasonably use different types of data during subsequent model training while protecting data privacy and security. The sensitivity and security requirements of data can be judged based on factors such as data source, content, usage, laws and regulations, etc.; for example, some data involving personal information, business secrets, national interests, etc. belong to data that cannot be made public and need to be processed and trained in a private cloud environment; some data without sensitive information or that has been desensitized can be made public or processed and trained in a public cloud environment.

[0168] Among them, after classifying industry data, data preprocessing operations such as data cleaning, missing value handling, deduplication, data augmentation, etc. can be further performed on the data to improve the quality of training data.

[0169] S1320. Select a basic model according to the requirements of the business scenario.

[0170] According to the requirements of the business scenario, select a suitable basic large model. By selecting a general basic large model, the time and hardware and other cost inputs for pre-training the large model from scratch can be reduced. The following factors can be considered when specifically selecting a basic large model: for example, according to the type of data to be processed, select a language large model, a vision large model, a speech large model, etc.; the data type is an important factor affecting the selection of the basic large model because different types of data require different model structures and algorithms for processing. For example, if the industry data is text data, then some language large models can be selected; if the industry data is image data, then some vision large models can be selected; if the industry data is speech data, then some speech large models can be selected.

[0171] According to the task type of the business scenario, select the corresponding task-based large model (such as a customer service large model, a translation large model, etc.); the task type is another important factor affecting the selection of the basic large model because different tasks require different model outputs and evaluation metrics. For example, if the business scenario of the industry data is customer service, then some large models suitable for dialogue generation can be selected; if the business scenario of the industry data is translation, then some large models suitable for machine translation can be selected.

[0172] According to the hardware resource conditions for model deployment, select a large model with a suitable parameter version, etc. The hardware resource conditions are another important factor affecting the selection of the basic large model because large models with different parameter scales require different computing capabilities and storage spaces. For example, if the hardware resources for industry model deployment are limited, then some large models with smaller parameter scales can be selected; if the hardware resources of the industry object are sufficient, then some large models with larger parameter scales can be selected.

[0173] S1330. Upload the publicly available data to the public cloud and deploy the training environment for the basic model in the public cloud environment.

[0174] Utilize the resources and services of the public cloud to facilitate and enhance the efficiency of fine-tuning the basic large model. The advantage of the public cloud is that the user does not need to purchase and maintain their own hardware devices. They only need to pay as needed to use the computing, storage, network and other resources provided by the public cloud, as well as various cloud services such as cloud databases, cloud storage, and cloud functions. In the public cloud, the user can upload their publicly available training data to the cloud storage.

[0175] Among them, deploying the training environment in the public cloud environment can be to deploy the training environment of the basic large model on a cloud server, including installing the required software, configuring the required parameters, loading the required model, etc.; or choosing the large model training hosting service of the public cloud provider (the user does not need to worry about the specific implementation of training tools, underlying storage, etc., which are all implemented by the public cloud provider); or purchasing the resources of the public cloud (such as container computing resources, object storage, etc.) and the user deploys the training tools by themselves.

[0176] S1340. Utilize the public cloud training environment to fine-tune the basic model based on the data in the public cloud to obtain the first-level industry model.

[0177] Utilize the resources and services of the public cloud to conduct industry-related training on the basic large model so that it can adapt to industry data and tasks.

[0178] During this model fine-tuning, an appropriate fine-tuning algorithm (i.e., the aforementioned adjustment strategy) should be selected and carried out using the publicly available data for specific industries and specific scenarios prepared in the early stage. Currently, commonly used fine-tuning algorithms include supervised fine-tuning algorithms and parameter-efficient fine-tuning algorithms, etc.

[0179] For example, according to the annotation situation of the data, select a supervised fine-tuning algorithm or an unsupervised fine-tuning algorithm. The annotation situation of the data is an important factor affecting the selection of the fine-tuning algorithm because different annotation situations require different training methods and evaluation metrics. For example, if the industry data is annotated data, then some supervised fine-tuning algorithms such as GLUE, SQuAD, etc. can be selected; if the industry data is unannotated data, then some unsupervised fine-tuning algorithms such as MLM, ELECTRA, etc. can be selected.

[0180] According to the type of the task, select the corresponding task-class fine-tuning algorithm. The type of the task is another important factor affecting the selection of the fine-tuning algorithm, because different tasks require different model outputs and evaluation metrics. For example, if the task of industry data is text classification, then some fine-tuning algorithms suitable for text classification can be selected, such as TextCNN, TextRNN, etc.; if the task of industry data is text generation, then some fine-tuning algorithms suitable for text generation can be selected, such as Seq2Seq, GPT-2, etc.

[0181] According to the parameter scale of the model, select the parameter-efficient fine-tuning algorithm or the parameter-inefficient fine-tuning algorithm. The parameter scale of the model is another important factor affecting the selection of the fine-tuning algorithm, because models with different parameter scales require different computing capabilities and storage spaces. For example, if the model is a model with a large parameter scale, then some parameter-efficient fine-tuning algorithms can be selected, such as ReZero, AdaBelief, etc.; if the model is a model with a small parameter scale, then some parameter-inefficient fine-tuning algorithms can be selected, such as Adam, SGD, etc.

[0182] The first-level industry large model obtained after training adjustment already has some industry knowledge understanding capabilities. In the public cloud environment, the model can already conduct preliminary tests and evaluations on the model effects and performance of the first-level industry large model for specific industry scenario tasks, and perform iterative training of the model.

[0183] In the public cloud environment, the industry model can use various testing and evaluation tools provided by the public cloud, such as cloud functions, cloud logs, etc., to conduct preliminary tests and evaluations on the model effects and performance of the first-level industry large model. The testing and evaluation methods can select appropriate evaluation metrics and data sets according to different tasks and fields, such as accuracy, recall rate, F1 value, human evaluation, etc. The results of the testing and evaluation can provide some feedback and suggestions for the object to help the object perform iterative training on the model, that is, adjust and optimize the parameters, structure, algorithm, etc. of the model according to the results of the testing and evaluation, so as to improve the performance and adaptability of the model.

[0184] S1350. Perform model compression processing on the first-level industry model to obtain the second-level industry model.

[0185] The first-level industry large model still retains the same scale of parameter quantity as the basic large model, which poses a great resource challenge to the private cloud environment of the object. In order to reduce the resource burden of the object's private cloud environment, the first-level industry large model can be compressed in the public cloud environment (such as quantization, pruning, and knowledge distillation) to reduce the model size and computing requirements, and output a model with a smaller parameter scale.

[0186] After compression processing, it is also necessary to evaluate the compressed secondary industry large model to ensure that the model is within a certain acceptable range of accuracy and performance loss.

[0187] S1360: Keep non-disclosable data in the private cloud environment, transfer the secondary industry model to the private cloud, and deploy the training environment of the secondary industry model in the private cloud environment.

[0188] In the private cloud, the object can store its non-disclosable training data in the storage service of the private cloud, and then transfer the secondary industry large model from the public cloud to the private cloud through a dedicated line or in an encrypted manner, or export it from the public cloud and then upload it offline to the private cloud.

[0189] Deploy the training environment of the secondary industry large model on the server of the private cloud. One can choose to purchase the large model training tools / products of the manufacturer (without the object implementing the development and deployment of the training tools, all implemented by the manufacturer); or the object can develop and deploy its own training tools in the private cloud.

[0190] S1370: Utilize the private cloud training environment to fine-tune the secondary industry model based on the data in the private cloud to obtain the tertiary industry model.

[0191] To further improve the performance and adaptability of the industry large model and enable it to better meet the business scenarios and requirements of the object, it is necessary to perform another fine-tuning on the secondary industry model based on non-disclosable data in the private cloud training environment to further adjust and optimize the parameters of the model, so that the model can better adapt to the specific data and tasks of the object.

[0192] In the process of the second fine-tuning, attention should be focused on the business effect, and iterative repeated training can be carried out according to the actual situation until the model accuracy and performance indicators that meet the business requirements are achieved. For the evaluation of the business effect, appropriate evaluation indicators and data sets can be selected according to different tasks and fields, such as accuracy rate, recall rate, F1 value, manual evaluation, etc.

[0193] The tertiary industry large model obtained after the second fine-tuning is the final industry large model (the aforementioned target industry model), which has the highest industry knowledge understanding ability. In the private cloud environment, the object can conduct final testing and evaluation on the model effect and performance of the tertiary industry large model for specific industry scenario tasks, and deploy and apply the model.

[0194] In one example, before deployment and inference, the tertiary industry large model can also perform a model compression operation to further reduce the scale of model parameters, so as to facilitate embedding the model into more scenarios and devices to serve more objects that call the model service.

[0195] The training of industry large models must be based on the cornerstone of industry data. Due to the need for data security protection by the object, it is often impossible to provide all industry data to the public cloud environment. This requires the object to have more large model training resources (such as high-specification GPUs, storage, etc.). The requirements for training resources directly increase the application threshold of large models in industry institutions, resulting in many objects being unable to use them.

[0196] In the embodiments of the present application, the elasticity of computing / storage and other resources on the public cloud is fully utilized to solve the problem that the object cannot build its own training environment. The object can rent the training environment of the public cloud platform and pay as needed, which not only saves the cost of building the training environment but also allows the object to focus on the training and application process of large models that are more closely related to the business level.

[0197] At the same time, the embodiments of the present application can enable the object not to provide a lot of sensitive data (data that cannot be made public after evaluation) externally. Instead, in the private cloud environment built by the object, while ensuring security, through technologies such as model compression, the scale of model parameters trained on the public cloud is further reduced, enabling a "smaller" parameter-scale model to be trained in the private cloud environment with lower hardware resources, further reducing the construction cost of the object's private cloud environment, and allowing the object to perform model fine-tuning in the private cloud environment at a lower cost and higher efficiency to meet the business needs.

[0198] The device embodiments of the present application are introduced, which can be used to execute the model training method based on the cloud environment in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the model training method based on the cloud environment above.

[0199] The embodiments of the present application provide a model training device based on the cloud environment, as Figure 14 shown, the device includes:

[0200] An acquisition module 1410, configured to acquire industry data of a target industry in a private cloud environment;

[0201] A division module 1420, configured to divide the industry data into public data and non-public data according to the security level of the industry data, and transmit the public data from the private cloud environment to the public cloud environment;

[0202] An adjustment module 1430, configured to acquire a target basic model, and in the public cloud environment, adjust the target basic model according to the public data to obtain a first-level industry model;

[0203] The adjustment module 1430 is further configured to adjust the primary industry model according to the non-public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process tasks in the target industry.

[0204] In an embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to perform model compression processing on the primary industry model in the public cloud environment to obtain a secondary industry model; transfer the secondary industry model from the public cloud environment to the private cloud environment, and build a training environment for the secondary industry model in the private cloud environment; in the training environment, adjust the secondary industry model according to the non-public data to obtain a target industry model.

[0205] In an embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to perform encryption processing on the secondary industry model to obtain an encrypted model, and transmit the encrypted model to the private cloud environment through a dedicated line; decrypt the encrypted model in the private cloud environment to obtain the secondary industry model; after verifying the transmitted secondary industry model, divide an isolation space for the secondary industry model in the private cloud environment, and build a training environment in the isolation space.

[0206] In an embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to remove connections or parameters with contribution degrees less than a preset contribution degree threshold in the primary industry model in the public cloud environment to obtain a primary industry model; perform quantization processing on the primary industry model to compress the model parameters of the primary industry model from floating-point representation to low-order integers; use the computing resources of the public cloud environment to evaluate the quantized primary industry model, and determine whether to adjust the compressed model parameters according to the evaluation result to obtain the secondary industry model.

[0207] In an embodiment of the present application, based on the foregoing solution, the number of private cloud environments includes multiple, and the non-public data in each private cloud environment is different; the adjustment module is further configured to respectively adjust the secondary industry model according to the local non-public data in each private cloud environment in the training environments deployed in each private cloud environment; aggregate the model update parameters of each private cloud environment for the secondary industry model to obtain target model update parameters; adjust the secondary industry model according to the target model update parameters to obtain a target industry model.

[0208] In one embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to perform data anonymization processing on specified words in the non-public data local to each private cloud environment, and add preset noise to the non-public data after the data anonymization processing to obtain target non-public data; in the training environment deployed in each private cloud environment, adjust the secondary industry model according to the target non-public data.

[0209] In one embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to deploy a training environment for the target basic model in the public cloud environment, and in the training environment, adjust the target basic model according to the public data to obtain an initial model; test the initial model according to multiple business scenario tasks to determine the coverage of the initial model for the business scenario; perform iterative adjustment on the initial model according to the coverage to obtain the primary industry model.

[0210] In one embodiment of the present application, based on the foregoing solution, the adjustment module is further configured to select a corresponding adjustment strategy according to the annotation situation of the public data and the task type corresponding to the target industry; in the training environment, adjust the target basic model according to the adjustment strategy and the public data to obtain an initial model.

[0211] In one embodiment of the present application, based on the foregoing solution, the division module is further configured to determine the security level of each industry data according to the source and content of the industry data; use the industry data with the security level greater than the preset security level threshold as non-public data; for the industry data with the security level less than the preset security level threshold as pending industry data, obtain the sensitivity of the pending industry data, and determine the public data in the pending industry data according to the sensitivity.

[0212] In one embodiment of the present application, based on the foregoing solution, the division module is further configured to determine the sensitivity of the pending industry data according to the number of target words included in the pending industry data; divide the pending industry data into data to be desensitized and regular data according to the sensitivity, and use the data obtained after desensitizing the data to be desensitized and the regular data as the public data.

[0213] In one embodiment of the present application, based on the foregoing solution, the acquisition module is further configured to acquire the data type of the industry data and the resource conditions for model deployment; determine a model structure corresponding to the data type according to the data type, and determine a parameter version corresponding to the resource conditions according to the resource conditions; acquire the task type to which the industry data belongs, and determine a model output corresponding to the task type according to the task type; and select a target base model from multiple base models according to the model structure, parameter version, and model output.

[0214] In one embodiment of the present application, based on the foregoing solution, the device further includes a deployment module, configured to acquire the deployment scenario of the target industry model, and determine whether to perform model compression processing on the target industry model according to the deployment scenario; if it is determined to perform model compression processing on the target industry model, deploy the model after the model compression processing from the private cloud environment to the deployment scenario.

[0215] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be elaborated here.

[0216] An embodiment of the present application further provides an electronic device, including one or more processors and a storage device. The storage device is configured to store one or more computer programs. When the one or more computer programs are executed by the one or more processors, the electronic device implements the model training method based on the cloud environment as described above.

[0217] Figure 15 The structural schematic diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown.

[0218] It should be noted that Figure 15 The computer system 1500 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0219] Such as Figure 15As shown, computer system 1500 includes a Central Processing Unit (CPU) 1501, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 1502 or the program loaded from the storage section 1508 into the Random Access Memory (RAM) 1503, such as executing the methods in the above embodiments. In the RAM 1503, various programs and data required for system operation are also stored. The CPU 1501, ROM 1502, and RAM 1503 are connected to each other via a bus 1504. An Input / Output (I / O) interface 1505 is also connected to the bus 1504.

[0220] In some embodiments, the following components are connected to the I / O interface 1505: an input section 1506 including a keyboard, a mouse, etc.; an output section 1507 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 1509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the I / O interface 1505 as needed. A removable medium 1511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1510 as needed so that a computer program read from it can be installed into the storage section 1508 as needed.

[0221] Specifically, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer programs. For example, the embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1509, and / or installed from the removable medium 1511. When the computer program is executed by the processor (CPU) 1501, various functions defined in the system of the present application are executed.

[0222] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0223] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and a computer program.

[0224] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware, and the described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.

[0225] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist alone without being assembled into the electronic device.

[0226] Another aspect of this application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the electronic device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the electronic device executes the method described above in each of the above embodiments.

[0227] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0228] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of this application. This application is intended to cover any variations, uses, or adaptations of this application, which follow the general principles of this application and include known common knowledge or conventional technical means in the technical field not disclosed in this application.

[0229] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation of this application. Those of ordinary skill in the art can easily make corresponding modifications or alterations according to the main concept and spirit of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.

Claims

1. A model training method based on a cloud environment, characterized in that, it includes: Obtain industry data of the target industry in the private cloud environment; Divide the industry data into public data and non - public data according to the security level of the industry data, and transmit the public data from the private cloud environment to the public cloud environment; Obtain a target basic model, and in the public cloud environment, adjust the target basic model according to the public data to obtain a primary industry model; Adjust the primary industry model according to the non - public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process the tasks of the target industry.

2. The method according to claim 1, characterized in that, Adjusting the primary industry model according to the non - public data in the private cloud environment to obtain a target industry model includes: In the public cloud environment, perform model compression processing on the primary industry model to obtain a secondary industry model; Transmit the secondary industry model from the public cloud environment to the private cloud environment, and build a training environment for the secondary industry model in the private cloud environment; In the training environment, adjust the secondary industry model according to the non - public data to obtain a target industry model.

3. The method according to claim 2, characterized in that, The step of transmitting the secondary industry model from the public cloud environment to the private cloud environment and building a training environment for the secondary industry model in the private cloud environment includes: Perform encryption processing on the secondary industry model to obtain an encrypted model, and transmit the encrypted model to the private cloud environment through a dedicated line; Decrypt the encrypted model in the private cloud environment to obtain the secondary industry model; After the transmitted secondary industry model passes the verification, divide an isolated space for the secondary industry model in the private cloud environment, and build a training environment in the isolated space.

4. The method according to claim 2, characterized in that, The step of performing model compression processing on the primary industry model in the public cloud environment to obtain a secondary industry model includes: In the public cloud environment, remove connections or parameters with a contribution degree less than a preset contribution degree threshold in the primary industry model to obtain a primary industry model; Perform quantization processing on the primary industry model to compress the model parameters of the primary industry model from floating - point representation to low - order integers; Use the computing resources of the public cloud environment to evaluate the quantized primary industry model, and determine whether to adjust the compressed model parameters according to the evaluation results to obtain the secondary industry model.

5. The method according to claim 2, characterized in that, The number of private cloud environments includes multiple, and the non - public data in each private cloud environment is different; the step of adjusting the secondary industry model according to the non - public data in the training environment to obtain a target industry model includes: In the training environments deployed in each private cloud environment, the secondary industry model is adjusted respectively according to the local non-public data in each private cloud environment; Aggregate the model update parameters of the secondary industry model for each private cloud environment to obtain target model update parameters; Adjust the secondary industry model according to the target model update parameters to obtain a target industry model.

6. The method according to claim 5, wherein, the adjusting the secondary industry model respectively according to the local non-public data in each private cloud environment in the training environments deployed in each private cloud environment includes: Performing data anonymization processing on specified words in the non-public data local to each private cloud environment, and adding preset noise to the non-public data after data anonymization processing to obtain target non-public data; In the training environments deployed in each private cloud environment, adjusting the secondary industry model according to the target non-public data.

7. The method according to claim 1, wherein, the adjusting the target basic model according to the public data in the public cloud environment to obtain a primary industry model includes: Deploying a training environment for the target basic model in the public cloud environment, and in the training environment, adjusting the target basic model according to the public data to obtain an initial model; Testing the initial model according to multiple business scenario tasks to determine the coverage of the initial model for business scenarios; Iteratively adjusting the initial model according to the coverage to obtain the primary industry model.

8. The method according to claim 7, wherein, the adjusting the target basic model according to the public data in the training environment to obtain an initial model includes: Selecting a corresponding adjustment strategy according to the annotation situation of the public data and the task type corresponding to the target industry; In the training environment, adjusting the target basic model according to the adjustment strategy and the public data to obtain an initial model.

9. The method according to claim 1, wherein, the dividing the industry data into public data and non-public data according to the security level of the industry data includes: Determining the security level of each industry data according to the source and content of the industry data; Regarding the industry data with a security level higher than a preset security level threshold as non-public data; Regarding the industry data with a security level lower than the preset security level threshold as pending industry data, obtaining the sensitivity of the pending industry data, and determining the public data in the pending industry data according to the sensitivity.

10. The method according to claim 9, wherein, the determining the public data according to the data type of the pending industry data includes: Determining the sensitivity of the pending industry data according to the number of target words included in the pending industry data; Divide the to-be-determined industry data into data to be desensitized and regular data according to the sensitivity level, and use the data obtained after desensitizing the data to be desensitized and the regular data as the public data.

11. The method according to claim 1, wherein, the obtaining of the target base model includes: obtaining the data type of the industry data and the resource conditions for model deployment; determining a model structure corresponding to the data type according to the data type, and determining a parameter version corresponding to the resource conditions according to the resource conditions; obtaining the task type to which the industry data belongs, and determining a model output corresponding to the task type according to the task type; selecting a target base model from multiple base models according to the model structure, parameter version and model output.

12. The method according to any one of claims 1 to 11, wherein, after adjusting the first-level industry model according to the non-public data in the private cloud environment to obtain a target industry model, the method further includes: obtaining the deployment scenario of the target industry model, and determining whether to perform model compression processing on the target industry model according to the deployment scenario; if it is determined to perform model compression processing on the target industry model, deploying the model after the model compression processing from the private cloud environment to the deployment scenario.

13. A model training device based on a cloud environment, wherein, it includes: an obtaining module, configured to obtain industry data of a target industry in a private cloud environment; a dividing module, configured to divide the industry data into public data and non-public data according to the security level of the industry data, and transmit the public data from the private cloud environment to a public cloud environment; an adjusting module, configured to obtain a target base model, and in the public cloud environment, adjust the target base model according to the public data to obtain a first-level industry model; the adjusting module is further configured to adjust the first-level industry model according to the non-public data in the private cloud environment to obtain a target industry model, and the target industry model is used to process the tasks of the target industry.

14. An electronic device, wherein, it includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enable the electronic device to execute the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, wherein, a computer program is stored thereon, and when the computer program is executed by a processor of an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 12.