Model development data anti-leakage method and device oriented to Plevatory intelligent computing center cloud platform

By allocating encrypted data to the CPU in the intelligent computing center cloud platform and decrypting and transmitting it to the GPU for model development, the problems of low computing resource utilization and high leasing cost are solved, and efficient model development and universal computing power applications are realized.

CN120387181AActive Publication Date: 2025-07-29DATACANVAS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510889602.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-29
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

In the prior art, the intelligent computing center has low computing resource utilization rate, low model development efficiency and high computing power leasing costs.

Method used

In the intelligent computing center cloud platform, after receiving the encrypted data sent by the client, it is allocated to the CPU of multiple nodes for decryption, and the decrypted data is transmitted to the GPU for model development, avoiding the occupation of GPU resources, and directly decrypting the data by the CPU to improve computing resource utilization and model development efficiency.

Benefits of technology

It improves the utilization rate of computing power resources, reduces the preparation time and leasing cost of model development, and realizes the widespread application of universal computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387181A_ABST
    Figure CN120387181A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for preventing leakage of model development data for a common computing power intelligent computing center cloud platform, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the following steps: S1, receiving encrypted data sent by a client; s2, distributing the encrypted data to a first node in a plurality of nodes, wherein the plurality of nodes are nodes deployed by an intelligent computing center cloud platform; s3, decrypting the encrypted data based on the first CPU of the first node to obtain decrypted data; s4, transmitting the decrypted data to a first GPU (Graphic Processing Unit) of the first node; and S5, carrying out model development based on the decrypted data through the first GPU. According to the method, the computing power resource utilization rate can be improved, and wide application of the general computing power is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for preventing model development data leakage in a cloud platform of an intelligent computing center for inclusive computing power. Background Technique

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios for artificial intelligence deep learning model development, model training, and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process parameters. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to process parameter data and achieve the output of the target result. It is a new type of productive force that integrates parameter computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] In the process of model development in an intelligent computing center, it is necessary to first upload the data required for model development to the intelligent computing center, and then use the computing power resources provided by the GPUs in the intelligent computing center to develop the model based on the data. In the prior art, in order to prevent data leakage, it is necessary to encrypt the data before uploading it; after the intelligent computing center receives the encrypted data, it stores the encrypted data and decrypts the data through the GPUs to facilitate subsequent model development. However, in the prior art, decrypting the data through the GPUs will consume a large amount of computing power resources, resulting in less computing power resources available for model development, making the utilization rate of computing power resources very low. Or, decrypting and storing the data through the CPUs and reading the decrypted data from storage during the model development stage takes a long time, resulting in very low model development efficiency. At the same time, since data decryption consumes computing power resources, this part of the computing power resources also needs to be borne by users when renting computing power services, resulting in a very high rental cost, making it difficult to achieve the widespread application of inclusive computing power.

[0008] It can be seen that in the prior art, there are problems such as very low utilization rate of computing power resources in the intelligent computing center or very low model development efficiency, and very high computing power rental costs. Summary of the Invention

[0009] Embodiments of the present invention provide a method and device for preventing data leakage in model development for an intelligent computing center cloud platform for inclusive computing power, so as to solve the problems in the prior art such as very low utilization rate of computing power resources in the intelligent computing center or very low model development efficiency, and very high computing power rental costs.

[0010] To solve the above problems, the present invention is implemented as follows: In a first aspect, an embodiment of the present invention provides a method for preventing data leakage in model development for an intelligent computing center cloud platform for inclusive computing power, including: Step S1, receiving encrypted data sent by a client; Step S2, allocating the encrypted data to a first node among multiple nodes, where the multiple nodes are nodes deployed by the intelligent computing center cloud platform; Step S3, decrypting the encrypted data based on a first CPU of the first node to obtain decrypted data; Step S4, transmitting the decrypted data to a first GPU of the first node; Step S5, developing a model based on the decrypted data through the first GPU.

[0011] In one embodiment, step S2 includes: Step S21, obtaining a first parameter of each node among the multiple nodes; Step S22: Determine the first node from the multiple nodes based on the first parameter and a preset rule; Step S23: Allocate the encrypted data to the first node; Wherein, the preset rule includes one of the following: When the first parameter is the number of times of receiving encrypted data, the first node is the node with the least number of times of receiving encrypted data among the multiple nodes; When the first parameter is a load parameter, the first node is the node with the smallest load parameter among the multiple nodes.

[0012] In one embodiment, the step S4 includes: Step S41: Obtain the load parameter of the first GPU; Step S42: When the load parameter is less than or equal to a first set load threshold, transmit the decrypted data to the first GPU.

[0013] In one embodiment, the method further includes: Step S43: When the load parameter is greater than the first set load threshold, transmit the decrypted data to the second GPU of a second node, where the second node is a node other than the first node among the multiple nodes.

[0014] In one embodiment, communication transmission of the second CPU included in the second node is abnormal, or the load parameter of the second GPU is less than or equal to a second set load threshold, and the second set load threshold is less than the first set load threshold.

[0015] In one embodiment, the step S4 includes: Step S41': When the data volume of the decrypted data accumulated in the memory of the first CPU reaches a preset data volume threshold, transmit the decrypted data accumulated in the memory of the first CPU to the first GPU.

[0016] In one embodiment, the preset data volume threshold is obtained by the following method: Obtain at least one of the batch size parameter, the number of model parameters, and the training parameters corresponding to the model to be developed; Based on at least one of the batch size parameter, the number of model parameters, and the training parameters, calculate the preset data volume threshold.

[0017] In a second aspect, an embodiment of the present invention further provides a model development data anti-leakage device for a cloud platform of an inclusive computing power intelligent computing center, including: A receiving module, configured to receive encrypted data sent by a client; A distribution module for distributing the encrypted data to a first node among a plurality of nodes, where the plurality of nodes are nodes deployed by an intelligent computing center cloud platform; A decryption module for decrypting the encrypted data based on a first CPU of the first node to obtain decrypted data; A transmission module for transmitting the decrypted data to a first GPU of the first node; A development module for performing model development based on the decrypted data through the first GPU.

[0018] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for preventing data leakage in model development for an intelligent computing center cloud platform for inclusive computing power as described in the first aspect above.

[0019] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for preventing data leakage in model development for an intelligent computing center cloud platform for inclusive computing power as described in the first aspect above.

[0020] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps in the method for preventing data leakage in model development for an intelligent computing center cloud platform for inclusive computing power as described in the first aspect above.

[0021] In the present invention, a method for preventing data leakage in model development for an intelligent computing center cloud platform oriented to inclusive computing power is provided, including: Step S1, receiving encrypted data sent by a client; Step S2, distributing the encrypted data to a first node among multiple nodes, where the multiple nodes are nodes deployed by the intelligent computing center cloud platform; Step S3, decrypting the encrypted data based on a first CPU of the first node to obtain decrypted data; Step S4, transmitting the decrypted data to a first GPU of the first node; Step S5, performing model development based on the decrypted data through the first GPU. In this way, by decrypting the encrypted data through the first GPU, it does not occupy the GPU resources of the intelligent computing center cloud platform, enabling the first GPU to provide more computing power resources for model development, greatly improving the utilization rate of computing power resources; during the model development process, there is no need to read decrypted data from storage, but directly obtain the decrypted data decrypted by the first CPU through the first GPU, thereby effectively reducing the preparation time and greatly improving the model development efficiency. At the same time, users do not need to bear the computing power resources consumed when decrypting encrypted data, thereby reducing the rental cost of computing power resources and realizing the wide application of inclusive computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of a method for preventing data leakage in model development for an intelligent computing center cloud platform oriented to inclusive computing power provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the data transmission process provided by an embodiment of the present invention; Figure 3 It is a structural diagram of a device for preventing data leakage in model development for an intelligent computing center cloud platform oriented to inclusive computing power provided by an embodiment of the present invention; Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0025] The "computing power" as described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0026] The "computational power" (Computational Power, CP) as described in the present invention refers to: the ability of a data center server to process data and achieve result output, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。

[0027] The "carrying capacity" (Network Power, NP) as described in the present invention refers to: the manifestation of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission within and between data centers, and a comprehensive indicator for measuring network transmission scheduling ability.

[0028] The "storage power" (Storage Power, SP) as described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0029] The "computing power infrastructure" as described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.

[0030] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0031] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0032] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0033] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.

[0034] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0035] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0036] The "intelligent computing center cloud platform" described in the present invention refers to: a cloud computing platform that comprehensively serves based on the hardware resources and software resources of the intelligent computing center.

[0037] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0038] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0039] The "Computing Power Center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0040] The "Supercomputing Center" described in the present invention refers to: that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0041] The "Computing Power Resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0042] The "Model" described in the present invention includes but is not limited to "Large Language Model" and "Multimodal Large Model".

[0043] The "Large Language Model" described in the present invention refers to the large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0044] The "Multimodal Large Model" described in the present invention (Multimodal Large Models) refers to: a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0045] The "Computing Power Operation Task" described in the present invention refers to: a specific workload or job executed on computing power resources and requiring a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.

[0046] The "Inclusive Computing Power" described in the present invention refers to: based on the requirements of equal opportunity and the principle of commercial sustainability, providing appropriate and effective computing power services to all social strata and groups with computing power service needs at an affordable cost.

[0047] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for preventing data leakage in model development for an inclusive computing power intelligent computing center cloud platform provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps: Step S1: Receive encrypted data sent by the client.

[0048] The above client is the client for the user to upload data. The client can be a terminal or a server. After the user collects or processes the data for model development and encrypts it, the encrypted data is uploaded to the intelligent computing center cloud platform through the client, and the intelligent computing center cloud platform decrypts the encrypted data, and then conducts model development based on the decrypted data.

[0049] The above encrypted data is the data encrypted by the client. It should be noted that the encrypted data can be obtained by the client encrypting the data in real time before transmission, or can be encrypted for all data in advance and then uploaded to the intelligent computing center cloud platform through the client.

[0050] Among them, the encrypted data is the data in the smallest unit, and the encrypted data cannot be further divided. If the intelligent computing center cloud platform receives multiple encrypted data, each encrypted data needs to be allocated in turn, and the nodes to which different encrypted data are allocated can be the same or different.

[0051] Step S2: Allocate the encrypted data to the first node among multiple nodes, and the multiple nodes are the nodes deployed by the intelligent computing center cloud platform.

[0052] The above multiple nodes are the nodes deployed by the intelligent computing center cloud platform. Each node is created based on the computing power resources of the intelligent computing center cloud platform and can provide computing power resources to execute computing power operation tasks to implement functions such as model development. Among them, each node is configured with CPU resources and GPU resources, specifically as Figure 2 shown, so that each node can realize data decryption and model development.

[0053] The above allocation of the encrypted data to the first node among multiple nodes can be that the gateway of the intelligent computing center cloud platform receives the encrypted data sent by the client, and then the gateway sends the encrypted data to different nodes (including the first node).

[0054] Step S3: Decrypt the encrypted data based on the first CPU of the first node to obtain decrypted data.

[0055] The above-mentioned first CPU can be the real CPU of the intelligent computing center cloud platform, or a virtual CPU created based on the CPU resources of the intelligent computing center cloud platform. It should be noted that the first CPU can provide CPU resources, so as to realize decrypting the encrypted data based on the first CPU of the first node.

[0056] The above-mentioned decrypting the encrypted data based on the first CPU of the first node is specifically that the gateway directly distributes the encrypted data to the memory of the first CPU, and then the first CPU decrypts the encrypted data in the memory. Among them, the first CPU decrypts the encrypted data in real time, that is, after the encrypted data is distributed to the memory of the first CPU, the first CPU starts to decrypt the distributed encrypted data, so as to avoid the encrypted data occupying the memory of the first CPU due to long-term non-processing of the encrypted data.

[0057] In the present invention, decrypting the encrypted data by the first CPU does not occupy the GPU resources of the intelligent computing center cloud platform, so that the intelligent computing center cloud platform can provide more GPU resources for model development, greatly improving the utilization rate of computing power resources.

[0058] Step S4: Transmit the decrypted data to the first GPU of the first node.

[0059] The above-mentioned first GPU can be the real GPU of the intelligent computing center cloud platform, or a virtual GPU created based on the GPU resources of the intelligent computing center cloud platform. It should be noted that the first GPU can provide GPU resources, so as to realize model development based on the decrypted data through the first GPU.

[0060] The above-mentioned transmitting the decrypted data to the first GPU of the first node is specifically that after the first CPU decrypts to obtain the decrypted data, the decrypted data in the memory of the first CPU is transmitted to the memory of the first GPU, and then the first GPU performs model development through the decrypted data in the memory.

[0061] In some embodiments, transmitting the decrypted data from the memory of the first CPU to the memory of the first GPU is achieved through an extended (Peripheral Component Interconnect Express, PCIE) interface for fast transmission, so that the first GPU can perform model development according to the received decrypted data.

[0062] In the present invention, the first CPU directly transmits the decrypted data to the first GPU without storing it in the storage space of the intelligent computing center cloud platform, thereby avoiding the first GPU obtaining the decrypted data from the storage space, reducing the time loss in the data preparation process, and greatly improving the efficiency of model development.

[0063] Step S5: Perform model development based on the decrypted data by means of the first GPU.

[0064] In the present invention, a method for preventing leakage of model development data for an intelligent computing center cloud platform for inclusive computing power is provided, including: Step S1: Receive encrypted data sent by a client; Step S2: Allocate the encrypted data to a first node among multiple nodes, where the multiple nodes are nodes deployed by the intelligent computing center cloud platform; Step S3: Decrypt the encrypted data by means of a first CPU of the first node to obtain decrypted data; Step S4: Transmit the decrypted data to a first GPU of the first node; Step S5: Perform model development based on the decrypted data by means of the first GPU. In this way, the first GPU decrypts the encrypted data without occupying the GPU resources of the intelligent computing center cloud platform, enabling the first GPU to provide more computing power resources for model development, thus greatly improving the utilization rate of computing power resources; during the model development process, there is no need to read the decrypted data from storage, but directly obtain the decrypted data decrypted by the first CPU through the first GPU, thereby effectively reducing the preparation time and greatly improving the model development efficiency. At the same time, the user does not need to bear the computing power resources consumed when decrypting the encrypted data, thus reducing the rental cost of computing power resources and realizing the wide application of inclusive computing power.

[0065] In one embodiment, Step S2 includes: Step S21: Obtain first parameters of each of the multiple nodes; Step S22: Determine the first node from the multiple nodes based on the first parameters and a preset rule; Step S23: Allocate the encrypted data to the first node; Wherein, the preset rule includes one of the following: When the first parameter is the number of times of receiving encrypted data, the first node is the node among the multiple nodes with the least number of times of receiving encrypted data; When the first parameter is a load parameter, the first node is the node among the multiple nodes with the smallest load parameter.

[0066] It should be noted that when allocating encrypted data to multiple nodes, there may be a situation where the load of different nodes varies significantly due to uneven allocation, resulting in a significant difference in the model development efficiency of different nodes, which may lead to a situation where the computing power resources of some nodes are not fully utilized and a problem of low utilization rate of computing power resources of some nodes.

[0067] To solve the above problems, a preset rule is set in the present invention. The first node to which the encrypted data needs to be allocated is determined from multiple nodes through the preset rule, and then the encrypted data is allocated. In this way, every time the encrypted data is received, a confirmation is made through the preset rule before allocation, so that multiple nodes can receive the encrypted data evenly, and the efficiency of model development for each node is close, avoiding the situation that the computing power resources of some nodes are not fully utilized, and greatly improving the utilization rate of computing power resources.

[0068] Among them, the preset rule can be set according to the quantity of the encrypted data or determined according to the load of different nodes.

[0069] In some embodiments, in the process of allocating the encrypted data to multiple nodes, the encrypted data can be sequentially allocated to different nodes, so that the number of times each node receives the encrypted data is the same or close, so that the load of each node is close, avoiding the situation that the computing power resources of some nodes are not fully utilized, and greatly improving the utilization rate of computing power resources.

[0070] In other embodiments, in the process of allocating the encrypted data to multiple nodes, first obtain the load parameters of each node among the multiple nodes, directly determine the current load conditions of different nodes through the load parameters, and then allocate the encrypted data to the node with the lowest load parameter, so that the load of each node is close, avoiding the situation that some nodes with low load are not allocated encrypted data, achieving the load balance of each node, and further improving the utilization rate of computing power resources.

[0071] In other embodiments, the first node can also be determined by the difference between the load parameters of different nodes. For example, when model development is carried out by 2 nodes, when the difference between the load parameters of the 2 nodes is less than or equal to the set threshold, the allocation is carried out in the way of the number of times of using the encrypted data; when the difference between the load parameters of the 2 nodes is greater than the set threshold, the encrypted data is allocated to the node with lower load parameter.

[0072] In one embodiment, step S4 includes: Step S41, obtain the load parameter of the first GPU; Step S42, when the load parameter is less than or equal to the first set load threshold, transmit the decrypted data to the first GPU.

[0073] It should be noted that the decrypted data can also be transmitted between the CPUs and GPUs of different nodes through the PICE interface. In this way, if the GPU of a certain node has a large load while the GPUs of other nodes have a small load, the decrypted data in the CPU memory of this node can be transmitted to the GPU memory of other nodes to achieve load balancing of the GPUs of different nodes and improve the utilization rate of computing power resources.

[0074] Specifically, when the load parameter is less than or equal to the first set load threshold, the load of the GPU of the first node is at a reasonable level and there is no situation of large load. At this time, the decrypted data is transmitted to the first GPU, and the model development is directly carried out by the GPU of the first node.

[0075] In one embodiment, the method further includes: Step S43, when the load parameter is greater than the first set load threshold, transmit the decrypted data to the second GPU of the second node, where the second node is a node other than the first node among the multiple nodes.

[0076] In the present invention, when the load parameter is greater than the first set load threshold, the first GPU of the first node has a large load. At this time, the decrypted data in the first CPU memory needs to be sent to the second GPU memory to improve the utilization rate of the computing power resources of the second GPU.

[0077] Among them, the second node can be the node with the smallest load parameter corresponding to the GPU. By transmitting the decrypted data to the GPU memory of this node, the utilization rate of the computing power resources of this node can be improved.

[0078] It should be noted that if the load parameter corresponding to the GPU of the second node is greater than or equal to the load parameter corresponding to the first GPU of the first node, or the load parameters corresponding to the GPUs of all nodes are greater than the first set load threshold, it means that the loads of all nodes are at a relatively high level at this time. At this time, it may not be necessary to transmit the decrypted data across nodes, and the decrypted data in the first CPU memory is directly transmitted to the first GPU memory, and the first GPU develops the model based on the decrypted data.

[0079] In one embodiment, the second node has an abnormal communication transmission of the second CPU, or the load parameter of the second GPU is less than or equal to the second set load threshold, and the second set load threshold is less than the first set load threshold.

[0080] In the present invention, the decrypted data in the CPU memory can be transmitted across nodes to the GPU memory, based on which the problem of abnormal communication transmission can be solved. It should be noted that during the distribution process of the encrypted data, there may be abnormal communication transmission between the CPUs of some nodes and the gateway. At this time, these nodes cannot obtain the encrypted data from the gateway, and the gateway can only distribute the encrypted data to the CPU memories of other nodes. This situation will result in the inability of the GPUs of the nodes that cannot obtain the encrypted data to perform model development, while the load of other nodes is relatively high. Therefore, in the present invention, the nodes with abnormal communication transmission are regarded as the second nodes, and the decrypted data in the first CPU memory is transmitted to the second GPU memory, so that the second GPU can perform model development, improving the utilization rate of the computing power resources of the second GPU.

[0081] Furthermore, a GPU with a relatively low load can also be used as the second GPU to improve the utilization rate of the computing power resources of the second GPU. Among them, whether the load of the GPU is relatively low is determined by the second set load threshold. Specifically, when the load parameter of the GPU is less than or equal to the second set load threshold, the load of the GPU is relatively low and can be used as the second GPU; when the load parameter of the GPU is greater than the second set load threshold, the load of the GPU is normal, and at this time, this GPU is not used as the second GPU.

[0082] Among them, the second set load threshold is less than the first set load threshold. The first set load threshold is used to determine whether the load of the CPU is too high, and the second set load threshold is used to determine whether the load of the CPU is too low.

[0083] In one embodiment, step S4 includes: Step S41’: When the data volume of the decrypted data accumulated in the memory of the first CPU reaches the preset data volume threshold, the decrypted data accumulated in the memory of the first CPU is transmitted to the first GPU.

[0084] It should be noted that during the process of transmitting the encrypted data to the CPU memory, the transmission is realized through the Internet, and its transmission rate is relatively low; while during the process of transmitting the decrypted data from the CPU memory to the GPU memory, the transmission is realized through the PCIE interface, and its transmission rate is relatively high. This results in that during the entire data transmission process, the GPU needs to wait for the CPU to transmit the decrypted data for model development. At this time, the GPU occupies the computing power resources but does not perform model development, resulting in waste of the computing power resources.

[0085] To solve the above problems, in the present invention, after the encrypted data is transmitted to the first CPU, the decrypted data is first saved in the CPU memory until the amount of the decrypted data accumulated in the memory reaches a preset data volume threshold, and then the accumulated decrypted data is transmitted to the first GPU. At this time, the first GPU can receive more decrypted data and can directly start model development without waiting, thereby maximizing the utilization rate of computing power resources.

[0086] In one embodiment, the preset data volume threshold is obtained in the following manner: Obtain at least one of the batch size parameter, the number of model parameters, and the training parameters corresponding to the model to be developed; Based on at least one of the batch size parameter, the number of model parameters, and the training parameters, calculate the preset data volume threshold.

[0087] In the present invention, the preset data volume threshold is positively correlated with the batch size parameter, the number of model parameters, and the training parameters, and the preset data volume threshold can be calculated based on at least one of the batch size parameter, the number of model parameters, and the training parameters through a formula or a model. For example, the batch size parameter, the number of model parameters, and the training parameters are weighted to calculate the preset data volume threshold.

[0088] In addition, the preset data volume threshold is also limited by the GPU memory, that is, the preset data volume threshold cannot exceed the upper limit value of the GPU memory to indicate a program error caused by data overflow, so that the model development process can proceed stably.

[0089] Please refer to Figure 3 , Figure 3 which is a structural diagram of a model development data anti-leakage device for a cloud platform of an inclusive computing power intelligent computing center provided by an embodiment of the present invention. As Figure 3 shown, the model development data anti-leakage device 300 for the cloud platform of the inclusive computing power intelligent computing center includes: A receiving module 301, configured to receive encrypted data sent by a client; An allocation module 302, configured to allocate the encrypted data to a first node among a plurality of nodes, where the plurality of nodes are nodes deployed by the intelligent computing center cloud platform; A decryption module 303, configured to decrypt the encrypted data based on a first CPU of the first node to obtain decrypted data; A transmission module 304, configured to transmit the decrypted data to a first GPU of the first node; A development module 305, configured to perform model development based on the decrypted data through the first GPU.

[0090] In one embodiment, the allocation module 302 includes: A first acquisition unit, configured to acquire a first parameter of each of the multiple nodes; A determination unit, configured to determine the first node from the multiple nodes based on the first parameter and a preset rule; An allocation unit, configured to allocate the encrypted data to the first node; Wherein, the preset rule includes one of the following: When the first parameter is the number of times of receiving encrypted data, the first node is the node with the least number of times of receiving encrypted data among the multiple nodes; When the first parameter is a load parameter, the first node is the node with the smallest load parameter among the multiple nodes.

[0091] In one embodiment, the transmission module 304 includes: A second acquisition unit, configured to acquire a load parameter of the first GPU; A first transmission unit, configured to transmit the decrypted data to the first GPU when the load parameter is less than or equal to a first set load threshold.

[0092] In one embodiment, the transmission module 304 further includes: A second transmission unit, configured to transmit the decrypted data to a second GPU of a second node when the load parameter is greater than the first set load threshold, and the second node is a node other than the first node among the multiple nodes.

[0093] In one embodiment, the second CPU included in the second node has abnormal communication transmission, or the load parameter of the second GPU is less than or equal to a second set load threshold, and the second set load threshold is less than the first set load threshold.

[0094] In one embodiment, the transmission module 304 includes: A third transmission unit, configured to transmit the decrypted data accumulated in the memory of the first CPU to the first GPU when the data volume of the decrypted data accumulated in the memory of the first CPU reaches a preset data volume threshold.

[0095] In one embodiment, the preset data volume threshold is obtained by the following method: Acquire at least one of a batch size parameter, a model parameter quantity, and a training parameter corresponding to the model to be developed; Based on at least one of the batch size parameter, the model parameter quantity, and the training parameter, calculate to obtain the preset data volume threshold.

[0096] The model development data leakage prevention device for the inclusive computing power intelligent computing center cloud platform provided by the embodiments of the present invention can implement each process of the above-mentioned model development data leakage prevention method for the inclusive computing power intelligent computing center cloud platform. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0097] It should be noted that the model development data leakage prevention device for the inclusive computing power intelligent computing center cloud platform in the embodiments of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0098] The present invention also provides an electronic device. Refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 401, a processor 402, and a program or instruction running on the memory 401. When the program or instruction is executed by the processor 402, it can implement Figure 1 any step in the corresponding model development data leakage prevention method embodiment for the inclusive computing power intelligent computing center cloud platform and achieve the same beneficial effects. They will not be elaborated here.

[0099] Among them, the processor 402 can be a CPU, ASIC, FPGA, or GPU.

[0100] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned model development data leakage prevention method embodiments for the inclusive computing power intelligent computing center cloud platform can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0101] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1 any step in the corresponding model development data leakage prevention method embodiment for the inclusive computing power intelligent computing center cloud platform and can achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium can be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0102] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned Figure 1 corresponding model development data leakage prevention method embodiment for the inclusive computing power intelligent computing center cloud platform and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0103] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, as used in this application, "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including the 7 cases of A alone, B alone, C alone, A and B both present, B and C both present, A and C both present, and A, B and C all present.

[0104] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.

[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of this application.

[0106] The embodiments of this application are described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them fall within the protection scope of this application.

Claims

1. A method for preventing data leakage in model development for an inclusive computing power intelligent computing center cloud platform, characterized in that, including: Step S1, receiving encrypted data sent by a client; Step S2, allocating the encrypted data to a first node among multiple nodes, where the multiple nodes are nodes deployed by an intelligent computing center cloud platform; Step S3, decrypting the encrypted data based on a first CPU of the first node to obtain decrypted data; Step S4, transmitting the decrypted data to a first GPU of the first node; Step S5, performing model development based on the decrypted data through the first GPU.

2. The method according to claim 1, characterized in that, The step S2 includes: Step S21, obtaining first parameters of each node among the multiple nodes; Step S22, determining the first node from the multiple nodes based on the first parameters and a preset rule; Step S23, allocating the encrypted data to the first node; wherein, the preset rule includes one of the following: When the first parameter is the number of times of receiving encrypted data, the first node is the node among the multiple nodes with the least number of times of receiving encrypted data; When the first parameter is a load parameter, the first node is the node among the multiple nodes with the smallest load parameter.

3. The method according to claim 1, characterized in that The step S4 includes: Step S41, obtaining a load parameter of the first GPU; Step S42, when the load parameter is less than or equal to a first set load threshold, transmitting the decrypted data to the first GPU.

4. The method according to claim 3, characterized in that, The method further includes: Step S43, when the load parameter is greater than the first set load threshold, transmitting the decrypted data to a second GPU of a second node, where the second node is a node other than the first node among the multiple nodes.

5. The method according to claim 4, characterized in that, There is an abnormal communication transmission of a second CPU included in the second node, or a load parameter of the second GPU is less than or equal to a second set load threshold, and the second set load threshold is less than the first set load threshold.

6. The method according to claim 1, wherein The step S4 includes: Step S41', when the data volume of the decrypted data accumulated in the memory of the first CPU reaches a preset data volume threshold, transmitting the decrypted data accumulated in the memory of the first CPU to the first GPU.

7. The method according to claim 6, wherein The preset data volume threshold is obtained through the following method: obtaining at least one of a batch size parameter, a model parameter quantity, and a training parameter corresponding to a model to be developed; calculating the preset data volume threshold based on at least one of the batch size parameter, the model parameter quantity, and the training parameter.

8. A model development data leakage prevention device for an inclusive computing power intelligent computing center cloud platform, characterized in that, including: a receiving module, configured to receive encrypted data sent by a client; an allocating module, configured to allocate the encrypted data to a first node among multiple nodes, where the multiple nodes are nodes deployed by an intelligent computing center cloud platform; a decrypting module, configured to decrypt the encrypted data based on a first CPU of the first node to obtain decrypted data; a transmitting module, configured to transmit the decrypted data to a first GPU of the first node; a developing module, configured to perform model development based on the decrypted data through the first GPU.

9. An electronic device, characterized in that, including: A processor, a memory, and a program stored on the memory and executable on the processor, the program, when executed by the processor, implementing the steps of the method for preventing data leakage in model development for the inclusive computing power intelligent computing center cloud platform according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for preventing data leakage in model development for the inclusive computing power intelligent computing center cloud platform according to any one of claims 1 to 7 are implemented.

11. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the method for preventing data leakage in model development for the inclusive computing power intelligent computing center cloud platform according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Privacy protection-based cloud computing method, financial data cloud computing method and device

    CN115051816A

  • Data processing method and device, computer equipment, storage medium and program product

    CN116975792A

  • Private cloud database encryption method and device, equipment and medium

    CN117725602A

  • Federal learning GPU acceleration ciphertext aggregation method based on homomorphic encryption

    CN119483906A

  • Method and apparatus for processing wafer inspection task, system, and storage medium

    WO2022052523A1