Self-adaptive compression transmission method and device for training task data of cross-provision computing power intelligent computing center model

By adaptively selecting compression algorithms and resource scheduling between intelligent computing centers, the problem of slow data transmission speed of model training tasks in intelligent computing centers is solved, and efficient data transmission and training efficiency are improved.

CN120342968AActive Publication Date: 2025-07-18DATACANVAS LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510787974.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-18
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the intelligent computing center, the nearby intelligent computing center cannot meet the training needs of large models due to the high task load, resulting in slow data transmission speed of model training tasks and reduced training efficiency.

Method used

The first intelligent computing center pre-executes the model training task, finds the second intelligent computing center with insufficient idle computing resources, selects the optimal compression algorithm based on comprehensive information, compresses the model training task data and transmits it to the second intelligent computing center, and uses the coordinated scheduling of computing resources and network resources to reduce transmission time.

Benefits of technology

It realizes the rapid transmission of model training task data, improves the transmission rate and training efficiency, reduces the actual transmission amount and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342968A_ABST
    Figure CN120342968A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive compression transmission method and device for model training task data of an intelligent computing center with cross-provision of computing power, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the following steps: a first intelligent computing center pre-executes a model training task based on model training task data; searching a second intelligent computing center under the condition that idle computing power resources of the first intelligent computing center are insufficient; determining a compression algorithm of the model training task data according to the first information; and based on a compression algorithm, compressing the model training task data and then transmitting the data to a second intelligent computing center. In the method, information such as idle computing power resources of a first intelligent computing center and a second intelligent computing center is comprehensively analyzed, and an optimal compression algorithm is adaptively selected, so that the actual transmission quantity can be reduced by compressing data, and the transmission time consumption can be reduced based on cooperative scheduling of the computing power resources and network resources; and finally, the data transmission rate of the model training task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for adaptively compressing and transmitting model training task data of an intelligent computing center that provides computing power across different locations. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, to mainly provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). An intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to execute a certain computing requirement together, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which mainly provides services to society through computing power infrastructure.

[0007] In an intelligent computing center, when a user conducts large model training, they usually prefer to select a nearby intelligent computing center to undertake the training task first. However, in actual applications, nearby intelligent computing centers are often unable to meet the training requirements due to excessive task loads. At this time, it is necessary to transmit the training task data to a remote intelligent computing center to complete the model training operation. Large model training usually requires a large number of iterations (such as tens of thousands of times). If the data transmission speed is slow, it will lead to an increase in the waiting time of the training task and a decrease in training efficiency. Therefore, since the emergence of intelligent computing centers, how to quickly transmit model training task data has become an urgent technical problem to be solved. Summary of the Invention

[0008] The present invention provides a method and device for adaptively compressing and transmitting model training task data of an intelligent computing center that provides computing power across different locations, so as to solve the problem of how to quickly transmit model training task data.

[0009] To solve the above problems, the present invention is implemented as follows: In a first aspect, the present invention provides a method for adaptively compressing and transmitting model training task data across intelligent computing centers that provide computing power, which is executed by a first intelligent computing center and includes: Step S1: Pre-execute a model training task based on the model training task data; Step S2: When the idle computing power resources of the first intelligent computing center are insufficient, search for a second intelligent computing center; Step S3: Determine the compression algorithm for the model training task data according to first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; Step S4: Based on the compression algorithm, compress and transmit the model training task data to the second intelligent computing center.

[0010] In one embodiment, after step S2 and before step S3, the method further includes: Step S5: Establish a communication link with the second intelligent computing center; Step S6: Receive second information transmitted by the second intelligent computing center through the communication link, where the second information includes the idle computing power resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0011] In one embodiment, the number of second intelligent computing centers is N, where N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S7: Determine a target second intelligent computing center among the N second intelligent computing centers according to third information of each of the N second intelligent computing centers, where the third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; Step S4 includes: Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, compress and transmit the model training task data to the target second intelligent computing center.

[0012] In one embodiment, the number of second intelligent computing centers is N, where N is an integer greater than 1; After the step S3 and before the step S4, the method further includes: Step S8: According to the third information of each of the N second intelligent computing centers, divide the model training task data into N portions of first model training task data, and determine the first model training task data that each second intelligent computing center needs to process. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center. The step S4 includes: Step S42: Based on the compression algorithm corresponding to each second intelligent computing center, compress the first model training task data that each second intelligent computing center needs to process and then transmit it to each second intelligent computing center.

[0013] In one embodiment, after the step S4, the method further includes: Step S9: Allocate the model training task to the second intelligent computing center so that the second intelligent computing center executes the model training task based on the model training data.

[0014] In one embodiment, the step S4 includes: Step S43: Based on the compression algorithm, compress the model training task data and the verification data corresponding to the model training task and then transmit them to the second intelligent computing center. The verification data corresponding to the model training task is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0015] In a second aspect, the present invention further provides an adaptive compression transmission device for model training task data across intelligent computing centers providing computing power, which is applied to a first intelligent computing center and includes: A first pre-execution module, configured to pre-execute a model training task based on model training task data; A first search module, configured to search for a second intelligent computing center when the idle computing power resources of the first intelligent computing center are insufficient; A first determination module, configured to determine the compression algorithm of the model training task data according to first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; A first transmission module, configured to compress and transmit the model training task data to the second intelligent computing center based on the compression algorithm.

[0016] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for adaptively compressing and transmitting model training task data of a cross-providing computing power intelligent computing center model as described in the first aspect above.

[0017] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for adaptively compressing and transmitting model training task data of a cross-providing computing power intelligent computing center model as described in the first aspect above.

[0018] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, it implements the steps in the method for adaptively compressing and transmitting model training task data of a cross-providing computing power intelligent computing center model as described in the first aspect above.

[0019] In the present invention, the first intelligent computing center pre-executes a model training task based on model training task data; in the case where the idle computing power resources of the first intelligent computing center are insufficient, a second intelligent computing center is searched for; according to first information, a compression algorithm for the model training task data is determined, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the online task data volume of the first intelligent computing center, and the online task quantity of the second intelligent computing center; based on the compression algorithm, the model training task data is compressed and then transmitted to the second intelligent computing center. In this method, by comprehensively analyzing information such as the idle computing power resources of the first and second intelligent computing centers, the model training task volume, the cross-center bandwidth resources, and the online task loads of the two intelligent computing centers, an optimal compression algorithm is adaptively selected, so that both the actual transmission volume can be reduced by compressing the model training task data, and the transmission time can be reduced based on the coordinated scheduling of computing power resources and network resources, ultimately achieving an improvement in the transmission rate of the model training task data. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for the description of the present invention. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a flowchart of a method for adaptively compressing and transmitting model training task data of an intelligent computing center that provides computing power across different entities according to the present invention; Figure 2 It is one of the schematic diagrams of data transmission between a first intelligent computing center and a second intelligent computing center provided by the present invention; Figure 3 It is a schematic diagram of idle computing power resources of a first intelligent computing center and a second intelligent computing center provided by the present invention; Figure 4 It is another schematic diagram of data transmission between a first intelligent computing center and a second intelligent computing center provided by the present invention; Figure 5 It is a schematic diagram of data attribute fields provided by the present invention; Figure 6 It is a structural diagram of a device for adaptively compressing and transmitting model training task data of an intelligent computing center that provides computing power across different entities according to the present invention; Figure 7 It is a structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0022] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0023] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0024] The "computational power" (ComputationalPower, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP通用 +CP 智能 +CP 超级 。

[0025] The "Network Power (NP)" as described in the present invention refers to: It is an indication of the data transmission capacity of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0026] The "Storage Power (SP)" as described in the present invention refers to: It is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server-internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0027] The "computing power infrastructure" as described in the present invention refers to: A new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized computing, storage, transmission, and application of information.

[0028] The "new type of information infrastructure" as described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, satellite Internet, etc., computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, etc., and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0029] The "computing power" as described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0030] The "general computing power" as described in the present invention refers to: The computing power provided by servers based on central processing unit (CPU) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0031] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPUs (Graphics Processing Units), field programmable gate arrays (FPGAs), and application specific integrated circuits (ASICs), such as natural language processing, machine vision, and so on.

[0032] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0033] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPUs) and intelligent computing power (GPUs, FPGAs, ASICs, etc.). The intelligent computing center covers facilities, hardware, and software and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0034] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".

[0035] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0036] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0037] The "supercomputing center" described in the present invention refers to: that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0038] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0039] In the prior art, when users conduct large model training in an intelligent computing center, they usually give priority to the nearest intelligent computing center to undertake the training task. However, in actual applications, the nearby intelligent computing center is often unable to meet the training needs due to excessive task load. At this time, the training task data needs to be transmitted to the remote intelligent computing center to complete the model training operation. Large model training usually requires a large number of iterations (such as tens of thousands of times). If the data transmission speed is slow, the waiting time for the training task will increase and the training efficiency will decrease. Therefore, since the emergence of intelligent computing centers, how to quickly transmit model training task data has become a technical problem that needs to be solved urgently. In order to realize the rapid transmission of model training task data, in the present invention, the first intelligent computing center pre-executes the model training task based on the model training task data; when the idle computing power resources of the first intelligent computing center are insufficient, the second intelligent computing center is searched; according to the first information, the compression algorithm of the model training task data is determined, and the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task amount of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the online task data amount of the first intelligent computing center and the number of online tasks of the second intelligent computing center; based on the compression algorithm, the model training task data is compressed and transmitted to the second intelligent computing center. In this method, by comprehensively analyzing the idle computing power resources of the first intelligent computing center and the second intelligent computing center, the model training task amount, the cross-center bandwidth resources, and the online task load of the two intelligent computing centers, the optimal compression algorithm is adaptively selected, so that the actual transmission amount can be reduced by compressing the data, and the transmission time can be reduced based on the coordinated scheduling of computing power resources and network resources, and finally the data transmission rate of the model training task is improved.

[0040] For details, see Figure 1 , Figure 1 : is a flowchart of a method for adaptively compressing and transmitting model training task data across intelligent computing centers providing computing power provided by the present invention. The method for adaptively compressing and transmitting model training task data across intelligent computing centers providing computing power provided by the present invention is executed by the first intelligent computing center, such as Figure 1 As shown, the following steps are included: Step S1, pre-execute a model training task based on model training task data; In this step, first, the user device initiates a model training task and gives priority to using the local intelligent computing center (i.e., the first intelligent computing center mentioned above) to execute the model training task. It should be noted that the above "pre-execution" means preparing to execute but not actually executing yet.

[0041] Step S2: When the idle computing power resources of the first intelligent computing center are insufficient, search for the second intelligent computing center; In this step, refer to Figure 2 , when the first intelligent computing center finds that its idle computing power resources are insufficient to execute the model training task, the first intelligent computing center can search for other intelligent computing centers nearby that have idle computing power resources, that is, the second intelligent computing center mentioned above.

[0042] Refer to Figure 3 , Figure 3 The black part in Figure 3 represents the used computing power resources of each intelligent computing center, and the white part represents the idle computing power resources of each intelligent computing center. As can be seen from

[0043] Step S3: Determine the compression algorithm for the model training task data according to the first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; In this step, in order to transmit the model training task data quickly later, the first intelligent computing center can automatically select the optimal compression algorithm to compress the model training task data according to some real-time information (i.e., the above first information). Specifically, the first information can include the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center. Exemplarily: When the idle computing power resources of the first intelligent computing center are relatively large and the number of online tasks is small, the idle computing power resources of the second intelligent computing center are relatively large and the number of online tasks is small, and the bandwidth resources between the first intelligent computing center and the second intelligent computing center are tight, a compression algorithm with high GPU consumption for compression, high GPU consumption for decompression, and high compression ratio can be selected; When there are relatively many idle computing power resources and few online tasks in the first intelligent computing center, relatively many idle computing power resources and few online tasks in the second intelligent computing center, and sufficient bandwidth resources between the first intelligent computing center and the second intelligent computing center, a compression algorithm with high GPU consumption during compression, high GPU consumption during decompression, and a high compression ratio can be selected; When there are severely insufficient idle computing power resources and many online tasks in the first intelligent computing center, relatively many idle computing power resources and few online tasks in the second intelligent computing center, and tight bandwidth resources between the first intelligent computing center and the second intelligent computing center, a compression algorithm with relatively low GPU consumption during compression, low GPU consumption during decompression, and a high compression ratio can be selected; When there are severely insufficient idle computing power resources and many online tasks in the first intelligent computing center, relatively many idle computing power resources and few online tasks in the second intelligent computing center, and sufficient bandwidth resources between the first intelligent computing center and the second intelligent computing center, a compression algorithm with low GPU consumption during compression, low GPU consumption during decompression, and a high compression ratio can be selected.

[0044] Exemplarily, the compression algorithm may include Deflate, Snappy, LZ4, Zstandard, etc. Which algorithm to specifically select can be determined in combination with actual requirements and the characteristics of the algorithm itself.

[0045] Step S4: Based on the compression algorithm, compress the model training task data and transmit it to the second intelligent computing center.

[0046] In this step, the first intelligent computing center compresses the model training task data according to the compression algorithm selected in step S3 and transmits it to the second intelligent computing center through the network. After that, the second intelligent computing center decompresses the model training task data based on the corresponding decompression algorithm.

[0047] Specifically, in the case of selecting a compression algorithm with high GPU consumption during compression, high GPU consumption during decompression, and a high compression ratio, the sliding window strategy can be used to transmit the model training task data. The sliding window strategy refers to regarding data transmission as a flow control process between the "sender" and the "receiver", and defining a "window" to limit the amount of data that the sender can send before receiving an acknowledgment. The window size will be dynamically adjusted according to factors such as network status and the processing capacity of the receiver, similar to a "sliding" change; In the case of selecting a compression algorithm with low GPU consumption during compression, low GPU consumption during decompression, and a high compression ratio, the model training task data can be transmitted through large data volume transmission. "Large data volume transmission" means transmitting model training task data with a huge scale (usually reaching the TB level or even the PB level) between different intelligent computing centers.

[0048] In the above embodiments, the first intelligent computing center pre-executes a model training task based on model training task data; when the idle computing power resources of the first intelligent computing center are insufficient, it searches for a second intelligent computing center; according to first information, it determines a compression algorithm for the model training task data, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the online task data volume of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; based on the compression algorithm, it compresses the model training task data and transmits it to the second intelligent computing center. In this embodiment, by comprehensively analyzing information such as the idle computing power resources of the first and second intelligent computing centers, the model training task volume, the cross-center bandwidth resources, and the online task loads of the two intelligent computing centers, an optimal compression algorithm is adaptively selected, which can not only reduce the actual transmission volume by compressing data, but also reduce the transmission time consumption based on the coordinated scheduling of computing power resources and network resources, and ultimately improve the transmission rate of the model training task data.

[0049] In one embodiment, after step S2 and before step S3, the method further includes: Step S5: Establish a communication link with the second intelligent computing center; Step S6: Receive second information transmitted by the second intelligent computing center through the communication link, where the second information includes the idle computing power resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0050] In the above embodiments, refer to Figure 4 , the first intelligent computing center establishes a communication link with the second intelligent computing center by shaking hands with the second intelligent computing center. Handshaking is a link establishment mechanism in network communication, similar to both parties confirming their identities and communication rules. The specific process is as follows: The first intelligent computing center sends a connection request to the second intelligent computing center, and after the second intelligent computing center receives it, it returns a confirmation message. The two parties complete protocol negotiation (such as transmission protocol, data format, etc.), and finally establish a communication link that can transmit data bidirectionally, laying a foundation for subsequent information interaction and data transmission.

[0051] After the communication link is established, the second intelligent computing center actively sends "second information" containing its own status to the first intelligent computing center. The core data is the idle computing power resources (such as the number of available GPUs, storage capacity, etc.) and the number of online tasks (the load situation of currently running tasks). This information is transmitted to the first intelligent computing center in real time through the communication link for its subsequent decision-making use.

[0052] In this embodiment, by obtaining the idle computing power resources and task loads of the second intelligent computing center, the first intelligent computing center can accurately grasp the availability of the computing resources of the other party, avoid the blindness of resource allocation caused by information asymmetry, and provide data support for subsequent task scheduling or data transmission strategies.

[0053] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S7: Determine a target second intelligent computing center among the N second intelligent computing centers according to the third information of each second intelligent computing center in the N second intelligent computing centers. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; Step S4 includes: Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, compress the model training task data and transmit it to the target second intelligent computing center.

[0054] In the above embodiment, when the first intelligent computing center finds multiple second intelligent computing centers nearby that can share the model training task for it, the first intelligent computing center can screen according to the third information of each second intelligent computing center. The third information includes two aspects: Idle computing power resources: such as the number of available GPU cores, CPU computing power, and memory capacity, etc., which reflect the computing ability of this intelligent computing center to undertake new tasks; Bandwidth resources: the network bandwidth between the first intelligent computing center and this second intelligent computing center, which determines the theoretical maximum rate of data transmission.

[0055] The first intelligent computing center can comprehensively evaluate the third information of each second intelligent computing center through an algorithm (such as the weighted scoring method). For example, a higher weight is given to the idle computing power (to ensure that tasks can be processed quickly), and the bandwidth is also considered (to avoid excessive transmission time). Finally, the one with the best comprehensive conditions is selected as the target second intelligent computing center. Then the first intelligent computing center uses the compression algorithm matched according to the target second intelligent computing center to compress the model training task data, and transmits the compressed model training task data to the target second intelligent computing center.

[0056] In this embodiment, when the first intelligent computing center finds multiple second intelligent computing centers nearby that can share the model training task for it, it can select the target second intelligent computing center with the optimal comprehensive conditions based on the idle computing power resources and bandwidth resources, so as to ensure that the model training task can be quickly started and executed at the target second intelligent computing center, thereby improving the overall training efficiency of the model.

[0057] In one embodiment, the number of the second intelligent computing centers is N, and N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S8: According to the third information of each of the N second intelligent computing centers, divide the model training task data into N pieces of first model training task data, and determine the first model training task data to be processed by each second intelligent computing center. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; Step S4 includes: Step S42: Based on the compression algorithm corresponding to each second intelligent computing center, compress the first model training task data to be processed by each second intelligent computing center and then transmit it to each second intelligent computing center.

[0058] In the above embodiment, when the first intelligent computing center finds N second intelligent computing centers nearby that can share the model training task for it, the first intelligent computing center can split the complete model training task data into N pieces of first model training task data, and according to the information such as the idle computing power resources and bandwidth resources of each second intelligent computing center, allocate the specific amount of model training task data to be processed for each second intelligent computing center. Exemplarily, for a second intelligent computing center with more idle computing power resources and bandwidth resources, the first intelligent computing center can allocate more model training task data to it.

[0059] Then, for the first model training task data allocated to each second intelligent computing center, the first intelligent computing center respectively selects an appropriate compression algorithm to compress the first model training task data, and then transmits the compressed first model training task data to the corresponding second intelligent computing center, thereby realizing distributed data processing.

[0060] In this embodiment, the first model training task data is dynamically allocated according to the third information of each second intelligent computing center, avoiding the problems of overload of some second intelligent computing centers and idle of some second intelligent computing centers caused by traditional "average distribution", enabling the computing power resources of the N second intelligent computing centers to be evenly utilized, and thus improving the overall training speed of the model training task.

[0061] In one embodiment, after the step S4, the method further includes: Step S9: Allocate the model training task to the second intelligent computing center so that the second intelligent computing center executes the model training task based on the model training data.

[0062] In the above embodiment, continue to refer to Figure 4 , after the first intelligent computing center transmits the compressed model training task data to the second intelligent computing center, allocate a specific model training task to the second intelligent computing center, so that the second intelligent computing center performs a training operation based on the received model training task data.

[0063] In this embodiment, associate the transmitted model training task data with the allocated model training task to ensure that the second intelligent computing center can execute the training based on the correct data set, and avoid the dislocation between the model training task and the model training data.

[0064] In one embodiment, the step S4 includes: Step S43: Based on the compression algorithm, compress the model training task data and the check data corresponding to the model training task data, and then transmit them to the second intelligent computing center. The check data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0065] In the above embodiment, refer to Figure 5 , Figure 5 The data attribute fields in include magic number, source data center, target data center, compression algorithm, checksum (i.e., the above-mentioned check data), data length, and data content. Among them, the types of check data may include hash values (such as MD5, SHA-256), parity check codes, cyclic redundancy checks, etc. It is generated by calculating the original data and can be used to detect whether the data has been tampered with, lost, or in error during transmission. The compression algorithm can process the model training task data and the check data corresponding to the model training task data at the same time, form a compressed package and transmit it to the second intelligent computing center. The second intelligent computing center uses the corresponding decompression algorithm to decompress the compressed package and uses the check data to verify the integrity of the model training task data.

[0066] In this implementation manner, the model training task data may be in error (such as bit flipping) during network transmission due to problems such as noise and packet loss. The check data can detect such problems in real time. For example, if there are pixel value errors in the transmitted image training data in the compressed package, the check data can compare the differences to avoid the second intelligent computing center training the model based on incorrect data, resulting in a decrease in the recognition accuracy.

[0067] Please refer to Figure 6 , Figure 6 which is the structural diagram of an adaptive compression transmission device for model training task data across intelligent computing centers providing computing power according to the present invention. As Figure 6 shown, the adaptive compression transmission device 600 for model training task data across intelligent computing centers providing computing power includes: A first pre-execution module 601 for pre-executing a model training task based on the model training task data; A first search module 602 for searching for a second intelligent computing center when the idle computing power resources of the first intelligent computing center are insufficient; A first determination module 603 for determining the compression algorithm for the model training task data according to first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; A first transmission module 604 for compressing and transmitting the model training task data to the second intelligent computing center based on the compression algorithm.

[0068] In one embodiment, the device further includes: A first establishment module for establishing a communication link with the second intelligent computing center; A first transmission module for receiving second information transmitted by the second intelligent computing center through the communication link, where the second information includes the idle computing power resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0069] In one embodiment, the number of second intelligent computing centers is N, where N is an integer greater than 1; The device further includes: A second determination module for determining a target second intelligent computing center among the N second intelligent computing centers according to third information of each of the N second intelligent computing centers, where the third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; The first transmission module includes: A first transmission unit for compressing and transmitting the model training task data to the target second intelligent computing center based on the compression algorithm corresponding to the target second intelligent computing center.

[0070] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1; The device further includes: A third determination unit, configured to split the model training task data into N pieces of first model training task data according to the third information of each of the N second intelligent computing centers, and determine the first model training task data that each second intelligent computing center needs to process, where the third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; The first transmission unit includes: A second transmission unit, configured to compress the first model training task data that each second intelligent computing center needs to process based on a compression algorithm corresponding to each second intelligent computing center, and then transmit it to each second intelligent computing center.

[0071] In one embodiment, the device further includes: A first allocation module, configured to allocate the model training task to the second intelligent computing center, so that the second intelligent computing center executes the model training task based on the model training data.

[0072] In one embodiment, the first transmission module includes: A third transmission unit, configured to compress and transmit the model training task data and the verification data corresponding to the model training task to the second intelligent computing center based on the compression algorithm, where the verification data corresponding to the model training task is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0073] The device for adaptively compressing and transmitting model training task data across intelligent computing centers providing computing power according to the present invention can implement each process of the above-described method for adaptively compressing and transmitting model training task data across intelligent computing centers providing computing power. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they are not described herein again.

[0074] It should be noted that the device for distributing computing power operation data of the intelligent computing center in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0075] The present invention further provides an electronic device. Refer to Figure 7 , Figure 7It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 701, a processor 702, and a program or instruction stored on the memory 701 and running. When the program or instruction is executed by the processor 702, it can implement Figure 1 Any steps in the corresponding embodiment of the adaptive compression transmission method for model training task data of the intelligent computing center that provides computing power and achieve the same beneficial effects will not be elaborated here.

[0076] Among them, the processor 702 can be a CPU, ASIC, FPGA or GPU.

[0077] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiment of the adaptive compression transmission method for model training task data of the intelligent computing center that provides computing power can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0078] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above Figure 1 Any steps in the corresponding embodiment of the adaptive compression transmission method for model training task data of the intelligent computing center that provides computing power, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. The storage medium such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disc, etc.

[0079] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, it implements the above Figure 1 Each process of the corresponding embodiment of the adaptive compression transmission method for model training task data of the intelligent computing center that provides computing power, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0080] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, the use of "and / or" in this application means at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, and A and B both exist, B and C both exist, A and C both exist, and A, B, and C all exist, a total of 7 cases.

[0081] It should be noted that in this text, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising such element.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of the present application.

[0083] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. An adaptive compression and transmission method for model training task data across intelligent computing centers providing computing power, characterized in that, Executed by the first intelligent computing center, including: Step S1: Pre-execute the model training task based on the model training task data; Step S2: When the idle computing power resources of the first intelligent computing center are insufficient, search for the second intelligent computing center; Step S3: Determine the compression algorithm for the model training task data according to the first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; Step S4: Based on the compression algorithm, compress and transmit the model training task data to the second intelligent computing center.

2. The method for adaptively compressing and transmitting task data of cross-providing computing power intelligent computing center model training according to claim 1, wherein After step S2 and before step S3, the method further includes: Step S5: Establish a communication link with the second intelligent computing center; Step S6: Receive the second information transmitted by the second intelligent computing center through the communication link, where the second information includes the idle computing power resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

3. The adaptive compression and transmission method for the intelligent computing center model training task data across the provision of computing power according to claim 1, characterized in that The number of the second intelligent computing centers is N, and N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S7: Determine the target second intelligent computing center among the N second intelligent computing centers according to the third information of each of the N second intelligent computing centers, where the third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; Step S4 includes: Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, compress and transmit the model training task data to the target second intelligent computing center.

4. The method for adaptively compressing and transmitting intelligent computing center model training task data across provided computing powers according to claim 1, wherein The number of the second intelligent computing centers is N, and N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S8: According to the third information of each of the N second intelligent computing centers, divide the model training task data into N pieces of first model training task data, and determine the first model training task data that each second intelligent computing center needs to process, where the third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources between the first intelligent computing center and the second intelligent computing center; Step S4 includes: Step S42: Based on the compression algorithm corresponding to each second intelligent computing center, compress and transmit the first model training task data that each second intelligent computing center needs to process to each second intelligent computing center.

5. The method for adaptively compressing and transmitting task data of a cross-providing computing power intelligent computing center model training according to any one of claims 1 to 4, characterized in that After step S4, the method further includes: Step S9: Allocate the model training task to the second intelligent computing center so that the second intelligent computing center executes the model training task based on the model training data.

6. The adaptive compression and transmission method for model training task data of an intelligent computing center that provides computing power across different entities, according to any one of claims 1 to 4, wherein Step S4 includes: Step S43: Based on the compression algorithm, compress the model training task data and the verification data corresponding to the model training task data, and then transmit them to the second intelligent computing center. The verification data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

7. An adaptive compression and transmission device for model training task data across intelligent computing centers providing computing power, characterized in that, Applied to the first intelligent computing center, it includes: A first pre-execution module, configured to pre-execute the model training task based on the model training task data; A first search module, configured to search for a second intelligent computing center when the idle computing power resources of the first intelligent computing center are insufficient; A first determination module, configured to determine the compression algorithm of the model training task data according to first information, where the first information includes the idle computing power resources of the first intelligent computing center, the idle computing power resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; A first transmission module, configured to compress the model training task data based on the compression algorithm and transmit it to the second intelligent computing center.

8. An electronic device, characterized in that, It includes: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the cross-providing computing power intelligent computing center model training task data adaptive compression transmission method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the cross-providing computing power intelligent computing center model training task data adaptive compression transmission method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions. When the computer instructions are executed by the processor, it implements the steps of the cross-providing computing power intelligent computing center model training task data adaptive compression transmission method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data compression method and equipment and computer readable storage medium

    CN108197168A

  • Federated learning architecture under dynamic bandwidth and unreliable network and compression algorithm of architecture

    CN111447083A

  • Data processing method and device of distributed assembly line and storage medium

    CN114428786A

  • Large model distributed training method oriented to cloud environment and related equipment

    CN116341652A

  • Computing power network data transmission method and device, electronic equipment and storage medium

    CN116582547A