Method and device for adaptive compression transmission of training task data of intelligent computing center model across computing power providers

By coordinating and scheduling computing and network resources among intelligent computing centers and selecting the optimal compression algorithm, the problem of slow data transmission caused by excessive task load in intelligent computing centers is solved, enabling rapid transmission of model training task data and improving training efficiency.

CN120342968BActive Publication Date: 2025-11-18DATACANVAS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510787974.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-11-18
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In intelligent computing centers, when users train large models, they usually choose the nearest intelligent computing center first. However, due to the high workload, the training needs cannot be met, resulting in slow data transmission speed and reduced training efficiency.

Method used

The first intelligent computing center pre-executes the model training task based on the model training task data. When there is insufficient idle computing resources, the second intelligent computing center is searched. The compression algorithm is determined based on comprehensive information, and the compressed model training task data is transmitted to the second intelligent computing center. By utilizing the coordinated scheduling of computing resources and network resources, the optimal compression algorithm is selected to reduce the amount of transmission and the time consumption.

Benefits of technology

It enables rapid transmission of model training task data, improves transmission rate, reduces transmission time, and enhances training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342968B_ABST
    Figure CN120342968B_ABST
Patent Text Reader

Abstract

The application provides a kind of across providing the intelligent computing center model training task data adaptive compression transmission method and device, it is related to intelligent computing center, wisdom calculation center and computing power infrastructure technical field, this method includes: first intelligent computing center is based on model training task data pre-execution model training task;In the case where the idle computing power resource of first intelligent computing center is insufficient, find second intelligent computing center;According to first information, determine the compression algorithm of model training task data;Based on compression algorithm, model training task data is compressed and transmitted to second intelligent computing center.In the method, by comprehensive analysis idle computing power resource information such as first intelligent computing center and second intelligent computing center, adaptive selection optimal compression algorithm, both can reduce actual transmission volume by compressing data, and can reduce transmission time consumption based on the collaborative scheduling of computing power resources and network resources, finally realize the improvement of model training task data transmission rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent computing centers, smart computing centers and computing infrastructure technology, and specifically to an adaptive compression transmission method and apparatus for model training task data across intelligent computing centers that provide computing power. Background Technology

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.

[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".

[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform a certain computing requirement. It is the computing power to achieve the target output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] In intelligent computing centers, users typically prioritize selecting the nearest center to handle training tasks when training large models. However, in practice, nearby centers often face excessive workloads and cannot meet training demands. In such cases, training task data must be transferred to a remote intelligent computing center to complete the model training. Large model training often requires numerous iterations (e.g., tens of thousands), and slow data transfer rates increase waiting times and reduce training efficiency. Therefore, since the emergence of intelligent computing centers, the rapid transfer of model training task data has become a pressing technical challenge. Summary of the Invention

[0008] This invention provides an adaptive compression transmission method and apparatus for model training task data across intelligent computing centers that provide computing power, in order to solve the problem of how to quickly transmit model training task data.

[0009] To solve the above problems, the present invention is implemented as follows:

[0010] In a first aspect, the present invention provides an adaptive compression transmission method for model training task data across intelligent computing centers providing computing power, executed by a first intelligent computing center, comprising:

[0011] Step S1: Pre-execute the model training task based on the model training task data;

[0012] Step S2: If the idle computing power resources of the first intelligent computing center are insufficient, search for the second intelligent computing center;

[0013] Step S3: Determine the compression algorithm for the model training task data based on the first information. The first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center.

[0014] Step S4: Based on the compression algorithm, the model training task data is compressed and then transmitted to the second intelligent computing center.

[0015] In one embodiment, after step S2 and before step S3, the method further includes:

[0016] Step S5: Establish a communication link with the second intelligent computing center;

[0017] Step S6: Receive second information transmitted by the second intelligent computing center through the communication link. The second information includes the idle computing resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0018] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0019] After step S3 and before step S4, the method further includes:

[0020] Step S7: Based on the third information of each of the N second intelligent computing centers, determine the target second intelligent computing center among the N second intelligent computing centers. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0021] Step S4 includes:

[0022] Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, the model training task data is compressed and transmitted to the target second intelligent computing center.

[0023] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0024] After step S3 and before step S4, the method further includes:

[0025] Step S8: Based on the third information of each of the N second intelligent computing centers, the model training task data is divided into N parts of first model training task data, and the first model training task data to be processed by each second intelligent computing center is determined. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0026] Step S4 includes:

[0027] Step S42: Based on the compression algorithm corresponding to each of the second intelligent computing centers, compress the first model training task data to be processed by each of the second intelligent computing centers and transmit it to each of the second intelligent computing centers.

[0028] In one embodiment, after step S4, the method further includes:

[0029] Step S9: Assign the model training task to the second intelligent computing center so that the second intelligent computing center can execute the model training task based on the model training data.

[0030] In one embodiment, step S4 includes:

[0031] Step S43: Based on the compression algorithm, the model training task data and the corresponding verification data are compressed and transmitted to the second intelligent computing center. The verification data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0032] Secondly, the present invention also provides an adaptive compression transmission device for model training task data across intelligent computing centers providing computing power, applied to a first intelligent computing center, comprising:

[0033] The first pre-execution module is used to pre-execute the model training task based on the model training task data.

[0034] The first search module is used to search for a second intelligent computing center when the idle computing resources of the first intelligent computing center are insufficient.

[0035] The first determining module is used to determine the compression algorithm for the model training task data based on the first information, wherein the first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center and the number of online tasks of the second intelligent computing center.

[0036] The first transmission module is used to compress the model training task data based on the compression algorithm and then transmit it to the second intelligent computing center.

[0037] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the adaptive compression transmission method for training task data across intelligent computing centers providing computing power as described in the first aspect above.

[0038] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the adaptive compression transmission method for model training task data across intelligent computing centers providing computing power as described in the first aspect above.

[0039] Fifthly, the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the adaptive compression transmission method for training task data across intelligent computing centers providing computing power as described in the first aspect above.

[0040] In this invention, a first intelligent computing center pre-executes model training tasks based on model training task data. If the first intelligent computing center lacks sufficient idle computing resources, a second intelligent computing center is located. Based on first information, a compression algorithm for the model training task data is determined. This first information includes the idle computing resources of the first and second intelligent computing centers, the workload of the model training tasks, the bandwidth resources between the first and second intelligent computing centers, the amount of online task data in the first intelligent computing center, and the number of online tasks in the second intelligent computing center. Based on the compression algorithm, the model training task data is compressed and transmitted to the second intelligent computing center. This method, by comprehensively analyzing information such as the idle computing resources of the first and second intelligent computing centers, the amount of model training tasks, the cross-center bandwidth resources, and the online task load of the two intelligent computing centers, adaptively selects the optimal compression algorithm. This reduces the actual transmission volume by compressing the model training task data and reduces transmission time based on the coordinated scheduling of computing and network resources, ultimately improving the data transmission rate of the model training tasks. Attached Figure Description

[0041] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of an adaptive compression transmission method for model training task data across intelligent computing centers that provide computing power, provided by the present invention.

[0043] Figure 2 This is one of the data transmission diagrams of a first intelligent computing center and a second intelligent computing center provided by the present invention;

[0044] Figure 3 This is a schematic diagram of the idle computing resources of a first intelligent computing center and a second intelligent computing center provided by the present invention;

[0045] Figure 4 This is the second of a data transmission schematic diagram of a first intelligent computing center and a second intelligent computing center provided by the present invention;

[0046] Figure 5 This is a schematic diagram of a data attribute field provided by the present invention;

[0047] Figure 6This is a structural diagram of an adaptive compression transmission device for model training task data across intelligent computing centers that provide computing power, provided by the present invention.

[0048] Figure 7 This is a structural diagram of an electronic device provided by the present invention. Detailed Implementation

[0049] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0050] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.

[0051] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values ​​indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0052] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.

[0053] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices within servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), while the commonly used unit of measurement for performance is the number of read / write operations per second (IOPS / TB). Disaster recovery ratio is an important indicator of security and reliability.

[0054] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.

[0055] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0056] The “computing power” mentioned in this invention includes: general computing power, intelligent computing power, and supercomputing power.

[0057] The "general-purpose computing power" mentioned in this invention refers to the computing power provided by servers based on central processing unit (CPU) chips, which is used to support basic general-purpose computing such as cloud computing and edge computing.

[0058] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPUs (Graphics Processing Units), Field Programmable Gate Arrays (FPGAs), and Application Specific Integrated Circuits (ASICs) for various innovative artificial intelligence applications, such as natural language processing and machine vision.

[0059] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0060] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as the development, training, and inference of deep learning models). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.

[0061] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".

[0062] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.

[0063] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0064] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.

[0065] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.

[0066] In existing technologies, when users train large models in intelligent computing centers, they typically prioritize choosing the nearest intelligent computing center to handle the training task. However, in practice, nearby intelligent computing centers are often overloaded and unable to meet training demands. In such cases, training task data needs to be transferred to a remote intelligent computing center to complete the model training job. Large model training usually requires numerous iterations (e.g., tens of thousands of times). If the data transfer speed is slow, it will increase the waiting time for training tasks and reduce training efficiency. Therefore, since the emergence of intelligent computing centers, how to quickly transfer model training task data has become a pressing technical problem to be solved. To achieve rapid transmission of model training task data, this invention involves a first intelligent computing center pre-executing model training tasks based on the data. If the first intelligent computing center lacks sufficient idle computing resources, a second intelligent computing center is located. Based on first information, a compression algorithm for the model training task data is determined. This first information includes the idle computing resources of the first and second intelligent computing centers, the workload of the model training tasks, the bandwidth resources between the first and second intelligent computing centers, the amount of online task data at the first intelligent computing center, and the number of online tasks at the second intelligent computing center. Based on the compression algorithm, the model training task data is compressed and transmitted to the second intelligent computing center. This method adaptively selects the optimal compression algorithm by comprehensively analyzing information such as the idle computing resources of the first and second intelligent computing centers, the workload of the model training tasks, the cross-center bandwidth resources, and the online task load of the two intelligent computing centers. This reduces the actual transmission volume through data compression and reduces transmission time through the coordinated scheduling of computing and network resources, ultimately improving the data transmission rate of the model training tasks.

[0067] For details, please see Figure 1 , Figure 1 This is a flowchart of an adaptive compression transmission method for model training task data across intelligent computing centers providing computing power, provided by the present invention. The method is executed by a first intelligent computing center. Figure 1 As shown, it includes the following steps:

[0068] Step S1: Pre-execute the model training task based on the model training task data;

[0069] In this step, the user device first initiates the model training task, prioritizing the use of the local intelligent computing center (i.e., the first intelligent computing center mentioned above) to execute the model training task. It should be noted that the "pre-execution" mentioned above refers to the preparation for execution but not yet the actual execution.

[0070] Step S2: If the idle computing power resources of the first intelligent computing center are insufficient, search for the second intelligent computing center;

[0071] In this step, see Figure 2 If the first intelligent computing center finds that its idle computing resources are insufficient to perform the model training task, it can search for other nearby intelligent computing centers that have idle computing resources, namely the second intelligent computing center mentioned above.

[0072] See Figure 3 , Figure 3 The black portion represents the used computing resources of each intelligent computing center, and the white portion represents the idle computing resources of each intelligent computing center. Figure 3 It is evident that the first intelligent computing center has relatively few idle computing resources, while the second intelligent computing center has relatively many idle computing resources.

[0073] Step S3: Determine the compression algorithm for the model training task data based on the first information. The first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center.

[0074] In this step, to facilitate the rapid transmission of model training task data later, the first intelligent computing center can automatically select the optimal compression algorithm to compress the model training task data based on some real-time information (i.e., the aforementioned first information). Specifically, the first information may include the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the workload of the model training task, the bandwidth resources between the first and second intelligent computing centers, and the number of online tasks at the first and second intelligent computing centers. For example:

[0075] When the first intelligent computing center has relatively more idle computing resources and fewer online tasks, the second intelligent computing center has relatively more idle computing resources and fewer online tasks, and the bandwidth resources between the first and second intelligent computing centers are tight, a compression algorithm that consumes a lot of GPU resources for compression and decompression and has a high compression ratio can be selected.

[0076] When the first intelligent computing center has relatively more idle computing resources and fewer online tasks, the second intelligent computing center has relatively more idle computing resources and fewer online tasks, and the bandwidth resources between the first and second intelligent computing centers are sufficient, a compression algorithm that consumes a lot of GPU resources for compression and decompression and has a high compression ratio can be selected.

[0077] When the first intelligent computing center has a severe shortage of idle computing resources and a large number of online tasks, the second intelligent computing center has relatively more idle computing resources and a smaller number of online tasks, and the bandwidth resources between the first and second intelligent computing centers are tight, a compression algorithm with low GPU consumption for compression and low GPU consumption for decompression with a high compression ratio can be selected.

[0078] When the first intelligent computing center has a severe shortage of idle computing resources and a large number of online tasks, the second intelligent computing center has relatively more idle computing resources and a smaller number of online tasks, and the bandwidth resources between the first and second intelligent computing centers are sufficient, a compression algorithm that consumes less GPU resources for compression, consumes less GPU resources for decompression, and has a high compression ratio can be selected.

[0079] For example, compression algorithms may include Deflate, Snappy, LZ4, Zstandard, etc. The specific algorithm to be selected can be determined based on actual needs and the characteristics of the algorithm itself.

[0080] Step S4: Based on the compression algorithm, the model training task data is compressed and then transmitted to the second intelligent computing center.

[0081] In this step, the first intelligent computing center compresses the model training task data according to the compression algorithm selected in step S3 and then transmits it to the second intelligent computing center via the network. After that, the second intelligent computing center decompresses the model training task data based on the corresponding decompression algorithm.

[0082] Specifically, when choosing a compression algorithm that is both GPU-intensive for compression and decompression, but also has a high compression ratio, a sliding window strategy can be used to transmit model training task data. The sliding window strategy treats data transmission as a flow control process between the "sender" and the "receiver," defining a "window" to limit the amount of data the sender can send before receiving acknowledgment. The window size dynamically adjusts based on factors such as network conditions and the receiver's processing capacity, similar to a "sliding" motion.

[0083] When selecting a compression algorithm that consumes less GPU resources for compression and decompression while achieving a high compression ratio, model training task data can be transferred via large-scale data transfer. "Large-scale data transfer" refers to the transfer of massive amounts of model training task data (usually reaching TB or even PB levels) between different intelligent computing centers.

[0084] In the above embodiment, the first intelligent computing center pre-executes the model training task based on the model training task data; when the idle computing resources of the first intelligent computing center are insufficient, the second intelligent computing center is searched; based on first information, a compression algorithm for the model training task data is determined, the first information including the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the workload of the model training task, the bandwidth resources between the first and second intelligent computing centers, the amount of online task data of the first intelligent computing center, and the number of online tasks of the second intelligent computing center; based on the compression algorithm, the model training task data is compressed and transmitted to the second intelligent computing center. In this embodiment, by comprehensively analyzing the idle computing resources of the first and second intelligent computing centers, the amount of model training tasks, the cross-center bandwidth resources, and the online task load of the two intelligent computing centers, the optimal compression algorithm is adaptively selected. This not only reduces the actual transmission volume by compressing the data, but also reduces the transmission time based on the coordinated scheduling of computing resources and network resources, ultimately improving the data transmission rate of the model training task.

[0085] In one embodiment, after step S2 and before step S3, the method further includes:

[0086] Step S5: Establish a communication link with the second intelligent computing center;

[0087] Step S6: Receive the second information transmitted by the second intelligent computing center through the communication link. The second information includes the idle computing resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0088] In the above embodiments, see Figure 4 The first intelligent computing center establishes a communication link with the second intelligent computing center by shaking hands. The handshake is a link establishment mechanism in network communication, similar to both parties confirming their identities and communication rules. Specifically, the first intelligent computing center sends a connection request to the second intelligent computing center, which receives it and returns confirmation information. Both parties then complete protocol negotiation (such as transmission protocols and data formats), ultimately establishing a bidirectional data transmission link, laying the foundation for subsequent information exchange and data transmission.

[0089] After the communication link is established, the second intelligent computing center proactively sends "secondary information" containing its own status to the first intelligent computing center. The core data includes idle computing resources (such as the number of available GPUs, storage capacity, etc.) and the number of online tasks (the load of currently running tasks). This information is transmitted to the first intelligent computing center in real time via the communication link for its subsequent decision-making.

[0090] In this embodiment, the first intelligent computing center can accurately grasp the availability of the other party's computing resources by obtaining the idle computing resources and task load of the second intelligent computing center, avoiding blind resource allocation caused by information asymmetry, and providing data support for subsequent task scheduling or data transmission strategies.

[0091] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0092] After step S3 and before step S4, the method further includes:

[0093] Step S7: Based on the third information of each of the N second intelligent computing centers, determine the target second intelligent computing center among the N second intelligent computing centers. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0094] Step S4 includes:

[0095] Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, the model training task data is compressed and transmitted to the target second intelligent computing center.

[0096] In the above embodiments, when the first intelligent computing center finds multiple second intelligent computing centers nearby that can share the model training task, the first intelligent computing center can filter based on the third information of each second intelligent computing center. The third information includes two aspects:

[0097] Idle computing resources: such as the number of available GPU cores, CPU computing power, and memory capacity, which reflect the computing capacity of the intelligent computing center to undertake new tasks; Bandwidth resources: the network bandwidth between the first intelligent computing center and the second intelligent computing center, which determines the theoretical maximum data transmission rate.

[0098] The first intelligent computing center can comprehensively evaluate the third information of each second intelligent computing center through algorithms (such as weighted scoring). For example, it can assign higher weights to idle computing power (to ensure that tasks can be processed quickly) while considering bandwidth (to avoid excessive transmission time), and finally select the one with the best overall conditions as the target second intelligent computing center. Then, the first intelligent computing center compresses the model training task data with a compression algorithm matched to the target second intelligent computing center, and transmits the compressed model training task data to the target second intelligent computing center.

[0099] In this embodiment, if the first intelligent computing center finds multiple second intelligent computing centers nearby that can share the model training task, the target second intelligent computing center with the best overall conditions can be selected based on the idle computing power resources and bandwidth resources, thereby ensuring that the model training task can be started and executed quickly in the target second intelligent computing center, thereby improving the overall training efficiency of the model.

[0100] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0101] After step S3 and before step S4, the method further includes:

[0102] Step S8: Based on the third information of each of the N second intelligent computing centers, the model training task data is divided into N parts of first model training task data, and the first model training task data to be processed by each second intelligent computing center is determined. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0103] Step S4 includes:

[0104] Step S42: Based on the compression algorithm corresponding to each of the second intelligent computing centers, compress the first model training task data to be processed by each of the second intelligent computing centers and transmit it to each of the second intelligent computing centers.

[0105] In the above embodiments, when the first intelligent computing center finds N second intelligent computing centers nearby that can share the model training task, the first intelligent computing center can split the complete model training task data into N parts of first model training task data, and allocate the specific amount of model training task data to be processed to each second intelligent computing center based on information such as the idle computing power resources and bandwidth resources of each second intelligent computing center. For example, for second intelligent computing centers with more idle computing power resources and bandwidth resources, the first intelligent computing center can allocate more model training task data to them.

[0106] Then, for the first model training task data allocated to each second intelligent computing center, the first intelligent computing center selects a compression algorithm that is suitable for it to compress the first model training task data, and then transmits the compressed first model training task data to the corresponding second intelligent computing center, thereby realizing distributed data processing.

[0107] In this embodiment, the first model training task data is dynamically allocated according to the third information of each second intelligent computing center, avoiding the problem of some second intelligent computing centers being overloaded and others being idle due to the traditional "average allocation". This ensures that the computing resources of N second intelligent computing centers can be used in a balanced manner, thereby improving the overall training speed of the model training task.

[0108] In one embodiment, after step S4, the method further includes:

[0109] Step S9: Assign the model training task to the second intelligent computing center so that the second intelligent computing center can execute the model training task based on the model training data.

[0110] In the above embodiments, see again Figure 4 After the compressed model training task data is transmitted from the first intelligent computing center to the second intelligent computing center, a specific model training task is assigned to the second intelligent computing center, enabling the second intelligent computing center to perform training operations based on the received model training task data.

[0111] In this embodiment, the transmitted model training task data is associated with the assigned model training task to ensure that the second intelligent computing center can perform training based on the correct dataset, thus avoiding misalignment between the model training task and the model training data.

[0112] In one embodiment, step S4 includes:

[0113] Step S43: Based on the compression algorithm, the model training task data and the corresponding verification data are compressed and transmitted to the second intelligent computing center. The verification data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0114] In the above embodiments, see Figure 5 , Figure 5The data attribute fields include the magic number, source data center, target data center, compression algorithm, checksum (i.e., the aforementioned verification data), data length, and data content. The verification data type may include hash values ​​(such as MD5, SHA-256), parity check codes, cyclic redundancy check, etc., which are generated by calculation on the original data and can be used to detect whether the data has been tampered with, lost, or erroneous during transmission. The compression algorithm can simultaneously process the model training task data and the corresponding verification data, forming a compressed package which is then transmitted to the second intelligent computing center. The second intelligent computing center uses a corresponding decompression algorithm to decompress the compressed package and uses the verification data to verify the integrity of the model training task data.

[0115] In this implementation, model training task data may be corrupted during network transmission due to noise, packet loss, or other issues (such as bit flipping). Verification data can detect such problems in real time. For example, if pixel values ​​in the transmitted image training data are incorrect in the compressed package, the verification data can compare the differences, preventing the second intelligent computing center from training the model based on erroneous data, which would lead to a decrease in recognition accuracy.

[0116] Please see Figure 6 , Figure 6 This is a structural diagram of an adaptive compression and transmission device for model training task data across intelligent computing centers providing computing power, as provided by the present invention. Figure 6 As shown, the adaptive compression transmission device 600 for model training task data across the intelligent computing center providing computing power includes:

[0117] The first pre-execution module 601 is used to pre-execute the model training task based on the model training task data;

[0118] The first search module 602 is used to search for a second intelligent computing center when the idle computing power resources of the first intelligent computing center are insufficient.

[0119] The first determining module 603 is used to determine the compression algorithm of the model training task data according to the first information. The first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center and the number of online tasks of the second intelligent computing center.

[0120] The first transmission module 604 is used to compress the model training task data based on the compression algorithm and then transmit it to the second intelligent computing center.

[0121] In one embodiment, the apparatus further includes:

[0122] The first establishment module is used to establish a communication link with the second intelligent computing center;

[0123] The first transmission module is used to receive second information transmitted by the second intelligent computing center through the communication link. The second information includes the idle computing resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

[0124] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0125] The device further includes:

[0126] The second determining module is used to determine a target second intelligent computing center among the N second intelligent computing centers based on the third information of each of the N second intelligent computing centers. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0127] The first transmission module includes:

[0128] The first transmission unit is used to compress the model training task data and transmit it to the target second intelligent computing center based on the compression algorithm corresponding to the target second intelligent computing center.

[0129] In one embodiment, the number of the second intelligent computing centers is N, where N is an integer greater than 1;

[0130] The device further includes:

[0131] The third determining unit is used to divide the model training task data into N parts of first model training task data according to the third information of each of the N second intelligent computing centers, and to determine the first model training task data that each second intelligent computing center needs to process. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center.

[0132] The first transmission unit includes:

[0133] The second transmission unit is used to compress the first model training task data to be processed by each second intelligent computing center and transmit it to each second intelligent computing center based on the compression algorithm corresponding to each second intelligent computing center.

[0134] In one embodiment, the apparatus further includes:

[0135] The first allocation module is used to allocate the model training task to the second intelligent computing center, so that the second intelligent computing center can execute the model training task based on the model training data.

[0136] In one embodiment, the first transmission module includes:

[0137] The third transmission unit is used to compress the model training task data and the corresponding verification data based on the compression algorithm and then transmit them to the second intelligent computing center. The verification data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

[0138] The adaptive compression and transmission device for training task data of intelligent computing centers providing computing power provided by this invention can realize the various processes of the above-mentioned adaptive compression and transmission method for training task data of intelligent computing centers providing computing power. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0139] It should be noted that the computing power operation data distribution device of the intelligent computing center in this invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0140] The present invention also provides an electronic device, see [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 701, a processor 702, and a program or instructions stored in the memory 701 that run on the memory. When the program or instructions are executed by the processor 702, they can achieve the following: Figure 1 The corresponding steps in the implementation of the adaptive compression transmission method for training task data of intelligent computing centers that provide computing power, and the achievement of the same beneficial effects, will not be elaborated here.

[0141] The processor 702 can be a CPU, ASIC, FPGA, or GPU.

[0142] Those skilled in the art will understand that all or part of the steps of the above-described embodiments of the adaptive compression transmission method for model training tasks across intelligent computing centers providing computing power can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.

[0143] The present invention also provides a readable storage medium on which a computer program is stored, and which, when executed by a processor, can perform the above-described functions. Figure 1Any step in the corresponding embodiment of the adaptive compression and transmission method for model training task data across intelligent computing centers providing computing power can achieve the same technical effect, and will not be described again here to avoid repetition. The storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0144] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The corresponding implementation of the adaptive compression and transmission method for training task data of intelligent computing centers that provide computing power is described in detail below to avoid repetition.

[0145] The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0146] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0148] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for adaptive compression and transmission of model training task data across intelligent computing centers providing computing power, characterized in that, Executed by the First Intelligent Computing Center, including: Step S1: Pre-execute the model training task based on the model training task data; Step S2: If the idle computing power resources of the first intelligent computing center are insufficient, search for the second intelligent computing center; Step S3: Determine the compression algorithm for the model training task data based on the first information. The first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center, and the number of online tasks of the second intelligent computing center. Step S4: Based on the compression algorithm, the model training task data is compressed and then transmitted to the second intelligent computing center.

2. The adaptive compression and transmission method for model training task data across intelligent computing centers providing computing power according to claim 1, characterized in that, After step S2 and before step S3, the method further includes: Step S5: Establish a communication link with the second intelligent computing center; Step S6: Receive second information transmitted by the second intelligent computing center through the communication link. The second information includes the idle computing resources of the second intelligent computing center and the number of online tasks of the second intelligent computing center.

3. The adaptive compression and transmission method for model training task data across intelligent computing centers providing computing power according to claim 1, characterized in that, The number of the second intelligent computing centers is N, where N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S7: Based on the third information of each of the N second intelligent computing centers, determine the target second intelligent computing center among the N second intelligent computing centers. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center. Step S4 includes: Step S41: Based on the compression algorithm corresponding to the target second intelligent computing center, the model training task data is compressed and transmitted to the target second intelligent computing center.

4. The adaptive compression and transmission method for model training task data across intelligent computing centers providing computing power according to claim 1, characterized in that, The number of the second intelligent computing centers is N, where N is an integer greater than 1; After step S3 and before step S4, the method further includes: Step S8: Based on the third information of each of the N second intelligent computing centers, the model training task data is divided into N parts of first model training task data, and the first model training task data to be processed by each second intelligent computing center is determined. The third information includes the idle computing power resources of the second intelligent computing center and the bandwidth resources of the first intelligent computing center and the second intelligent computing center. Step S4 includes: Step S42: Based on the compression algorithm corresponding to each of the second intelligent computing centers, compress the first model training task data to be processed by each of the second intelligent computing centers and transmit it to each of the second intelligent computing centers.

5. The adaptive compression transmission method for model training task data across intelligent computing centers providing computing power according to any one of claims 1 to 4, characterized in that, After step S4, the method further includes: Step S9: Assign the model training task to the second intelligent computing center so that the second intelligent computing center can execute the model training task based on the model training data.

6. The adaptive compression transmission method for model training task data across intelligent computing centers providing computing power according to any one of claims 1 to 4, characterized in that, Step S4 includes: Step S43: Based on the compression algorithm, the model training task data and the corresponding verification data are compressed and transmitted to the second intelligent computing center. The verification data corresponding to the model training task data is used by the second intelligent computing center to verify the accuracy of the model training task data.

7. An adaptive compression and transmission device for model training task data across intelligent computing centers providing computing power, characterized in that, Applied to the first intelligent computing center, including: The first pre-execution module is used to pre-execute the model training task based on the model training task data. The first search module is used to search for a second intelligent computing center when the idle computing resources of the first intelligent computing center are insufficient. The first determining module is used to determine the compression algorithm for the model training task data based on the first information, wherein the first information includes the idle computing resources of the first intelligent computing center, the idle computing resources of the second intelligent computing center, the task volume of the model training task, the bandwidth resources between the first intelligent computing center and the second intelligent computing center, the number of online tasks of the first intelligent computing center and the number of online tasks of the second intelligent computing center. The first transmission module is used to compress the model training task data based on the compression algorithm and then transmit it to the second intelligent computing center.

8. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the adaptive compressed transmission method for model training tasks across intelligent computing centers providing computing power as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the adaptive compressed transmission method for model training task data across intelligent computing centers providing computing power as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of the adaptive compressed transmission method for model training task data across intelligent computing centers providing computing power as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data processing method and device of distributed assembly line and storage medium

    CN114428786A

  • Method and device for intelligent computing center to provide computing power resources through computing power package

    CN119739442A