Data processing system and method, and electronic device, storage medium and program product

By acquiring and sorting the LA group identification information of cloud computing devices through terminal devices, the problem of high communication latency between cloud computing devices is solved, enabling more efficient computing and resource utilization, and ensuring the security of the computing cluster.

WO2025260948A1PCT designated stage Publication Date: 2025-12-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/089484
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-04-17
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In existing technologies, the communication network latency between cloud computing devices is high, resulting in low computing efficiency, and communication across LA groups consumes a lot of network resources.

Method used

By obtaining the LA group identification information of cloud computing devices through terminal devices, and sorting the cloud computing devices according to the LA group identification information, computing tasks can be arranged within the same LA group, reducing communication across LA groups.

Benefits of technology

It reduces network latency between cloud computing devices, reduces network transmission resource consumption, improves computing efficiency, and ensures the information security of the computing cluster by providing LA group identifier characterization information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025089484_26122025_PF_FP_ABST
    Figure CN2025089484_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a data processing system and method, and an electronic device, a storage medium and a program product. In the present application, the data processing system comprises a terminal device and a computing cluster. The computing cluster comprises at least two layers-for-core (LCs), at least two layer-for-access (LA) groups associated with each LC, and a plurality of cloud computing devices associated with each LA group, wherein each LA group comprises at least two LAs, a plurality of cloud computing devices associated with a same LA group communicate with each other on the basis of the LAs in the LA group, and pluralities of cloud computing devices associated with different LA groups communicate with each other on the basis of the LAs of the corresponding LA groups and the associated LCs. The terminal device is used for sending a first acquisition request to at least two cloud computing devices in the computing cluster; receiving LA group identifier information returned by the cloud computing devices; and sorting the at least two cloud computing devices according to the LA group identifier information, and according to the order of the at least two cloud computing devices after sorting, using the at least two cloud computing devices for computing.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing systems, methods, electronic devices, storage media and program products

[0001] This application claims priority to Chinese Patent Application No. 202410781428.8, filed on June 17, 2024, entitled “Data Processing Method, Apparatus, Electronic Device, Storage Medium and Program Product”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the fields of cloud computing, big data and other technologies, and specifically to a data processing system, method, electronic device, storage medium and program product.

[0003] Background of the Invention

[0004] High-performance computing (HPC) clusters use high-performance cloud computing devices as nodes, interconnected via RDMA (Remote Direct Memory Access) technology, to provide high-bandwidth and extremely low-latency network services. Users can utilize HPC clusters for large-scale high-performance computing, artificial intelligence, big data recommendation, and other parallel computing applications.

[0005] In related technologies, cloud administrators can manually select a certain number of cloud computing devices and related resources based on user needs, and then orchestrate these resources to provide computing services for users. Cloud resource orchestration refers to the combination, allocation, and management of various computing resources in the cloud environment to achieve efficient resource utilization and smooth business operation. Cloud resource orchestration can include resource allocation, resource scheduling, network topology orchestration, and resource management, etc. Summary of the Invention

[0006] This application provides a data processing system, method, apparatus, electronic device, storage medium, and program product that can reduce network latency in communication between orchestrated cloud computing devices. The technical solution is as follows:

[0007] This application provides a data processing system, which includes terminal devices and a computing cluster; the computing cluster includes at least two aggregation devices (LC), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group; wherein, each LA group includes at least two LAs, and the multiple cloud computing devices associated with the same LA group communicate with each other based on the LAs of that LA group, and the multiple cloud computing devices associated with different LA groups communicate with each other based on the LAs of the corresponding LA group and the associated LCs;

[0008] The terminal device is used to send a first acquisition request to at least two cloud computing devices in the computing cluster, wherein the first acquisition request is used to request the acquisition of LA group identification information corresponding to the cloud computing device;

[0009] Receive LA group identification information returned by the at least two cloud computing devices, wherein the LA group identification information of the cloud computing devices is the representation information of the identification of the LA group associated with the cloud computing devices;

[0010] The at least two cloud computing devices are sorted according to the LA group identification information, and then used for computation in the sorted order.

[0011] This application also provides a data processing method, executed by a terminal device, including:

[0012] A first acquisition request is sent to at least two cloud computing devices in the computing cluster. The first acquisition request is used to request the acquisition of the access device (LA) group identification information corresponding to the cloud computing device. The computing cluster includes at least two aggregation devices (LC), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes at least two LAs. Multiple cloud computing devices associated with the same LA group communicate based on the LAs of the LA group. Multiple cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LC.

[0013] Receive LA group identification information returned by the at least two cloud computing devices, wherein the LA group identification information of the cloud computing devices is determined based on the identification of the LA group associated with the cloud computing devices;

[0014] The at least two cloud computing devices are sorted according to the LA group identification information, and the at least two cloud computing devices are used for computing in the sorted order. This application also provides a data processing method executed by a cloud management device of a computing cluster. The computing cluster includes at least two LCs, at least two LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes at least two LAs. Multiple cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, and multiple cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs.

[0015] The method includes:

[0016] Based on the association between the device identifier of the pre-configured cloud computing device and the identifier of the LA group, determine the identifier of the LA group associated with the cloud computing device;

[0017] Based on the identifier of the LA group, the LA group identifier information associated with the cloud computing device is determined, and the LA group identifier information is provided to the cloud computing device; wherein, the LA group identifier information is the representation information of the identifier of the LA group associated with the cloud computing device, which is used by the terminal device to sort at least two cloud computing devices to determine the order in which the terminal device uses the at least two cloud computing devices for calculation.

[0018] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described data processing method.

[0019] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described data processing method.

[0020] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described data processing method.

[0021] The technical solution provided in this application sorts each cloud computing device according to the LA group identifier information of the cloud computing devices, and uses each cloud computing device for computation in the sorted order. This arranges computing tasks as much as possible within the same LA group, greatly reducing communication between cloud computing devices across LA groups, minimizing transmission resource consumption, reducing network latency, and improving computing efficiency. Furthermore, instead of providing the original LA group identifier, the terminal device receives representation information of the LA group identifier, which helps ensure the information security of the computing cluster.

[0022] Brief description of the attached figures

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0024] Figure 1 is a schematic diagram of the implementation environment of a data processing method provided in an embodiment of this application;

[0025] Figure 2 is a schematic diagram of signaling interaction of a data processing method provided in an embodiment of this application;

[0026] Figure 3 is a schematic diagram of a computing cluster system architecture provided in an embodiment of this application;

[0027] Figure 4 is a schematic diagram of a data processing flow provided in an embodiment of this application;

[0028] Figure 5 is a schematic diagram of an allocation process provided in an embodiment of this application;

[0029] Figure 6 is a schematic diagram of an allocation process provided in an embodiment of this application;

[0030] Figure 7 is a schematic diagram of a data processing flow provided in an embodiment of this application;

[0031] Figure 8 is a schematic diagram of a computing cluster system architecture provided in an embodiment of this application;

[0032] Figure 9 is a schematic diagram of a computing cluster system architecture provided in an embodiment of this application;

[0033] Figure 10 is a schematic diagram of the structure of a data processing system provided in an embodiment of this application;

[0034] Figure 11 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0035] Figure 12 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0036] Figure 13 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0037] Figure 14 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0038] Methods of implementing the present invention

[0039] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0040] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.

[0041] It is understood that, in the specific embodiments of this application, any user-related data, such as LA group identification information, cloud computing device identification, user computing clusters associated with user clients, or associated cloud computing devices, requires user permission or consent when applied to specific products or technologies. Furthermore, the collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any user-related data is involved in the embodiments of this application, such data must be obtained with the user's authorization and consent, and in accordance with the relevant laws, regulations, and standards of the country and region.

[0042] The following section introduces and explains the terminology and related technologies involved in this application.

[0043] The computing cluster comprises multiple aggregation devices (Layer for core, LC), multiple access device (Layer for access, LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate via the LAs of that LA group, while cloud computing devices associated with different LA groups communicate via the LAs of the corresponding LA group and the associated LCs.

[0044] For example, a computing cluster can be an HCC (Hyper Computing Cluster), in which various cloud computing devices can communicate with each other via RDMA (Remote Direct Memory Access); providing users with high-bandwidth and extremely low-latency network services to meet the large-scale, high-performance parallel computing needs of related businesses.

[0045] LA Group: An LA group includes multiple LAs. In a computing cluster, LAs in the same LA group can connect to at least two cloud computing devices, so that cloud computing devices in the same LA group can communicate with each other based solely on LAs via RDMA.

[0046] LC: Used to aggregate traffic from multiple LA groups, enabling RDMA communication between various cloud computing devices across LA groups. It connects the various access devices within an LA group to facilitate inter-device communication; for example, an LC can be a switch.

[0047] Cloud computing equipment: Cloud servers used for parallel computing in a computing cluster. For example, cloud computing equipment can be a GPU server in a high-performance computing cluster.

[0048] Various embodiments provide a data processing system. The system may include terminal devices and a computing cluster. The computing cluster includes at least two aggregation devices (LCs), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes at least two LAs. Multiple cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, and multiple cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and their associated LCs.

[0049] In this system, a terminal device can be used to send a first acquisition request to at least two cloud computing devices in a computing cluster (e.g., cloud computing devices in a user computing cluster corresponding to the user of the terminal device), wherein the first acquisition request is used to request the acquisition of access device (LA) group identification information corresponding to the cloud computing device; receive LA group identification information returned by the at least two cloud computing devices, wherein the LA group identification information of the cloud computing device is the representation information of the identification of the LA group associated with the cloud computing device; sort the at least two cloud computing devices according to the LA group identification information, and use the at least two cloud computing devices for computation in the sorted order.

[0050] In this way, the terminal device sorts the cloud computing devices according to their LA group identifiers and uses these devices sequentially for computation. This allocates computational tasks to cloud computing devices within the same LA group as much as possible, thereby reducing cross-LA group communication between cloud computing devices, minimizing transmission resource consumption, reducing network latency, and improving computational efficiency. Furthermore, instead of providing the original LA group identifier, the terminal device receives a representation of the LA group identifier, which helps ensure the information security of the computing cluster.

[0051] In various embodiments, the data processing system may further include a cloud management device. The cloud management device is used to transform the identifier of the LA group associated with the cloud computing device to obtain the LA group identifier information of the cloud computing device, and then provide the LA group identifier information to the cloud computing device.

[0052] In various embodiments, the terminal device may also send an allocation request to the cloud management device. This allocation request is used to request the allocation of a cloud computing device for the terminal device. The terminal device may receive device identifiers of at least two cloud computing devices returned by the cloud management device, and determine the at least two cloud computing devices based on the device identifiers. The cloud management device may determine the number of idle cloud computing devices associated with each LA group; based on the determined number of idle devices, determine the order of at least two LA groups, and sequentially select the at least two cloud computing devices from the cloud computing devices associated with at least one of the at least two LA groups according to the determined order of the at least two LA groups, and provide the device identifiers of the at least two cloud computing devices to the terminal device.

[0053] In this way, selecting cloud computing devices from each LA group according to the number of available devices helps optimize the use of cloud resources.

[0054] In each embodiment, when the allocation request is used to request the creation of a user computing cluster, the cloud management device can sort the at least two LA groups in descending order based on the number of idle devices in the at least two LA groups to obtain a first LA group sequence; and select at least two cloud computing devices from the cloud computing devices associated with at least one LA group in the order of the at least two LA groups in the first LA group sequence to build the user computing cluster.

[0055] In this way, when creating user computing clusters, cloud computing devices are selected first from LA groups with more idle devices, thereby distributing each user computing cluster as evenly as possible to different LA groups and improving the processing efficiency of each user computing cluster.

[0056] In each embodiment, when there are at least two LA groups with the same number of idle devices in the first LA group sequence, the cloud management device can sort the LA groups in the first LA group sequence in ascending order based on the total number of user computing clusters to which the cloud computing devices in each LA group belong, to obtain the second LA group sequence; and select the cloud computing devices from the cloud computing devices associated with at least one of the at least two LA groups in the order of the LA groups in the second LA group sequence.

[0057] In this way, when creating a user computing cluster, for LA groups with the same number of idle devices, cloud computing devices are selected first from the LA group with the fewest user computing clusters served, so as to distribute each user computing cluster as evenly as possible to different LA groups and improve the processing efficiency of each user computing cluster.

[0058] In each embodiment, the cloud management device can search for an idle LA group in which all cloud computing devices are idle among the above at least two LA groups; when there is at least one idle LA group, cloud computing devices are selected from the cloud computing devices associated with the at least one idle LA group to build a user computing cluster.

[0059] In this way, when creating a user computing cluster, cloud computing devices are selected first from the idle LA group, thereby distributing each user computing cluster as evenly as possible to different LA groups and improving the processing efficiency of each user computing cluster.

[0060] In each embodiment, when the allocation request is used to request the addition of cloud computing devices to the user computing cluster, the cloud management device can sort at least one LA group to which the cloud computing devices in the user computing cluster belong in ascending order based on the number of idle devices in each LA group to obtain a third LA group sequence; and select cloud computing devices from the at least one LA group in the third LA group sequence according to the order of the LA groups to add them to the user computing cluster.

[0061] In this way, when expanding the capacity of a user computing cluster, cloud computing devices are selected first from the LA group of that user computing cluster, thereby concentrating the same user computing cluster into the same LA group as much as possible and improving the processing efficiency of the user computing cluster.

[0062] In each embodiment, when the selected cloud computing device cannot meet the requirements of the allocation request, the cloud management device can sort at least two LA groups in descending order based on the number of idle devices to obtain a first LA group sequence; according to the sorting order of the LA groups in the first LA group sequence, at least one cloud computing device is selected from the cloud computing devices associated with at least one of the at least two LA groups.

[0063] Figure 1 is a schematic diagram of the implementation environment of a data processing method provided in this application. As shown in Figure 1, the implementation environment includes: a cloud computing device 101, a terminal device 102, and a cloud management device 103.

[0064] The cloud computing device 101 can be a computing device in a computing cluster, and users can use the cloud computing device 101 to perform calculations.

[0065] For example, the computing cluster can be an HCC high-performance computing cluster, and the various cloud computing devices can communicate with each other via RDMA to provide users with high-bandwidth and extremely low-latency network services to meet the large-scale, high-performance parallel computing needs of related businesses.

[0066] In one possible scenario example, in the field of artificial intelligence, cloud computing equipment can be a GPU server in an HCC computing cluster. The HCC computing cluster can be used for large-scale AI training, which can meet the business's computing requirements for high computing performance, high stability, and high real-time performance of GPU servers.

[0067] Another possible scenario is in the field of industrial simulation, such as the automotive industry. HCC computing clusters can be used to perform simulation calculations to drive design, which can quickly respond to the real-time changing simulation needs of industrial manufacturing companies and promote product development in a timely manner.

[0068] For example, the terminal device 102, also known as the client, is a device that runs an application client for the computing cluster. The cloud management device 103 is the backend management terminal for the computing cluster. Users can interact with the cloud computing device 101 and the cloud management device 103 through the client 102.

[0069] For example, a user can trigger an allocation request for cloud computing devices on terminal device 102. For instance, the allocation request could be a creation request for establishing a new user computing cluster, or it could be an expansion request for scaling up an existing user computing cluster. Terminal device 102 can send the allocation request to cloud management device 103. Based on the allocation request, cloud management device 103 can allocate cloud computing devices to terminal device 102 from the computing cluster, enabling terminal device 102 to use the allocated cloud computing devices for relevant business computations.

[0070] In this embodiment, cloud computing devices within the same LA group only need to communicate based on LAs within that LA group. However, cloud computing devices in different LA groups need to communicate across LA groups, specifically based on LAs from at least two different LA groups and LCs associated with those at least two LA groups. Therefore, compared to communication between cloud computing devices within the same LA group, communication between cloud computing devices across LA groups significantly increases network latency and consumes more network transmission resources. In other words, communication within the same LA group is more efficient, faster, and requires fewer network transmission resources than communication across LA groups.

[0071] In this embodiment, the cloud computing device 101 may send a second acquisition request to the cloud management device 103. The second acquisition request is used to acquire the LA group identification information of the cloud computing device 101. Based on the second acquisition request, the cloud management device 103 may determine the identifier of the LA group associated with the cloud computing device 101 from the pre-stored association relationship between the device identifiers of each cloud computing device 101 and the identifiers of the LA groups, generate LA group identification information based on the determined LA group identifier, and return the LA group identification information to the cloud computing device 101.

[0072] After the cloud management device 103 allocates a cloud computing device 101 to the terminal device 102, the terminal device 102 becomes associated with the allocated cloud computing device 101. The terminal device 102 can send a first acquisition request to each associated cloud computing device 101, and each cloud computing device 101 returns its respective LA group identification information to the terminal device 102 based on the first acquisition request. The terminal device 102 can sort the associated cloud computing devices based on the LA group identification information, so as to use the associated cloud computing devices for calculation in the sorted order.

[0073] It should be noted that in this embodiment, among the cloud computing devices sorted according to LA group identifier information, those belonging to the same LA group have adjacent identifiers, thus grouping cloud computing devices in the same LA group together. Based on this, when a client uses cloud computing devices for computation in this order, it can significantly reduce RDMA communication across LA groups between cloud computing devices, minimize traffic routing, reduce network latency, and reduce traffic consumption.

[0074] It should be noted that the cloud computing device 101 or the cloud management device 103 can be a server in the cloud. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, and big data and artificial intelligence platforms.

[0075] The terminal device 102 is a user-side management platform, which users can use on the terminal device 102. For example, users can log in to their accounts on the terminal device 102 and use the computing cluster services. The terminal device 102 can be a web-based management platform, such as a management platform accessed through a webpage in a browser, or it can be a standalone application or a program plugin within a standalone application; this application does not limit the specific form of the terminal device 102. The terminal can be a smartphone, tablet computer, laptop computer, digital radio receiver, desktop computer, in-vehicle terminal (e.g., in-vehicle navigation terminal, in-vehicle computer, etc.), smart speaker, smartwatch, etc.

[0076] Figure 2 is a schematic diagram of a data processing method provided in an embodiment of this application. This data processing method can be executed interactively between terminal devices of a computing cluster, cloud management devices, and cloud computing devices. As shown in Figure 2, the data processing method includes the following steps.

[0077] Step 201: When the cloud management device receives a second acquisition request sent by any cloud computing device, the cloud management device determines the identifier of the LA group associated with any cloud computing device based on the association relationship between the device identifier of each cloud computing device and the identifier of the LA group in the pre-configured configuration.

[0078] The second request is used to request the LA group identification information of any cloud computing device. The LA group identification information is used to identify the LA group associated with any cloud computing device.

[0079] In this embodiment, the computing cluster includes multiple computing centers (LCs), multiple computing area (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LCs.

[0080] In this context, "each LC associated with multiple LA groups" means that each LC communicates with the LAs of multiple LA groups, and communication between the multiple LA groups connected to the LC passes through the LC. Similarly, "each LA group associated with multiple cloud computing devices" means that each LA group includes LAs that communicate with multiple cloud computing devices, and communication between the various cloud computing devices connected to the LAs passes through the LA.

[0081] As shown in Figure 3, multiple LAs within the same dashed line can form an LA group. Each LA in the LA group can connect to an LC (Limited Logic Controller) upstream and to a cloud computing device downstream. The cloud computing device can be a GPU server as shown in Figure 3, and each GPU server can include up to 8 GPU cards. GPU servers within the same LA group can communicate using RDMA technology without going through an LC. However, GPU servers in different LAs need to go through both an LC and an LA to communicate via RDMA. RDMA communication between GPU servers within the same LA group has lower network latency and consumes less bandwidth than RDMA communication between GPU servers in different LA groups.

[0082] For example, any cloud computing device can be any one of the various cloud computing devices associated with the terminal device. Once the cloud management device assigns a cloud computing device to the terminal device, the terminal device becomes associated with the assigned cloud computing device.

[0083] In each embodiment, prior to step 201, the cloud management device may proceed through the following step A1:

[0084] Step A1: The cloud management device can collect the device identifiers of each cloud computing device and the identifiers of the associated LA groups, and store the association relationship between the device identifiers of each cloud computing device and the identifiers of the LA groups.

[0085] For example, when new cloud computing devices are brought online in a computing cluster—that is, when the cloud computing devices are assigned to different LA groups in the computing cluster but have not yet been assigned to clients—the cloud management device can collect the device identifier of each cloud computing device and the identifier of its associated LA group, and store the device identifier and LA group identifier in association. For example, the device identifier and LA group identifier can be stored in the database of the cloud management device in a key-value pair.

[0086] For example, the device identifier of a cloud computing device can be its IP (Internet Protocol) address. The identifier of the LA group can be its ID (Identity Number). Of course, the cloud management device can collect other information about the associated LA group, such as the name of the LA group, the total number of cloud computing devices associated with the LA group, etc., which are not limited in this application.

[0087] In this step, each cloud computing device can send a second retrieval request to the cloud management device. For example, the cloud computing device can send a second retrieval request to the cloud management device via the Metadata interface.

[0088] For a second acquisition request from any cloud computing device, the cloud management device can, based on the device identifier of the cloud computing device, retrieve the identifier of the LA group associated with the device identifier of the cloud computing device from a pre-stored association relationship.

[0089] Step 202: The cloud management device determines the LA group identification information based on the identifier of the LA group associated with any cloud computing device, and returns the LA group identification information to the cloud computing device.

[0090] In this step, the cloud management device can transform the identifier of the LA group associated with any cloud computing device to obtain LA group identifier information. This transformation process can be used to desensitize the LA group identifier (i.e., process (security-sensitive) information into non-sensitive information). This LA group identifier information is used to indicate the identifier of the LA group associated with the cloud computing device. The LA group identifier information is also different for cloud computing devices in different LA groups.

[0091] In each embodiment, for any cloud computing device, the LA group identification information corresponding to the cloud computing device is obtained by the cloud management device through step B1. That is, the method of determining the LA group identification information based on the identifier of the LA group associated with any cloud computing device includes the following step B1.

[0092] Step B1: Transform the identifier of the LA group associated with the cloud computing device to obtain the implicit indicator value of the LA group identifier, and use the implicit indicator value as the LA group identifier information corresponding to the cloud computing device.

[0093] For example, the cloud management device can perform transformation processing on the identifier of the LA group based on a pre-configured mapping transformation algorithm to obtain the LA group identifier information. For instance, a hash algorithm can be used to calculate the hash value of the LA group's ID.

[0094] Based on this, by using the LA group identification information, the original ID of the LA group is hidden while distinguishing different LA groups. This achieves the desensitization of the real ID of the LA group in the cloud, thus ensuring the network security of the cloud computing cluster as much as possible.

[0095] Correspondingly, in the process of steps 201-202 above, on the side of any cloud computing device, before step 201, the cloud computing device may execute step C1.

[0096] Step C1: The cloud computing device sends a second acquisition request to the cloud management device of the computing cluster.

[0097] After step 202, any cloud computing device may execute step D1:

[0098] Step D1: The cloud computing device receives the LA group identification information returned by the cloud management device based on the second acquisition request.

[0099] As shown in Figure 4, during the deployment phase of cloud computing devices, such as when a GPU server is deployed, a device import process can be implemented. When the backend management device retrieves and executes the device import process, it collects the identifier of the LA group to which the cloud computing device belongs and writes the device identifier of the cloud computing device and the identifier of the LA group into the database DB in a key-value format. In Figure 4, the cloud computing device can access the backend through the database DB via the Metadata interface. After reading and transforming (e.g., encrypting) the data through the Metadata interface, the LA group identifier information, such as the hash value of the LA group identifier, is obtained; this transformed LA group identifier information is then returned to the cloud computing device.

[0100] When any cloud computing device is assigned to a terminal device, such as when a user joins the user computing cluster of that terminal device, the terminal device is associated with that cloud computing device. The terminal device can obtain the LA group identification information of that cloud computing device through the interaction process of steps 203-205. Based on this, the terminal device obtains the LA group identification information of each cloud computing device associated with the terminal device.

[0101] Step 203: The terminal device sends a first acquisition request to each cloud computing device associated with the terminal device.

[0102] The first acquisition request is used to request the LA group identification information corresponding to the cloud computing device. For example, the terminal device can send the first acquisition request to each associated cloud computing device via the SSH (Secure Shell) protocol.

[0103] In each embodiment, the various cloud computing devices associated with the terminal device may be pre-assigned to the terminal device by the cloud management device. The assignment process may include the following interactive steps E1-E5.

[0104] Step E1: The terminal device sends an allocation request to the cloud management device. This allocation request is used to request the allocation of cloud computing devices to the terminal device.

[0105] In some embodiments, the allocation request is a new request, which is used to request the creation of a user computing cluster for the terminal device.

[0106] For example, creating a user computing cluster refers to the process of allocating at least one cloud computing device to a terminal device and building the allocated cloud computing device into a user computing cluster associated with the terminal device.

[0107] For example, users can trigger a creation operation on their terminal devices, such as configuring the number of devices, cluster name, availability zone, etc., for the user computing cluster to be created. If a user does not yet have a user computing cluster, they can apply to create a new one through the creation operation; of course, if a user already has a user computing cluster, they can also create one as needed. In one possible application example, instance RDMA networks within the same user computing cluster can be interconnected, while instance RDMA networks across user computing clusters are isolated from each other.

[0108] In other embodiments, the allocation request is a scaling request, which is used to request the addition of cloud computing devices to the user computing cluster of the terminal device.

[0109] For example, if a user already has a user computing cluster, the user can also trigger an expansion operation on the terminal device to increase the number of cloud computing devices in the existing user computing cluster.

[0110] Step E2: The cloud management device receives the allocation request from the terminal device.

[0111] After receiving the allocation request, the cloud management device can allocate cloud computing devices to the terminal device through the following steps E3-E4. That is, the various cloud computing devices associated with the terminal device are determined by the cloud management device through the following steps E3-E4.

[0112] Step E3: The cloud management device determines the number of idle cloud computing devices associated with each LA group in an idle state.

[0113] The idle state refers to the fact that the cloud computing device has not yet been assigned to any client.

[0114] Step E4: Based on the number of idle devices in each LA group, the cloud management device determines the allocation order of each LA group, and in accordance with the allocation order of each LA group, sequentially selects cloud computing devices from at least one LA group associated with cloud computing devices and allocates them to the terminal device.

[0115] In this step, based on the two possible scenarios of the allocation request, the implementation of step E4 is divided into the following two scenarios: scenario one and scenario two.

[0116] In case 1, the allocation request is a new request, and the implementation of step E4 may include steps E41-E42.

[0117] Step E41: Based on the number of idle devices in each LA group, sort each LA group in descending order to obtain the first LA group sequence.

[0118] Here, the number of idle devices refers to the number of cloud computing devices in each cloud computing device group that are in an idle state. In this step, the LA groups can be arranged in descending order of the number of idle devices to obtain the first LA group sequence. That is, in the first LA group sequence, the earlier the LA group appears in the sequence, the more idle devices it has.

[0119] Step E42: Using the order of the LA groups in the first LA group sequence as the allocation order, select cloud computing devices from the cloud computing devices associated with at least one LA group in sequence, and build a user computing cluster for the terminal device based on the selected cloud computing devices.

[0120] In this step, cloud computing devices can be selected sequentially from each LA group according to the number of available devices, from largest to smallest. For example, if the number of devices to be created in a new user computing cluster is the first number, then cloud computing devices are selected first from the LA groups with the most available devices, in descending order of the number of available devices, until the first number of cloud computing devices are selected. In this way, when creating a user computing cluster, cloud computing devices are selected first from the LA groups with the most available devices, thereby distributing the user computing clusters as evenly as possible across different LA groups and improving the processing efficiency of each user computing cluster.

[0121] In some embodiments, when the number of idle devices in an LA group is the same as the total number of cloud computing devices in that LA group, all associated cloud computing devices are in an idle state, and that LA group is an idle LA group with the most idle devices. Cloud computing devices are selected preferentially from the idle LA group.

[0122] In each embodiment, for LA groups with the same number of idle devices, cloud computing devices can be preferentially selected from LA groups with fewer corresponding clients (i.e., LA groups with fewer total user computing clusters to which the cloud computing devices belong). Accordingly, step E42, which selects cloud computing devices sequentially from at least one LA group-associated cloud computing device using the order of the LA groups in the first LA group sequence as the allocation order, includes the following steps E42-1 and E42-2:

[0123] Step E42-1: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, sort the at least two LA groups in the first LA group sequence in ascending order based on the number of user computing clusters corresponding to each of the at least two LA groups to obtain the second LA group sequence.

[0124] The number of user computing clusters (also known as the number of clients) corresponding to the LA group refers to the number of user computing clusters to which each cloud computing device in the LA group belongs.

[0125] Step E42-2: Using the order of the LA groups in the second LA group sequence as the allocation order, select cloud computing devices from at least one LA group associated cloud computing devices in sequence.

[0126] For example, in the first LA group sequence, LA group a and LA group b are in the 1st and 2nd positions respectively. Both LA group a and LA group b are associated with 100 cloud computing devices, and each has 60 idle devices. Among them, of the 40 cloud computing devices already allocated in LA group a, 3 clients have been assigned; of the 40 cloud computing devices already allocated in LA group b, 2 clients have been assigned. Therefore, in the first LA sequence, the order of LA group b is moved before LA group a. After the adjustment, LA group a is in the 2nd position, and LA group b is in the 1st position. This is to prioritize the selection of cloud computing devices for terminal devices from the 60 idle cloud computing devices associated with LA group b.

[0127] In various embodiments, cloud computing devices may first be selected from the idle LA group. Only if there is no idle LA group or insufficient devices in the idle LA group will step E41 be executed. Accordingly, step E41 may be implemented using the following steps: E41-1, E41-2, and E41-3.

[0128] Step E41-1: Locate the idle LA group in each LA group. All cloud computing devices associated with the idle LA group are in an idle state.

[0129] Step E41-2: If there are idle LA groups in each LA group, select a cloud computing device from the cloud computing devices associated with the idle LA group, and build a user computing cluster for the terminal device based on the selected cloud computing device;

[0130] Step E41-3: If there are no idle LA groups in each LA group, sort each LA group in descending order based on the number of idle devices in each LA group to obtain the first LA group sequence.

[0131] For example, if the total number of devices associated with each LA group is the same, then the idle LA group has the largest number of idle devices. In this case, the idle LA group can be searched first, and cloud computing devices can be selected from the idle LA group first. If no idle LA group exists, then the step of sorting the LA groups in descending order based on the number of idle devices is performed.

[0132] In some embodiments, if there are idle LA groups in each LA group, steps E41-2 and E41-3 can be executed based on the total number of devices in the idle LA groups and the number of devices required for the user computing cluster to be created; the number of devices required for the user computing cluster to be created is a first number. Accordingly, the process may include the following steps:

[0133] If there are idle LA groups in each LA group, and the total number of devices in the idle LA groups is not less than a first number, select a first number of cloud computing devices from the cloud computing devices associated with the idle LA groups to construct a user computing cluster including the first number of cloud computing devices;

[0134] If there are idle LA groups in each LA group, and the total number of devices in the idle LA groups is less than a first number, all cloud computing devices are selected from the cloud computing devices associated with the idle LA groups. Based on the number of idle devices in the non-idle LA groups in each LA group, the LA groups are sorted in descending order to obtain a first LA group sequence. The remaining number of cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, using the sorting order of the LA groups in the first LA group sequence as the allocation order, to construct a user computing cluster including the first number of cloud computing devices. The sum of the remaining number and the total number of devices in the idle LA groups is the first number.

[0135] If there are no idle LA groups in each LA group, the LA groups are sorted in descending order based on the number of idle devices in each LA group to obtain the first LA group sequence; and the first number of cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, based on the sorting order of each LA group in the first LA group sequence, to construct a user computing cluster including the first number of cloud computing devices.

[0136] As shown in Figure 5, a dashed box represents an LA group, and the squares within the dashed box represent cloud computing devices. Squares with different fill patterns or colors represent cloud computing devices in different user computing clusters, while blank squares without any fill pattern or color represent idle cloud computing devices. Currently, each LA group does not include idle LA groups. Two sorting processes can be performed: the first sort is in descending order of the number of idle devices, placing those with more idle devices first. The second sort is in ascending order of the number of user computing clusters, placing those with fewer user computing clusters first. Based on this, the allocation strategy for newly created scenarios can be configured as follows:

[0137] Prioritize selecting cloud computing devices from the idle LA group;

[0138] If there are no available LA groups, cloud computing devices will be selected from LA groups whose number of available devices can meet the needs of the user's computing cluster.

[0139] If there are multiple LA groups that can meet the number of devices required by the user's computing cluster, cloud computing devices shall be selected first from the LA group with the fewer user computing clusters.

[0140] If there is no LA group that meets the number of devices required for the user's computing cluster, cloud computing devices will be selected from the LA with the most idle devices.

[0141] Scenario 2: The allocation request is an expansion request. Accordingly, the implementation of step E4 may include steps E43-E44:

[0142] Step E43: Based on the number of idle devices in each associated LA group, sort each associated LA group in ascending order to obtain the third LA group sequence. The cloud computing devices associated with each associated LA group include the cloud computing devices that have been allocated to the terminal device.

[0143] Step E44: Using the order of the associated LA groups in the third LA group sequence as the allocation order, select cloud computing devices from at least one associated LA group and allocate them to the terminal device.

[0144] For example, in a capacity expansion allocation scenario, the LA groups (also known as associated LA groups) of all cloud computing devices in the user's computing cluster can be located and sorted in ascending order of the number of idle devices in the associated LA groups. Cloud computing devices are allocated preferentially from LA groups with fewer idle devices, so that as many cloud computing devices of a client as possible fall into the same LA group, that is, to try to make the cloud computing devices in the LA group belong to the same user. In this way, when expanding the user's computing cluster, cloud computing devices are preferentially selected from the LA groups of the user's computing cluster, thereby concentrating the same user's computing cluster into the same LA group as much as possible, improving the processing efficiency of the user's computing cluster.

[0145] For example, if it is necessary to expand user A's user computing cluster by 3 cloud computing devices, the LA group where user A's cloud computing devices are located includes group 1, group 2, and group 3. Group 1 has 6 idle devices, group 2 has 8 idle devices, and group 3 has 3 idle devices. The devices are arranged in ascending order of the number of idle devices from least to most: group 3 < group 1 < group 2. Therefore, 3 cloud computing devices should be selected first from group 3.

[0146] It should be noted that steps E41-E42 and steps E43-E44 are allocation steps for two different allocation scenarios. This application does not limit the order of steps E41-E42 or steps E43-E44. That is, steps E41-E42 can be executed before or after steps E43-E44.

[0147] Step E5: The cloud management device returns the device identifiers of each assigned cloud computing device to the terminal device.

[0148] In this process, the cloud management device can deduct the allocated cloud computing devices from the cloud computing devices in the computing cluster and send the device identifier of the cloud computing device allocated to the terminal device to the terminal device. Therefore, this process can also be called the deduction and packing process.

[0149] Step E6: The terminal device receives the device identifiers of each cloud computing device returned by the cloud management device, and obtains the cloud computing devices associated with the terminal device based on the device identifiers of each cloud computing device.

[0150] The allocation process of steps E1-E6 above will be described below using the flowchart shown in Figure 6.

[0151] As shown in Figure 6, the allocation process includes the following steps:

[0152] (1) The terminal device initiates an allocation request;

[0153] (2) The cloud management device can obtain the current resource status of the user of the terminal device, such as whether the terminal device has a user computing cluster or has not created a user computing cluster.

[0154] (3) The cloud management device determines whether the allocation request is a new request;

[0155] (4) If it is a new request, the cloud management device can execute the new process, which can be sorted twice: the first sort is sorted in descending order according to the number of idle devices, with the LA group with more idle devices at the front; the second sort is sorted in ascending order according to the number of user computing clusters, with the user computing clusters with fewer clusters at the front; and execute (6);

[0156] (5) For expansion requests, the cloud management device can execute the expansion process, specifically obtain the LA group resource status of the user, specifically obtain the LA group associated with the user's cloud computing device, obtain the number of idle devices in the associated LA group, and sort them: sort them in ascending order according to the number of idle devices, with those having fewer idle devices listed first; and execute (6);

[0157] (6) The cloud management device selects cloud computing devices and assigns them to terminal devices according to the sorted LA group order, which is the process of packing and deducting.

[0158] (7) The cloud management device determines whether the deduction is complete. If the deduction is complete, the binning is successful, which means the allocation is successful. If the deduction is not complete, it means that some cloud computing devices have been successfully selected. Furthermore, the corresponding process can be executed again for allocation until all bins are successfully packed.

[0159] Step 204: The cloud computing device receives the first acquisition request sent by the terminal device, and based on the first acquisition request, returns the LA group identification information of any cloud computing device to the terminal device.

[0160] The LA group identification information is determined based on the identifier of the LA group associated with the cloud computing device. In this step, the cloud computing device can return its LA group identification information to the terminal device via the SSH protocol.

[0161] In each embodiment, the terminal device may execute step 203 in response to the triggering of a target event, the target event including at least one of the following:

[0162] Receives the device identifier for the cloud computing device assigned by the cloud management device;

[0163] Receives the changed device identifier of the cloud computing device sent by the cloud management device;

[0164] The target time period has been reached.

[0165] In one possible example, when a terminal device receives a device identifier for a cloud computing device assigned by a cloud management device, the terminal device can send a first acquisition request to each cloud computing device assigned by the cloud management device. For example, when creating a new user computing cluster, the first acquisition request can be sent to the cloud computing devices in the newly created user computing cluster. Alternatively, when expanding an existing user computing cluster, the first acquisition request can be sent to the newly added cloud computing devices.

[0166] In one possible example, the cloud computing device that changes could be a cloud computing device that malfunctions. For example, if device A malfunctions, the cloud management device can return the device identifier of device A and the device identifier of its replacement device B to the terminal device. The terminal device can send a first acquisition request to the replacement device B and can also delete the LA group identifier information of device A from its local storage.

[0167] In one possible example, the terminal device can also periodically update the LA group identification information of the cloud computing devices. For example, it can periodically send a first acquisition request to the cloud computing devices at each target time period; or it can periodically detect the LA group identification information of each cloud computing device and periodically send a first acquisition request to the cloud computing devices whose LA group identification information has expired or has not been acquired.

[0168] Step 205: The terminal device receives the LA group identification information returned by each cloud computing device associated with the terminal device.

[0169] The terminal device can receive LA group identification information returned by each cloud computing device associated with it via the SSH protocol. Based on this, the terminal device can obtain the distribution of LA groups for each associated cloud computing device.

[0170] In the sending and receiving processes in steps 204 and 205, since the LA group identification information is the information after the cloud management device has transformed the LA group identification, it is the information after the transformation processing of the original LA group identification. Therefore, the security of information in the cloud network during network transmission is greatly improved, while meeting the client's need to obtain the LA group identification information in a timely manner.

[0171] Step 206: The terminal device sorts the various cloud computing devices associated with the terminal device based on the LA group identification information, so as to use the various cloud computing devices associated with the terminal device for calculation in the sorted order.

[0172] Terminal devices can arrange cloud computing devices in ascending order according to LA group identification information, such as the implicit indication value of the LA group identifier, so that cloud computing devices in the same LA group are clustered together and adjacent to each other. Based on this, the cloud computing devices can be used sequentially for relevant business calculations according to the sorted order. For example, in relevant businesses using the Ring-Allreduce algorithm, the computing nodes are arranged in a ring, and data interaction is required between the circularly arranged computing nodes during the calculation process, such as transferring intermediate calculation data. The cloud computing devices can be used in the sorted order, which helps to obtain cloud computing devices clustered together in the same LA group for Ring-Allreduce algorithm calculations. That is, it ensures that the circularly arranged computing nodes are in the same LA group, thereby greatly reducing network latency during the Ring-Allreduce algorithm calculation process, significantly saving traffic consumption, and ensuring the efficiency of the calculation process.

[0173] As shown in Figure 7, each cloud computing device can obtain the transformed LA group identification information through the Metadata interface and return the LA group identification information to the client. The client sorts the IP addresses of the cloud computing devices according to the LA group identification information, and the IP addresses of cloud computing devices with the same LA group identification information are arranged together.

[0174] As shown in Figure 8, rectangles with the same fill pattern represent devices associated with the same client. In Figure 8, GPU servers associated with the same client are distributed in the same LA group. However, over time, resource adjustments associated with the client, such as expansion, downsizing, or machine relocation due to machine failure, gradually fragment the distribution of GPU servers for the same client across different LA groups. Specifically, GPU servers for the same client may be distributed across different LA groups. Therefore, communication between GPU servers for the same client requires crossing LA groups, which may result in network latency. Through the embodiments of this application, the terminal device can arrange GPU devices according to LA group identification information, thus grouping cloud computing devices in the same LA group together. Based on this, when the client uses cloud computing devices for computation in this order, it can greatly reduce RDMA communication across LA groups between cloud computing devices, minimize traffic routing, reduce network latency, and reduce traffic consumption.

[0175] The data processing method provided in this application includes a computing cluster comprising multiple cloud computing nodes (LCs), multiple cloud computing nodes (LAs) associated with each LC, and multiple cloud computing devices associated with each LA group. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of their respective LA groups and their associated LCs. A terminal device obtains the LA group identification information of each cloud computing device associated with it by sending a first acquisition request to the associated cloud computing device. Furthermore, the cloud computing devices are sorted based on the LA group identification information so that they can be used for computation in the sorted order. Since the sorting group concentrates cloud computing devices within the same LA group, it facilitates computation using these devices, significantly reduces cross-LA group communication between cloud computing devices, minimizes traffic routing, reduces network latency, and reduces traffic consumption.

[0176] Figure 10 is a schematic diagram of the structure of a data processing system provided in an embodiment of this application. As shown in Figure 10, the system includes: a terminal device 1001 of a computing cluster, a cloud management device 1002, and any cloud computing device 1003 associated with the terminal device.

[0177] The computing cluster includes multiple aggregation devices (LCs), multiple access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate with each other based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate with each other based on the LAs of the corresponding LA group and the associated LCs.

[0178] The cloud management device 1002 is used to determine the identifier of the LA group associated with any cloud computing device when it receives a second acquisition request sent by any cloud computing device, based on the association relationship between the device identifier of each cloud computing device and the identifier of the LA group.

[0179] The cloud management device 1002 is used to determine the LA group identification information based on the identifier of the LA group associated with any cloud computing device, and return the LA group identification information to the cloud computing device.

[0180] The terminal device 1001 is used to send a first acquisition request to each cloud computing device associated with the terminal device. The first acquisition request is used to request the acquisition of LA group identification information corresponding to the cloud computing device.

[0181] The cloud computing device 1003 is configured to receive a first acquisition request sent by a terminal device, and based on the first acquisition request, return the LA group identification information of the cloud computing device to the terminal device. The LA group identification information is determined based on the identification of the LA group associated with the cloud computing device.

[0182] The terminal device 1001 is used to receive LA group identification information returned by each cloud computing device associated with the terminal device;

[0183] The terminal device 1001 is used to sort the various cloud computing devices associated with the terminal device based on the LA group identification information, so as to use the various cloud computing devices associated with the terminal device for calculation in the sorted order.

[0184] The data processing method provided in this application includes a computing cluster comprising multiple cloud computing nodes (LCs), multiple cloud computing nodes (LAs) associated with each LC, and multiple cloud computing devices associated with each LA group. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of their respective LA groups and their associated LCs. A terminal device obtains the LA group identification information of each cloud computing device associated with it by sending a first acquisition request to the associated cloud computing device. Furthermore, the cloud computing devices are sorted based on the LA group identification information so that they can be used for computation in the sorted order. Since the sorting group concentrates cloud computing devices within the same LA group, it facilitates computation using these devices, significantly reduces cross-LA group communication between cloud computing devices, minimizes traffic routing, reduces network latency, and reduces traffic consumption.

[0185] Figure 11 is a schematic diagram of a data processing device provided in an embodiment of this application. The device is applied to a terminal device, which is an application client of a computing cluster. The computing cluster includes multiple aggregation devices (LCs), multiple access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes multiple LAs; cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, and cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs. As shown in Figure 11, the device includes:

[0186] The first sending module 1101 is used to send a first acquisition request to each cloud computing device associated with the terminal device. The first acquisition request is used to request the acquisition of LA group identification information corresponding to the cloud computing device.

[0187] The first receiving module 1102 is used to receive LA group identification information returned by each cloud computing device associated with the terminal device. The LA group identification information is determined based on the identification of the LA group associated with the cloud computing device.

[0188] The sorting module 1103 is used to sort the various cloud computing devices associated with the terminal device based on the LA group identification information, so as to use the various cloud computing devices associated with the terminal device for calculation in the sorted order.

[0189] In each embodiment, the data processing device applied to the cloud management device of the computing cluster includes an LA group identification information determination module. For any cloud computing device, the LA group identification information determination module, when determining the LA group identification information corresponding to the cloud computing device, is specifically used for:

[0190] The identifier of the LA group associated with the cloud computing device is transformed to obtain the implicit indication value of the LA group identifier, and this implicit indication value is used as the LA group identifier information corresponding to the cloud computing device.

[0191] In each embodiment, the device further includes:

[0192] The second sending module is used to send an allocation request to the cloud management device of the computing cluster. The allocation request is used to request the allocation of cloud computing devices to the terminal devices.

[0193] The third receiving module is used to receive the device identifiers of each cloud computing device returned by the cloud management device, and obtain the various cloud computing devices associated with the terminal device based on the device identifiers of each cloud computing device.

[0194] The data processing device applied to the cloud management device includes a cloud computing device determination module, which, when determining the various cloud computing devices associated with the terminal device, includes:

[0195] The first determining unit is used to determine the number of idle cloud computing devices associated with each LA group in an idle state;

[0196] The allocation unit is used to determine the allocation order of each LA group based on the number of idle devices in each LA group, and to select cloud computing devices from at least one LA group associated with cloud computing devices and allocate them to the terminal device in sequence according to the allocation order of each LA group.

[0197] In each embodiment, the allocation request is a new creation request, which is used to request the creation of a new user computing cluster for the terminal device:

[0198] This allocation unit is used for:

[0199] Based on the number of idle devices in each LA group, the LA groups are sorted in descending order to obtain the first LA group sequence;

[0200] The cloud computing devices associated with at least one LA group are selected sequentially from the cloud computing devices associated with at least one LA group, and a user computing cluster is built for the terminal device based on the selected cloud computing devices, using the order of arrangement of each LA group in the first LA group sequence as the allocation order.

[0201] In each embodiment, when the allocation unit selects cloud computing devices sequentially from at least one cloud computing device associated with an LA group, using the order of the LA groups in the first LA group sequence as the allocation order, it is specifically used for:

[0202] If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are sorted in ascending order based on the number of user computing clusters corresponding to each of the at least two LA groups to obtain the second LA group sequence. The number of user computing clusters refers to the number of clients associated with each cloud computing device associated with the LA group.

[0203] The cloud computing devices are selected sequentially from at least one cloud computing device associated with an LA group, based on the order of arrangement of each LA group in the second LA group sequence.

[0204] In each embodiment, when the allocation unit sorts the LA groups in descending order based on the number of idle devices in each LA group to obtain the first LA group sequence, it is specifically used for:

[0205] Find the idle LA group in each LA group, where all cloud computing devices associated with the idle LA group are in an idle state;

[0206] If there are idle LA groups in each LA group, select a cloud computing device from the cloud computing devices associated with that idle LA group, and build a user computing cluster for the terminal device based on the selected cloud computing device;

[0207] If there are no idle LA groups in each LA group, the LA groups are sorted in descending order based on the number of idle devices in each LA group to obtain the first LA group sequence.

[0208] In each embodiment, the allocation request is a scaling request, which is used to request the addition of cloud computing devices to the user computing cluster of the terminal device:

[0209] This allocation unit is used for:

[0210] Based on the number of idle devices in each associated LA group, the associated LA groups are sorted in ascending order to obtain the third LA group sequence. The cloud computing devices associated with each associated LA group include the cloud computing devices that have been allocated to the terminal device.

[0211] The cloud computing devices are selected from at least one associated LA group and allocated to the terminal device in sequence, based on the order of arrangement of each associated LA group in the third LA group sequence.

[0212] The data processing method provided in this application includes a computing cluster comprising multiple cloud computing nodes (LCs), multiple cloud computing nodes (LAs) associated with each LC, and multiple cloud computing devices associated with each LA group. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of their respective LA groups and their associated LCs. A terminal device obtains the LA group identification information of each cloud computing device associated with it by sending a first acquisition request to the associated cloud computing device. Furthermore, the cloud computing devices are sorted based on the LA group identification information so that they can be used for computation in the sorted order. Since the sorting group concentrates cloud computing devices within the same LA group, it facilitates computation using these devices, significantly reduces cross-LA group communication between cloud computing devices, minimizes traffic routing, reduces network latency, and reduces traffic consumption.

[0213] Figure 12 is a schematic diagram of a data processing device provided in an embodiment of this application. This device is applied to any cloud computing device associated with a terminal device in a computing cluster, the computing cluster including multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0214] Each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LCs.

[0215] As shown in Figure 12, the device includes:

[0216] The second receiving module 1201 is used to receive a first acquisition request sent by the terminal device, the first acquisition request being used to request the acquisition of LA group identification information corresponding to any cloud computing device;

[0217] The first return module 1202 is used to return LA group identification information of any cloud computing device to the terminal device based on the first acquisition request. The LA group identification information is determined based on the identifier of the LA group associated with the cloud computing device, so that the terminal device can sort the various cloud computing devices associated with the terminal device based on the LA group identification information, and use the various cloud computing devices associated with the terminal device for calculation in the sorted order.

[0218] In each embodiment, the device further includes:

[0219] The third sending module is used to send a second acquisition request to the cloud management device of the computing cluster. The second acquisition request is used to request the acquisition of the LA group identification information of any cloud computing device.

[0220] The fourth receiving module is used to receive the LA group identification information returned by the cloud management device based on the second acquisition request.

[0221] The data processing method provided in this application includes a computing cluster comprising multiple cloud computing nodes (LCs), multiple cloud computing nodes (LAs) associated with each LC, and multiple cloud computing devices associated with each LA group. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of their respective LA groups and their associated LCs. A terminal device obtains the LA group identification information of each cloud computing device associated with it by sending a first acquisition request to the associated cloud computing device. Furthermore, the cloud computing devices are sorted based on the LA group identification information so that they can be used for computation in the sorted order. Since the sorting group concentrates cloud computing devices within the same LA group, it facilitates computation using these devices, significantly reduces cross-LA group communication between cloud computing devices, minimizes traffic routing, reduces network latency, and reduces traffic consumption.

[0222] Figure 13 is a schematic diagram of a data processing device provided in an embodiment of this application. This application provides a cloud management device for a computing cluster, the computing cluster including multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0223] Each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LCs.

[0224] As shown in Figure 13, the device includes:

[0225] The first determining module 1301 is used to determine the identifier of the LA group associated with any cloud computing device based on the association relationship between the device identifier of each cloud computing device and the identifier of the LA group when a second acquisition request is received from any cloud computing device.

[0226] The second determining module 1302 is configured to determine LA group identification information based on the identifier of the LA group associated with any cloud computing device, and return the LA group identification information to any cloud computing device, so that the cloud computing device performs the following steps:

[0227] Upon receiving the first acquisition request sent by the associated target client segment, the LA group identification information of any cloud computing device is returned to the terminal device, so that the terminal device can sort the various cloud computing devices associated with the terminal device based on the LA group identification information, and use the various cloud computing devices associated with the terminal device for calculation in the sorted order.

[0228] In each embodiment, when determining the LA group identifier information based on the identifier of the LA group associated with any cloud computing device, the second determining module is specifically used for:

[0229] The identifier of the LA group associated with the cloud computing device is transformed to obtain the implicit indication value of the LA group identifier, and this implicit indication value is used as the LA group identifier information corresponding to the cloud computing device.

[0230] In each embodiment, the device further includes:

[0231] The idle device number determination module is used to determine the number of idle cloud computing devices associated with each LA group when receiving an allocation request from a terminal device. The allocation request is used to request the allocation of cloud computing devices to the terminal device.

[0232] The allocation module is used to determine the allocation order of each LA group based on the number of idle devices in each LA group, and to select cloud computing devices from at least one LA group associated with cloud computing devices and allocate them to the terminal device in the order of allocation of each LA group.

[0233] The second return module is used to return the device identifiers of the various cloud computing devices assigned to the terminal device.

[0234] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0235] The data processing method provided in this application includes a computing cluster comprising multiple cloud computing nodes (LCs), multiple cloud computing nodes (LAs) associated with each LC, and multiple cloud computing devices associated with each LA group. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, while cloud computing devices associated with different LA groups communicate based on the LAs of their respective LA groups and their associated LCs. A terminal device obtains the LA group identification information of each cloud computing device associated with it by sending a first acquisition request to the associated cloud computing device. Furthermore, the cloud computing devices are sorted based on the LA group identification information so that they can be used for computation in the sorted order. Since the sorting group concentrates cloud computing devices within the same LA group, it facilitates computation using these devices, significantly reduces cross-LA group communication between cloud computing devices, minimizes traffic routing, reduces network latency, and reduces traffic consumption.

[0236] Figure 14 is a schematic diagram of an electronic device provided in an embodiment of this application. As shown in Figure 14, the electronic device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the method described above.

[0237] In one embodiment, an electronic device is provided, as shown in FIG14. The electronic device 1400 shown in FIG14 includes a processor 1401 and a memory 1403. The processor 1401 and the memory 1403 are connected, for example, via a bus 1402. In some embodiments, the electronic device 1400 may further include a transceiver 1404, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 1404 is not limited to one type, and the structure of the electronic device 1400 does not constitute a limitation on the embodiments of this application.

[0238] Processor 1401 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1401 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0239] Bus 1402 may include a pathway for transmitting information between the aforementioned components. Bus 1402 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 1402 may be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 14, but this does not indicate that there is only one bus or one type of bus.

[0240] The memory 1403 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0241] The memory 1403 is used to store computer programs that execute the embodiments of this application, and the execution is controlled by the processor 1401. The processor 1401 is used to execute the computer programs stored in the memory 1403 to implement the steps shown in the foregoing method embodiments.

[0242] Electronic devices include, but are not limited to, servers, terminals, or cloud computing devices.

[0243] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.

[0244] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0245] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0246] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.

[0247] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.

[0248] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0249] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0250] The above description is only one implementation method of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application are also within the protection scope of the embodiments of this application.

Claims

1. A data processing system, the system comprising terminal devices and a computing cluster; the computing cluster comprising at least two aggregation devices (LCs), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group; wherein, Each LA group includes at least two LAs. Multiple cloud computing devices associated with the same LA group communicate based on the LAs of that LA group. Multiple cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LCs. The terminal device is used to send a first acquisition request to at least two cloud computing devices in the computing cluster, wherein the first acquisition request is used to request the acquisition of the access device (LA) group identification information corresponding to the cloud computing device; Receive LA group identification information returned by the at least two cloud computing devices, wherein the LA group identification information of the cloud computing devices is the representation information of the identification of the LA group associated with the cloud computing devices; The at least two cloud computing devices are sorted according to the LA group identification information, and the at least two cloud computing devices are used for calculation in the sorted order.

2. The method according to claim 1, wherein, Further includes: cloud-managed devices; The cloud management device is used to transform the identifier of the LA group associated with the cloud computing device to obtain the LA group identifier information of the cloud computing device, and to provide the LA group identifier information of the cloud computing device to the cloud computing device.

3. The system according to claim 1, further comprising a cloud management device; The terminal device is used to send an allocation request to the cloud management device, the allocation request being used to request the allocation of a cloud computing device to the terminal device; Receive device identifiers of at least two cloud computing devices returned by the cloud management device, and determine the at least two cloud computing devices based on the device identifiers; The cloud management device is used to determine the number of idle cloud computing devices associated with each LA group; based on the number of idle devices, determine the order of the at least two LA groups, and in accordance with the order of the at least two LA groups, sequentially select the at least two cloud computing devices from the cloud computing devices associated with at least one of the at least two LA groups, and provide the device identifiers of the at least two cloud computing devices to the terminal device.

4. The system according to claim 3, wherein, The allocation request is used to request the creation of a user computing cluster: The cloud management device is used for, Based on the number of idle devices in the at least two LA groups, the at least two LA groups are sorted in descending order to obtain the first LA group sequence; In accordance with the arrangement order of the at least two LA groups in the first LA group sequence, at least two cloud computing devices are selected sequentially from the cloud computing devices associated with at least one of the at least two LA groups to construct the user computing cluster.

5. The system according to claim 4, wherein, The cloud management device is used for: When there are at least two LA groups with the same number of idle devices in the first LA group sequence, the LA groups in the first LA group sequence are sorted in ascending order based on the total number of user computing clusters to which the cloud computing devices in each LA group belong, so as to obtain the number of user computing clusters in the second LA group sequence. According to the order of the LA groups in the second LA group sequence, at least two cloud computing devices are selected sequentially from at least one LA group associated cloud computing device in the at least two LA groups.

6. The system according to claim 4 or 5, wherein, The cloud management device is used for: Among the at least two LA groups, find the idle LA group in which all cloud computing devices are in an idle state; When at least one idle LA group exists, a cloud computing device is selected from the cloud computing devices associated with the at least one idle LA group to build the user computing cluster.

7. The system according to claim 3, wherein, The allocation request is used to request the addition of cloud computing devices to the user's computing cluster. The cloud management device is used for: Based on the number of idle devices, at least one LA group to which the cloud computing devices in the user computing cluster belong is sorted in ascending order to obtain a third LA group sequence; According to the order of the LA groups in the third LA group sequence, cloud computing devices are selected from at least one LA group and added to the user computing cluster in sequence.

8. A data processing method, executed by a terminal device, comprising: A first acquisition request is sent to at least two cloud computing devices in the computing cluster. The first acquisition request is used to request the acquisition of the access device (LA) group identification information corresponding to the cloud computing device. The computing cluster includes at least two aggregation devices (LC), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes at least two LAs. Multiple cloud computing devices associated with the same LA group communicate based on the LAs of the LA group. Multiple cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LC. Receive LA group identification information returned by the at least two cloud computing devices, wherein the LA group identification information of the cloud computing devices is determined based on the identification of the LA group associated with the cloud computing devices; The at least two cloud computing devices are sorted according to the LA group identification information, and the at least two cloud computing devices are used for calculation in the sorted order.

9. The method of claim 8, further comprising: Send an allocation request to the cloud management device, the allocation request being used to request the allocation of a cloud computing device to the terminal device; The system receives device identifiers of at least two cloud computing devices returned by the cloud management device, and determines the at least two cloud computing devices based on the device identifiers.

10. A data processing method, executed by a cloud management device of a computing cluster, wherein, The computing cluster includes at least two aggregation devices (LC), at least two access device (LA) groups associated with each LC, and multiple cloud computing devices associated with each LA group; wherein, each LA group includes at least two LAs, and the multiple cloud computing devices associated with the same LA group communicate with each other based on the LAs of that LA group, and the multiple cloud computing devices associated with different LA groups communicate with each other based on the LAs of the corresponding LA group and the associated LCs. The method includes: Based on the association between the device identifier of the pre-configured cloud computing device and the identifier of the LA group, determine the identifier of the LA group associated with the cloud computing device; Based on the identifier of the LA group, the LA group identifier information associated with the cloud computing device is determined, and the LA group identifier information is provided to the cloud computing device; wherein, the LA group identifier information is the representation information of the identifier of the LA group associated with the cloud computing device, which is used by the terminal device to sort at least two cloud computing devices to determine the order in which the terminal device uses the at least two cloud computing devices for calculation.

11. The method according to claim 10, wherein, The LA group identifier information associated with the cloud computing device includes: The LA group identifier information is obtained by transforming the identifier of the LA group associated with the cloud computing device.

12. The method of claim 10, further comprising: When a request to allocate cloud computing devices is received from the terminal device, the number of idle cloud computing devices associated with each LA group in an idle state is determined. Based on the number of idle devices, the order of the at least two LA groups is determined, and in accordance with the order, at least two cloud computing devices are selected sequentially from the cloud computing devices associated with at least one of the at least two LA groups; Provide the terminal device with the device identifiers of the at least two cloud computing devices.

13. The method according to claim 12, wherein, The allocation request is used to request the creation of a user computing cluster: The selection of at least two cloud computing devices includes: The first LA group sequence is obtained by sorting the at least two LA groups in descending order based on the number of idle devices; In accordance with the order of the LA groups in the first LA group sequence, at least two cloud computing devices are selected sequentially from at least one of the LA groups associated with cloud computing devices to construct the user computing cluster.

14. The method according to claim 13, wherein, Selecting the at least two cloud computing devices sequentially from at least one cloud computing device associated with one of the at least two LA groups includes: When there are at least two LA groups with the same number of idle devices in the first LA group sequence, the LA groups in the first LA group sequence are sorted in ascending order based on the total number of user computing clusters to which the cloud computing devices in each LA group belong, thereby obtaining the second LA group sequence; In accordance with the order of the LA groups in the second LA group sequence, at least two cloud computing devices are selected sequentially from at least one cloud computing device associated with at least one of the at least two LA groups.

15. The method according to claim 13 or 14, wherein, The first LA group sequence is obtained as follows: Among the at least two LA groups, find the idle LA group in which all cloud computing devices are in an idle state; When at least one idle LA group exists, a cloud computing device is selected from the cloud computing devices associated with the at least one idle LA group to build the user computing cluster.

16. The method according to claim 12, wherein, The allocation request is used to request the addition of cloud computing devices to the user's computing cluster. The selection of at least two cloud computing devices includes: The third LA group sequence is obtained by sorting at least one LA group to which the cloud computing devices in the user computing cluster belong in ascending order based on the number of idle devices; In accordance with the order of at least one LA group in the third LA group sequence, at least one cloud computing device is selected from the at least one LA group and added to the user computing cluster.

17. The method according to claim 15 or 16, wherein, Further includes: When the selected cloud computing devices cannot meet the requirements of the allocation request, the at least two LA groups are sorted in descending order based on the number of idle devices to obtain the first LA group sequence; In accordance with the order of the LA groups in the first LA group sequence, at least one cloud computing device is selected sequentially from the cloud computing devices associated with at least one of the at least two LA groups.

18. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein, The processor executes the computer program to implement the steps of the method according to any one of claims 8 to 17.

19. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 8 to 17.

20. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 8 to 17.

Citation Information

Patent Citations

  • Job allocation method, device, electronic equipment and readable storage medium

    CN113190358A

  • Cluster load balancing method and device

    CN117278567A

  • Server allocation method and device, electronic equipment and computer readable storage medium

    CN118200321A

  • Data processing method and device, electronic equipment, storage medium and program product

    CN118869735A

  • Instruction processing apparatus, acceleration unit, and server

    US20220350598A1