Data Processing Method, Apparatus, Electronic Device, Storage Medium and Program Product

By obtaining and sorting the LA group identification information of cloud computing devices, the order of use of cloud computing devices is optimized, the network delay problem caused by manual orchestration is solved, and a more efficient computing process is achieved.

CN118869735BActive Publication Date: 2025-06-17TENCENT CLOUD COMPUTING (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410781428.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2025-06-17
Estimated Expiration
2044-06-17

AI Technical Summary

Technical Problem

In the prior art, there are problems of low efficiency and unreasonable orchestration when manually operating orchestrating cloud computing devices, resulting in an increase in network latency between cloud computing devices.

Method used

By sending a acquisition request to the cloud computing device associated with the target client, it obtains the LA group identification information and sorts the cloud computing devices based on this information to optimize its usage order.

Benefits of technology

It reduces network latency between cloud computing devices, reduces traffic detours, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869735B_ABST
    Figure CN118869735B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, apparatus, electronic device, storage medium and program product, which relate to technical fields such as cloud computing and intelligent transportation. In the present application, a computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group; a target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client; and, based on the LA group identification information, the cloud computing devices are sorted, so as to use each cloud computing device for calculation in the sorted order. Since the cloud computing devices in the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices in the same LA group for calculation, greatly reducing the communication across LA groups between cloud computing devices, minimizing traffic detours, reducing network latency, and reducing traffic consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to technical fields such as cloud computing and big data, and provides a data processing method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] A high-performance computing cluster uses high-performance cloud computing devices as nodes, interconnected through RDMA (Remote Direct Memory Access) technology to provide high-bandwidth and extremely low-latency network services. Users can use the high-performance computing cluster for parallel computing such as large-scale high-performance computing, artificial intelligence, and big data recommendation.

[0003] In related technologies, cloud managers can, according to user requirements, manually select a certain number of cloud computing devices for orchestration through manual operations, and return the selected cloud computing devices to the user so that the user can use the cloud computing devices for relevant business calculations.

[0004] However, the above process involves manually orchestrating cloud computing devices. Manual orchestration has problems such as low efficiency and unreasonable orchestration. If the orchestration is unreasonable, it is very likely to cause network latency in communication between the orchestrated cloud computing devices. Summary of the Invention

[0005] The present application provides a data processing method, apparatus, electronic device, storage medium, and program product, which can reduce the network latency in communication between the cloud computing devices after orchestration. The technical solutions are as follows:

[0006] On the one hand, the present application provides a data processing method, which is applied to a target client, and the target client is an application client of a computing cluster; the computing cluster includes a plurality of aggregation devices LC, a plurality of access device groups LA associated with each LC, and a plurality of cloud computing devices associated with each LA group;

[0007] Wherein, each LA group includes a plurality of LAs, and the cloud computing devices associated with the same LA group communicate based on the LAs of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA group and the associated LC;

[0008] The method includes:

[0009] Sending a first acquisition request to each cloud computing device associated with the target client, where the first acquisition request is used to request to acquire the LA group identification information corresponding to the cloud computing device;

[0010] Receiving LA group identification information returned by each cloud computing device associated with the target client, where the LA group identification information is determined based on the identification of the LA group associated with the cloud computing device;

[0011] Sorting each cloud computing device associated with the target client based on the LA group identification information, so as to perform calculations using each cloud computing device associated with the target client in the sorted order.

[0012] On the other hand, the present application provides a data processing method, which is applied to any cloud computing device associated with the target client in a computing cluster. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0013] Wherein, each LA group includes multiple LAs. Communication between each cloud computing device associated with the same LA group is based on the LAs of the LA group, and communication between each cloud computing device associated with different LA groups is based on the LAs of the corresponding LA group and the associated LC;

[0014] The method includes:

[0015] Receiving a first acquisition request sent by the target client, where the first acquisition request is used to request to acquire the LA group identification information corresponding to the any cloud computing device;

[0016] Based on the first acquisition request, returning the LA group identification information of the any cloud computing device to the target client, where the LA group identification information is determined based on the identification of the LA group associated with the cloud computing device, so that the target client sorts each cloud computing device associated with the target client based on the LA group identification information, so as to perform calculations using each cloud computing device associated with the target client in the sorted order.

[0017] On the other hand, the present application provides a data processing method, which is applied to the cloud management device of the computing cluster. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0018] Wherein, each LA group includes multiple LAs. Communication between each cloud computing device associated with the same LA group is based on the LAs of the LA group, and communication between each cloud computing device associated with different LA groups is based on the LAs of the corresponding LA group and the associated LC;

[0019] The method includes:

[0020] When receiving a second acquisition request sent by any cloud computing device, based on the pre-configured association relationship between the device identifiers of each cloud computing device and the identifier of the LA group, determine the identifier of the LA group associated with the any cloud computing device;

[0021] Based on the identifier of the LA group associated with the any cloud computing device, determine LA group identifier information, and return the LA group identifier information to the any cloud computing device, so that the any cloud computing device performs the following steps:

[0022] When receiving a first acquisition request sent by the associated target client segment, return the LA group identifier information of the any cloud computing device to the target client, so that the target client sorts each cloud computing device associated with the target client based on the LA group identifier information, and calculates using each cloud computing device associated with the target client in the sorted order.

[0023] In a possible implementation manner, the determining LA group identifier information based on the identifier of the LA group associated with the any cloud computing device includes:

[0024] Perform a transformation process on the identifier of the LA group associated with the cloud computing device to obtain an implicit indication value of the identifier of the LA group, and use the implicit indication value as the LA group identifier information corresponding to the cloud computing device.

[0025] On the other hand, the present application provides a data processing system, and the data processing system includes a target client of a computing cluster, a cloud management device, and any cloud computing device associated with the target client;

[0026] Wherein, the computing cluster includes a plurality of aggregation devices LC, a plurality of access device LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; each LA group includes a plurality of LAs, and each cloud computing device associated with the same LA group communicates based on the LAs of the LA group, and each cloud computing device associated with different LA groups communicates based on the LAs of the corresponding LA group and the associated LC;

[0027] The cloud management device is configured to, when receiving a second acquisition request sent by any cloud computing device, determine the identifier of the LA group associated with the any cloud computing device based on the pre-configured association relationship between the device identifiers of each cloud computing device and the identifier of the LA group;

[0028] The cloud management device is configured to determine LA group identifier information based on the identifier of the LA group associated with the any cloud computing device, and return the LA group identifier information to the any cloud computing device;

[0029] The target client is used to send a first acquisition request to each cloud computing device associated with the target client, and the first acquisition request is used to request the acquisition of the LA group identification information corresponding to the cloud computing device;

[0030] Any one of the cloud computing devices is used to receive the first acquisition request sent by the target client, and based on the first acquisition request, return the LA group identification information of any one of the cloud computing devices to the target client, and the LA group identification information is determined based on the identification of the LA group associated with the cloud computing device;

[0031] The target client is used to receive the LA group identification information returned by each cloud computing device associated with the target client;

[0032] The target client is used to sort each cloud computing device associated with the target client based on the LA group identification information, so as to use each cloud computing device associated with the target client for calculation in the sorted order.

[0033] On the other hand, the present application provides a data processing device, and the device is applied to a target client, and the target client is an application client of a computing cluster; the computing cluster includes a plurality of aggregation devices LC, a plurality of access device LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group;

[0034] Wherein, each LA group includes a plurality of LAs, and the cloud computing devices associated with the same LA group communicate based on the LAs of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs;

[0035] The device includes:

[0036] A first sending module, configured to send a first acquisition request to each cloud computing device associated with the target client, and the first acquisition request is used to request the acquisition of the LA group identification information corresponding to the cloud computing device;

[0037] A first receiving module, configured to receive the LA group identification information returned by each cloud computing device associated with the target client, and the LA group identification information is determined based on the identification of the LA group associated with the cloud computing device;

[0038] A sorting module, configured to sort each cloud computing device associated with the target client based on the LA group identification information, so as to use each cloud computing device associated with the target client for calculation in the sorted order.

[0039] In a possible implementation manner, the data processing device of the cloud management device applied to the computing cluster includes an LA group identification information determination module. For any cloud computing device, when the LA group identification information determination module is used to determine the LA group identification information corresponding to the cloud computing device, it is specifically configured to:

[0040] Perform transformation processing on the identification of the LA group associated with the cloud computing device to obtain an implicit indication value of the identification of the LA group, and use the implicit indication value as the LA group identification information corresponding to the cloud computing device.

[0041] In a possible implementation manner, the device further includes:

[0042] A second sending module, configured to send an allocation request to the cloud management device of the computing cluster, where the allocation request is used to request to allocate a cloud computing device for a target client;

[0043] A third receiving module, configured to receive the device identifiers of each cloud computing device returned by the cloud management device, and obtain each cloud computing device associated with the target client based on the device identifiers of each cloud computing device;

[0044] Among them, the data processing device applied to the cloud management device includes a cloud computing device determination module. When the cloud computing device determination module is used to determine each cloud computing device associated with the target client, it includes:

[0045] A first determination unit, configured to determine the number of idle devices of the cloud computing devices in the idle state associated with each LA group;

[0046] An allocation unit, configured to determine the allocation order of each LA group based on the number of idle devices of each LA group, and sequentially select cloud computing devices from the cloud computing devices associated with at least one LA group according to the allocation order of each LA group, and allocate the cloud computing devices to the target client.

[0047] In a possible implementation manner, the allocation request is a new request, and the new request is used to request to create a user computing cluster for a target client:

[0048] The allocation unit is configured to:

[0049] Perform a descending order arrangement on each LA group based on the number of idle devices of each LA group to obtain a first LA group sequence;

[0050] Take the arrangement order of each LA group in the first LA group sequence as the allocation order, sequentially select cloud computing devices from the cloud computing devices associated with at least one LA group, and construct a user computing cluster for the target client based on the selected cloud computing devices.

[0051] In a possible implementation, when the allocation unit selects cloud computing devices from at least one cloud computing device associated with an LA group in the order of arrangement of each LA group in the first LA group sequence, it is specifically configured to:

[0052] If there are at least two LA groups with the same number of idle devices in the first LA group sequence, based on the number of clients corresponding to each LA group in the at least two LA groups, the at least two LA groups in the first LA group sequence are sorted in ascending order to obtain a second LA group sequence, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group;

[0053] Select cloud computing devices from at least one cloud computing device associated with an LA group in the order of arrangement of each LA group in the second LA group sequence.

[0054] In a possible implementation, when the allocation unit sorts each LA group in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence, it is specifically configured to:

[0055] Find the idle LA groups in each LA group, where each cloud computing device associated with the idle LA group is in an idle state;

[0056] If there are idle LA groups in each LA group, select cloud computing devices from the cloud computing devices associated with the idle LA groups, and construct a user computing cluster for the target client based on the selected cloud computing devices;

[0057] If there are no idle LA groups in each LA group, sort each LA group in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence.

[0058] In a possible implementation, the allocation request is an expansion request, and the expansion request is used to request adding cloud computing devices to the user computing cluster of the target client:

[0059] The allocation unit is configured to:

[0060] Sort each associated LA group in ascending order based on the number of idle devices in the associated LA groups in each LA group to obtain a third LA group sequence, where each cloud computing device associated with the associated LA group includes the cloud computing devices already allocated to the target client;

[0061] Select cloud computing devices from at least one associated LA group in the order of arrangement of each associated LA group in the third LA group sequence and allocate them to the target client.

[0062] On the other hand, the present application provides a data processing device, which is applied to any cloud computing device associated with a target client in a computing cluster. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0063] Wherein, each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, and cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs;

[0064] The device includes:

[0065] A second receiving module, configured to receive a first acquisition request sent by the target client, where the first acquisition request is used to request to acquire the LA group identification information corresponding to the any cloud computing device;

[0066] A first returning module, configured to, based on the first acquisition request, return the LA group identification information of the any cloud computing device to the target client. The LA group identification information is determined based on the identification of the LA group associated with the cloud computing device, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client in the sorted order for calculation.

[0067] In a possible implementation manner, the device further includes:

[0068] A third sending module, configured to send a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to acquire the LA group identification information of the any cloud computing device;

[0069] A fourth receiving module, configured to receive the LA group identification information returned by the cloud management device based on the second acquisition request.

[0070] On the other hand, the present application provides that the device is applied to the cloud management device of the computing cluster. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0071] Wherein, each LA group includes multiple LAs. Cloud computing devices associated with the same LA group communicate based on the LAs of that LA group, and cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs;

[0072] The device includes:

[0073] A first determination module, configured to, when receiving a second acquisition request sent by any cloud computing device, determine an identifier of an LA group associated with the any cloud computing device based on an association relationship between device identifiers of each pre-configured cloud computing device and an identifier of the LA group;

[0074] A second determination module, configured to determine LA group identifier information based on the identifier of the LA group associated with the any cloud computing device, and return the LA group identifier information to the any cloud computing device, so that the any cloud computing device performs the following steps:

[0075] When receiving a first acquisition request sent by an associated target client segment, return the LA group identifier information of the any cloud computing device to the target client, so that the target client sorts each cloud computing device associated with the target client based on the LA group identifier information, and uses each cloud computing device associated with the target client for calculation in the sorted order.

[0076] In a possible implementation manner, when determining the LA group identifier information based on the identifier of the LA group associated with the any cloud computing device, the second determination module is specifically configured to:

[0077] Perform a transformation process on the identifier of the LA group associated with the cloud computing device to obtain an implicit indication value of the identifier of the LA group, and use the implicit indication value as the LA group identifier information corresponding to the cloud computing device.

[0078] In a possible implementation manner, the apparatus further includes:

[0079] An idle device number determination module, configured to, when receiving an allocation request from a target client, determine the number of idle devices of cloud computing devices in an idle state associated with each LA group, where the allocation request is used to request to allocate cloud computing devices to the target client;

[0080] An allocation module, configured to determine an allocation order of each LA group based on the number of idle devices of each LA group, and sequentially select cloud computing devices from cloud computing devices associated with at least one LA group according to the allocation order of each LA group, and allocate the cloud computing devices to the target client;

[0081] A second return module, configured to return the device identifiers of the allocated cloud computing devices to the target client.

[0082] On the other hand, provided is an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the above data processing method.

[0083] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned data processing method is implemented.

[0084] On the other hand, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned data processing method is implemented.

[0085] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are as follows:

[0086] For the data processing method provided in the present application, the computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group; the cloud computing devices associated with the same LA group communicate with each other based on the LA of the LA group, and the cloud computing devices associated with different LA groups communicate with each other based on the LA of the corresponding LA group and the associated LC; the target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client; and, the cloud computing devices are sorted based on the LA group identification information, so as to use each cloud computing device for calculation in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices of the same LA group for calculation, greatly reducing the communication between cloud computing devices across LA groups, minimizing traffic detours, reducing network latency, and reducing traffic consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0088] Figure 1 It is a schematic diagram of the implementation environment of a data processing method provided in an embodiment of the present application;

[0089] Figure 2 It is a schematic diagram of signaling interaction of a data processing method provided in an embodiment of the present application;

[0090] Figure 3 It is a schematic diagram of the system architecture of a computing cluster provided in an embodiment of the present application;

[0091] Figure 4 It is a schematic diagram of a data processing flow provided in an embodiment of the present application;

[0092] Figure 5 It is a schematic diagram of an allocation process provided in an embodiment of the present application;

[0093] Figure 6 It is a schematic diagram of an allocation flow provided in an embodiment of the present application;

[0094] Figure 7 A schematic diagram of a data processing flow provided by an embodiment of the present application;

[0095] Figure 8 A schematic diagram of the architecture of a computing cluster system provided by an embodiment of the present application;

[0096] Figure 9 A schematic diagram of the architecture of a computing cluster system provided by an embodiment of the present application;

[0097] Figure 10 A schematic diagram of the structure of a data processing system provided by an embodiment of the present application;

[0098] Figure 11 A schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0099] Figure 12 A schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0100] Figure 13 A schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0101] Figure 14 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0102] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.

[0103] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. The terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the art of the present technology.

[0104] It should be understood that in the specific embodiments of the present application, any user-related data such as LA group identification information, the device identification of cloud computing devices, the user computing cluster associated with the user's client, or the associated cloud computing devices needs to obtain user permission or consent when the above embodiments of the present application are applied to specific products or technologies, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. That is to say, if any of the above user-related data is involved in the embodiments of the present application, this data needs to be obtained with the authorization and consent of the subject and comply with the relevant laws, regulations, and standards of the country and region.

[0105] The following introduces and explains the terms and related technologies involved in the present application:

[0106] Computing cluster: It includes multiple aggregation devices LC, multiple access device LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Among them, each LA group includes multiple LAs, and the cloud computing devices associated with the same LA group communicate based on the LAs of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs.

[0107] For example, the computing cluster can be HCC (Hyper Computing Cluster, high-performance computing cluster), and the cloud computing devices in the computing cluster can perform RDMA (Remote Direct Memory Access) communication; providing high-bandwidth and extremely low-latency network services for users to meet the large-scale and high-performance parallel computing requirements of related services.

[0108] LA group: An LA group includes multiple LAs (Layer for access, access devices). In the computing cluster, the LAs of the same LA group can be connected to at least two cloud computing devices, so that the cloud computing devices of the same LA group can perform RDMA communication only based on the LAs.

[0109] LC (Layer for core, aggregation device): It is used to converge the traffic of multiple LA groups to achieve RDMA communication between the cloud computing devices across LA groups. Connecting the access devices of the LA groups to achieve communication between devices; for example, LC can be a switch.

[0110] Cloud computing device: A cloud server used for parallel computing in the computing cluster. For example, the cloud computing device can be a GPU server of a high-performance computing cluster.

[0111] Figure 1 It is a schematic diagram of the implementation environment of a data processing method provided by the present application. AsFigure 1 As shown in the figure, the implementation environment includes: a cloud computing device 101, a client 102, and a cloud management device 103.

[0112] The cloud computing device 101 can be a computing device in a computing cluster, and users can use the cloud computing device 101 for computing.

[0113] For example, the computing cluster can be an HCC high-performance computing cluster, and RDMA communication can be performed between cloud computing devices to provide users with high-bandwidth and extremely low-latency network services to meet the large-scale and high-performance parallel computing requirements of related services.

[0114] In an example of a possible scenario, in the field of artificial intelligence, the cloud computing device can be a GPU server in an HCC computing cluster, and the HCC computing cluster can be used for large-scale AI training, which can meet the computing requirements of the business for high computing performance, high stability, and high real-time performance of the GPU server.

[0115] In another example of a possible scenario, in the field of industrial simulation, such as the automotive industry, which needs to utilize, the HCC computing cluster can be used for simulation calculations to drive design, and can quickly respond to the real-time changing simulation requirements of enterprises such as industrial manufacturing, and promote product R & D in a timely manner.

[0116] Exemplarily, the client 102 is an application client of the computing cluster. The cloud management device 103 is the background management end of the computing cluster. Users can perform data interaction with the cloud computing device 101 and the cloud management device 103 respectively through the client 102.

[0117] Exemplarily, users can trigger an allocation request for allocating cloud computing devices on the client 102. For example, the allocation request can be a new request for creating a user computing cluster, or the allocation request can also be an expansion request for expanding an existing user computing cluster. The client 102 can send the allocation request to the cloud management device 103. The cloud management device 103 can allocate a cloud computing device for the client 102 from the computing cluster based on the allocation request. So that the client 102 can use the allocated cloud computing device for computing processing of related services.

[0118] In the embodiments of the present application, each cloud computing device within the same LA group only needs to communicate based on the LA within the LA group. For cloud computing devices in different LA groups, cross-LA group communication is required. Specifically, cross-LA group communication can be achieved only based on the LAs in at least two different LA groups and the LCs associated with the at least two LA groups. Therefore, compared with the communication between cloud computing devices in the same LA group, cross-LA group communication between cloud computing devices will significantly increase network latency and consume more traffic. In other words, communication within the same LA group is more efficient, faster, and requires less traffic than cross-LA group communication.

[0119] In the embodiments of the present application, the cloud computing device 101 may send a second acquisition request to the cloud management device 103. The second acquisition request is used to acquire the LA group identification information of the cloud computing device 101. The cloud management device 103 may, based on the second acquisition request, determine the identification of the LA group associated with the cloud computing device 101 from the pre-stored association relationship between the device identifications of each cloud computing device 101 and the identifications of the LA groups, generate LA group identification information based on the determined identification of the LA group, and return the LA group identification information to the cloud computing device 101.

[0120] After the cloud management device 103 assigns the cloud computing device 101 to the client 102, the client 102 is associated with the assigned cloud computing device 101. The client 102 may send a first acquisition request to each associated cloud computing device 101, and each cloud computing device 101 returns its respective LA group identification information to the client 102 based on the first acquisition request. The client 102 may sort the associated cloud computing devices based on the LA group identification information, and use the associated cloud computing devices for calculation in the sorted order.

[0121] It should be noted that in the embodiments of the present application, among the cloud computing devices sorted according to the LA group identification information, the identifications of cloud computing devices belonging to the same LA group are adjacent, so that cloud computing devices in the same LA group are concentrated together. Based on this, when the client uses the cloud computing devices for calculation in this order, cross-LA group RDMA communication between cloud computing devices can be greatly reduced, traffic detours can be minimized, network latency can be reduced, and traffic consumption can be reduced.

[0122] It should be noted that the cloud computing device 101 or the cloud management device 103 may be a server in the cloud. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, and big data and artificial intelligence platforms.

[0123] The client 102 is a management platform on the user side, and the user can use the client 102 on the terminal. For example, the user logs in to their account on the client 102 and uses the services of the computing cluster. The client 102 can be a management platform in the form of a web page, for example, a management platform logged in through a web page of a browser, or can be an independent application, or can also be a program plug-in located in an independent application; the specific manifestation form of the client 102 in this application is not limited. Among them, the terminal can be a smart phone, a tablet computer, a notebook computer, a digital broadcast receiver, a desktop computer, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc.

[0124] Figure 2 It is a signaling interaction schematic diagram of a data processing method provided by an embodiment of this application. The data processing method can be interactively executed among the target client of the computing cluster, the cloud management device, and the cloud computing device. As Figure 2 shown, the data processing method includes the following steps.

[0125] Step 201, when the cloud management device receives a second acquisition request sent by any cloud computing device, the cloud management device determines the identifier of the LA group associated with the any cloud computing device based on the association relationship between the device identifiers of each cloud computing device and the identifier of the LA group pre-configured.

[0126] Among them, the second acquisition request is used to request to acquire the LA group identifier information of the any cloud computing device. The LA group identifier information is used to identify the LA group associated with the any cloud computing device.

[0127] In an embodiment of this application, the computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Among them, each LA group includes multiple LAs, and the cloud computing devices associated with the same LA group communicate based on the LAs of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs.

[0128] Among them, each LC being associated with multiple LA groups means that each LC is communicatively connected to the LAs of multiple LA groups, and the communication between the multiple LA groups connected by the LC passes through the LC. Each LA group being associated with multiple cloud computing devices means that the LAs included in each LA group are communicatively connected to multiple cloud computing devices, and the communication between the LAs connecting each cloud computing device passes through the LA.

[0129] As Figure 3 shown, multiple LAs within the same dotted line can form an LA group, and the LAs of each LA group can be connected to the LC above and the cloud computing device below. The cloud computing device can be Figure 3In the GPU server, each GPU server may include 8 GPU cards. GPU servers under the same LA group can communicate with each other using the RDMA technology without going through the LC. However, GPU servers under different LAs need to go through the LC and LA to perform RDMA communication. The RDMA communication of GPU servers under the same LA group has a smaller network latency and consumes less traffic than the RDMA communication of GPU servers under different LA groups.

[0130] Exemplarily, any one of the cloud computing devices is any one of the cloud computing devices associated with the target client. After the cloud management device assigns a cloud computing device to the target client, the target client is associated with the assigned cloud computing device.

[0131] In a possible implementation manner, before step 201, the cloud management device may perform the following step A1:

[0132] Step A1: The cloud management device may collect the device identifiers of each cloud computing device and the identifiers of the associated LA groups, and store the association relationship between the device identifiers of each cloud computing device and the identifiers of the LA groups.

[0133] Exemplarily, when each cloud computing device in the computing cluster is newly launched, that is, when the cloud computing device is divided into each LA group of the computing cluster and has not been assigned to a client yet, the cloud management device may collect the device identifiers of each cloud computing device and the identifiers of the associated LA groups, and perform associated storage on the device identifiers of each cloud computing device and the identifiers of the LA groups. For example, in a key-value manner, the device identifiers and the LA group identifiers are associated and stored in the DB (database) of the cloud management device.

[0134] Exemplarily, the device identifier of the cloud computing device may be the IP (Internet Protocol) address of the cloud computing device. The identifier of the LA group may be the ID (Identity) of the LA group. Of course, the cloud management device may collect other information of the associated LA group, such as the name of the LA group, the total number of device identifiers of the cloud computing devices associated with the LA group, etc., and this application does not make any limitations on this.

[0135] In this step, each cloud computing device may send a second acquisition request to the cloud management device. For example, the cloud computing device may send a second acquisition request to the cloud management device through the Metadata interface.

[0136] For a second acquisition request of any cloud computing device, the cloud management device may, based on the device identifier of the any cloud computing device, obtain the identifier of the LA group associated with the device identifier of the any cloud computing device from the pre-stored association relationship.

[0137] Step 202: The cloud management device determines LA group identifier information based on the identifier of the LA group associated with the any cloud computing device, and returns the LA group identifier information to the any cloud computing device.

[0138] In this step, the cloud management device may perform a transformation process on the identifier of the LA group associated with the any cloud computing device to obtain LA group identifier information. The transformation process may be a process for desensitizing the identifier of the LA group. The LA group identifier information is used to indicate the identifier of the LA group associated with the cloud computing device. The LA group identifier information of cloud computing devices in different LA groups is also different.

[0139] In a possible implementation manner, for any cloud computing device, the LA group identifier information corresponding to the cloud computing device is obtained by the cloud management device by executing step B1. That is, the implementation manner of determining the LA group identifier information based on the identifier of the LA group associated with the any cloud computing device includes the following step B1:

[0140] Step B1: Perform a transformation process on the identifier of the LA group associated with the cloud computing device to obtain an implicit indication value of the identifier of the LA group, and use the implicit indication value as the LA group identifier information corresponding to the cloud computing device.

[0141] Exemplarily, the cloud management device may perform a change process on the identifier of the LA group based on a pre-configured mapping transformation algorithm to obtain the LA group identifier information. For example, a hash algorithm may be used to calculate the hash value of the ID of the LA group.

[0142] Based on this, through the LA group identifier information, on the premise of being able to distinguish different LA groups, the original ID identifier of the LA group is hidden, realizing the desensitization process of the true ID identifier of the LA group in the cloud, and as much as possible ensuring the security of the network of the computing cluster in the cloud.

[0143] Correspondingly, corresponding to the process of the above steps 201-202, on the side of any cloud computing device, before step 201, the any cloud computing device may execute step C1:

[0144] Step C1: The any cloud computing device sends a second acquisition request to the cloud management device of the computing cluster.

[0145] After step 202, the any cloud computing device may execute step D1:

[0146] Step D1. Any one of the cloud computing devices receives the LA group identification information returned by the cloud management device based on the second acquisition request.

[0147] As Figure 4 shown, during the online deployment stage of the cloud computing device, for example, when the GPU server goes online, the device import process can be carried out. When the backend management device retrieves and executes the device import process, it collects the identification of the LA group where the cloud computing device is located, and writes the device identification of the cloud computing device and the identification of the LA group into the database DB in the form of key-value. Figure 4 In

[0148] it, the cloud computing device can access the backend through the Metadata interface via the database DB. After reading and performing desensitization processing through the Metadata interface, the LA group identification information is obtained, such as the hash value of the identification of the LA group; and the desensitized LA group identification information is returned to the cloud computing device.

[0149] Step 203. The target client sends a first acquisition request to each of the cloud computing devices associated with the target client.

[0150] Among them, the first acquisition request is used to request to obtain the LA group identification information corresponding to the cloud computing device. Exemplarily, the target client can send the first acquisition request to each of the associated cloud computing devices through the SSH (Secure Shell) protocol.

[0151] In a possible implementation manner, each of the cloud computing devices associated with the target client can be pre-allocated for the target client by the cloud management device. The allocation process may include the following interaction steps E1 - step E5:

[0152] Step E1. The target client sends an allocation request to the cloud management device, and the allocation request is used to request to allocate cloud computing devices for the target client.

[0153] In a possible case, the allocation request is a new request, and the new request is used to request to create a user computing cluster for the target client.

[0154] Exemplarily, creating a user computing cluster means the process of allocating at least one cloud computing device for the target client and constructing the allocated cloud computing devices into a user computing cluster associated with the target client.

[0155] For example, a user can trigger a new creation operation on the target client. For example, the user can trigger configuration operations on the target client for the number of devices, cluster name, available zone, etc. of the user computing cluster to be newly created. If the user does not have a user computing cluster yet, a new user computing cluster can be applied for and created through the new creation operation; of course, if the user already has a user computing cluster, it can also be created as needed. In one possible application example, the instance RDMA networks under the same user computing cluster can be interconnected, and the instance RDMA networks across user computing clusters are isolated from each other.

[0156] In another possible case, the allocation request is an expansion request, which is used to request to add cloud computing devices to the user computing cluster of the target client.

[0157] Exemplarily, if the user already has a user computing cluster, the user can also trigger an expansion operation on the target client to expand the number of cloud computing devices of the existing user computing cluster.

[0158] Step E2: The cloud management device receives the allocation request from the target client.

[0159] After receiving the allocation request, the cloud management device can allocate cloud computing devices to the target client through the following steps E3 - E4. That is, each cloud computing device associated with the target client is determined by the cloud management device through the following steps E3 - E4:

[0160] Step E3: The cloud management device determines the number of idle cloud computing devices associated with each LA group in the idle state.

[0161] Herein, the idle state means that the cloud computing device has not been allocated to any client.

[0162] Step E4: The cloud management device determines the allocation order of each LA group based on the number of idle devices of each LA group, and sequentially selects cloud computing devices from the cloud computing devices associated with at least one LA group according to the allocation order of each LA group, and allocates them to the target client.

[0163] In this step, based on the two possible cases of the allocation request, the implementation method of step E4 is correspondingly divided into the following case one and case two.

[0164] Case one: The allocation request is a new creation request. Correspondingly, the implementation method of step E4 may include steps E41 - E42:

[0165] Step E41: Based on the number of idle devices of each LA group, perform a descending order arrangement on each LA group to obtain the first LA group sequence.

[0166] Among them, the number of idle devices refers to the number of cloud computing devices in an idle state among the cloud computing devices associated with each LA group. In this step, each LA group can be arranged in descending order of the number of idle devices to obtain the first LA group sequence. That is, in the first LA group sequence, the earlier the order of the LA group, the more idle devices the LA group has.

[0167] Step E42: Taking the arrangement order of each LA group in the first LA group sequence as the allocation order, successively select cloud computing devices from the cloud computing devices associated with at least one LA group, and construct a user computing cluster for the target client based on the selected cloud computing devices.

[0168] In this step, cloud computing devices can be successively selected from each LA group in descending order of the number of idle devices. For example, if the number of devices in the user computing cluster to be newly built is the first quantity, then in descending order of the number of idle devices, cloud computing devices are preferentially selected from the LA groups with more idle devices until the first quantity of cloud computing devices is selected.

[0169] In a possible way, if the total number of devices of the cloud computing devices associated with each LA group is the same, the idle LA group with all the associated cloud computing devices in an idle state has the most idle devices, and cloud computing devices are preferentially selected from the idle LA group.

[0170] In a possible implementation manner, for LA groups with the same number of idle devices, cloud computing devices can also be preferentially selected from the LA group with fewer corresponding client numbers. Correspondingly, in step E42, taking the arrangement order of each LA group in the first LA group sequence as the allocation order, successively selecting cloud computing devices from the cloud computing devices associated with at least one LA group includes the following steps E42-1 and E42-2:

[0171] Step E42-1: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, based on the number of clients corresponding to each LA group in the at least two LA groups, arrange the at least two LA groups in the first LA group sequence in ascending order to obtain the second LA group sequence.

[0172] Among them, the number of clients refers to the number of clients associated with the cloud computing devices associated with each LA group.

[0173] Step E42-2: Taking the arrangement order of each LA group in the second LA group sequence as the allocation order, successively select cloud computing devices from the cloud computing devices associated with at least one LA group.

[0174] For example, in the first LA group sequence, LA group a and LA group b are located at the 1st and 2nd positions respectively. Both LA group a and LA group b are associated with 100 cloud computing devices, and the number of idle devices is 60 for each. Among the 40 cloud computing devices already allocated in LA group a, they are allocated to 3 clients in total; among the 40 cloud computing devices already allocated in LA group b, they are allocated to 2 clients in total. Then, in the first LA sequence, move the order of LA group b to the front of LA group a. After adjustment, LA group a is located at the 2nd position and LA group b is located at the 1st position; so as to preferentially select cloud computing devices for the target client from the 60 idle cloud computing devices associated with LA group b.

[0175] In a possible implementation manner, cloud computing devices can also be first selected from the idle LA groups. Only when there is no idle LA group or the number of devices in the idle LA group is insufficient, step E41 is executed. Correspondingly, the implementation manner of step E41 can include the following steps E41-1, E41-2, and E41-3:

[0176] Step E41-1: Search for the idle LA groups in each LA group, where all the cloud computing devices associated with the idle LA group are in the idle state;

[0177] Step E41-2: If there is an idle LA group in each LA group, select cloud computing devices from the cloud computing devices associated with the idle LA group, and build a user computing cluster for the target client based on the selected cloud computing devices;

[0178] Step E41-3: If there is no idle LA group in each LA group, perform a descending order arrangement of each LA group based on the number of idle devices in each LA group to obtain the first LA group sequence.

[0179] For example, if the total number of devices associated with each LA group is the same, and the number of idle devices in the idle LA group is the largest, the idle LA group can be searched first, and cloud computing devices can be preferentially selected from the idle LA group. If there is no idle LA group, then perform the step of arranging each LA group in descending order based on the number of idle devices.

[0180] In a possible manner, if there is an idle LA group in each LA group, steps E41-2 and E41-3 can also be executed based on the total number of devices in the idle LA group and the number of devices required for the user computing cluster to be newly built; the number of devices required for the user computing cluster to be newly built is the first quantity. Correspondingly, this process can include the following steps:

[0181] If there is an idle LA group in each LA group and the total number of devices in the idle LA group is not less than the first quantity, select the first quantity of cloud computing devices from the cloud computing devices associated with the idle LA group to build a user computing cluster including the first quantity of cloud computing devices;

[0182] If there is an idle LA group among all LA groups, and the total number of devices in the idle LA group is lower than the first quantity, select all cloud computing devices from the cloud computing devices associated with the idle LA group, and based on the number of idle devices in non-idle LA groups among all LA groups, sort all LA groups in descending order to obtain a first LA group sequence; taking the arrangement order of each LA group in the first LA group sequence as the allocation order, sequentially select the remaining quantity of cloud computing devices from the cloud computing devices associated with at least one LA group to construct a user computing cluster including the first quantity of cloud computing devices; wherein, the sum of the remaining quantity and the total number of devices in the idle LA group is the first quantity.

[0183] If there is no idle LA group among all LA groups, based on the number of idle devices in all LA groups, sort all LA groups in descending order to obtain a first LA group sequence; and taking the arrangement order of each LA group in the first LA group sequence as the allocation order, sequentially select the first quantity of cloud computing devices from the cloud computing devices associated with at least one LA group to construct a user computing cluster including the first quantity of cloud computing devices.

[0184] As Figure 5 shown, a dashed box represents an LA group, each square within the dashed box represents a cloud computing device, squares with different filling patterns or filling colors represent cloud computing devices associated with different clients, and blank squares without any filled pattern and color represent cloud computing devices in an idle state. Currently, there is no idle LA group among all LA groups, and two sorts can be performed. The first sort is in descending order according to the number of idle devices, with the LA group having more idle devices ranked first. The second sort is in ascending order according to the number of clients, with the LA group having fewer clients ranked first. Based on this, the allocation strategy in the newly created scenario can be configured as:

[0185] Give priority to selecting cloud computing devices from idle LA groups;

[0186] If there is no idle LA group, give priority to selecting cloud computing devices from the LA group whose number of idle devices can meet the device quantity required by the user computing cluster;

[0187] If there are multiple LA groups whose number of idle devices can meet the device quantity required by the user computing cluster, give priority to selecting cloud computing devices from the LA group with fewer corresponding clients;

[0188] If there is no LA group whose number of idle devices can meet the device quantity required by the user computing cluster, give priority to selecting cloud computing devices from the LA group with the most idle devices.

[0189] Case 2: The allocation request is an expansion request. Correspondingly, the implementation manner of step E4 may include steps E43 - E44:

[0190] Step E43: Based on the number of idle devices in the associated LA groups within each LA group, sort the associated LA groups in ascending order to obtain a third LA group sequence. Each cloud computing device associated with the associated LA group includes the cloud computing devices that have been allocated to the target client.

[0191] Step E44: Using the arrangement order of the associated LA groups in the third LA group sequence as the allocation order, sequentially select cloud computing devices from at least one associated LA group and allocate them to the target client.

[0192] Exemplarily, for the allocation scenario of expansion, all the associated LA groups where the cloud computing devices associated with the target client are located can be searched, and they are sorted in ascending order according to the number of idle devices in the associated LA groups. Cloud computing devices are preferentially allocated from the LA groups with fewer idle devices, so that the cloud computing devices of a client can fall into the same LA group as much as possible, that is, the cloud computing devices of the LA group belong to one user as much as possible.

[0193] For example, if 3 cloud computing devices need to be added to the user computing cluster of user A, the LA groups where the cloud computing devices of user A are located are found to include group 1, group 2, and group 3. The number of idle devices in group 1 is 6, the number of idle devices in group 2 is 8, and the number of idle devices in group 3 is 3. They are sorted in ascending order according to the number of idle devices from less to more as: group 3 < group 1 < group 2; then 3 cloud computing devices are preferentially selected from group 3.

[0194] It should be noted that steps E41 - E42 and steps E43 - E44 are the allocation steps for two allocation scenarios respectively. This application does not limit the order of steps E41 - E42 and the order of steps E43 - E44. That is, steps E41 - E42 can be executed before steps E43 - E44 or after steps E43 - E44.

[0195] Step E5: The cloud management device returns the device identifiers of the allocated cloud computing devices to the target client.

[0196] Among them, the cloud management device can deduct the allocated cloud computing devices from the cloud computing devices in the computing cluster and send the device identifiers of the cloud computing devices allocated to the target client to the target client. Therefore, this process can also be called the process of deducting and packing.

[0197] Step E6: The target client receives the device identifiers of the cloud computing devices returned by the cloud management device and obtains the cloud computing devices associated with the target client based on the device identifiers of the cloud computing devices.

[0198] Next, the allocation process of the above steps E1 - E6 will be introduced using the Figure 6 shown process.

[0199] like Figure 6 As shown, the allocation process includes the following steps:

[0200] (1) The target client initiates an allocation request;

[0201] (2) The cloud management device can obtain the current resource status of the user of the target client, for example, the target client already has a user computing cluster or has not created a user cluster;

[0202] (3) The cloud management device determines whether the allocation request is a new request;

[0203] (4) If it is a new creation request, the cloud management device can execute the new creation process, which can be sorted twice: the first sorting is to sort in descending order according to the number of idle devices, and the LA group with more idle devices is at the front; the second sorting is to sort in ascending order according to the number of clients, and the LA group with fewer clients is at the front; and execute (6);

[0204] (5) For the capacity expansion request, the cloud management device can execute the capacity expansion process, specifically, obtain the LA group resource status of the user, specifically, obtain the LA group associated with the user's cloud computing device, and obtain the number of idle devices in the associated LA group, and sort them: sort in ascending order according to the number of idle devices, with the number of idle devices being the smallest at the front; and execute (6);

[0205] (6) The cloud management device selects cloud computing devices and allocates them to the target client in the order of the sorted LA groups, which is the process of entering the packing deduction process;

[0206] (7) The cloud management device determines whether the deduction is complete. If the deduction is complete, the packing is successful, that is, the allocation is successful. If the deduction is not complete, it means that some cloud computing devices have been successfully selected. Furthermore, the corresponding process can be executed again for allocation until all packing is successful.

[0207] Step 204: The cloud computing device receives the first acquisition request sent by the target client, and based on the first acquisition request, returns the LA group identification information of any cloud computing device to the target client.

[0208] The LA group identification information is determined based on the identification of the LA group associated with the cloud computing device. In this step, the cloud computing device may return the LA group identification information of the cloud computing device to the target client based on the SSH protocol.

[0209] In a possible implementation, the target client may execute step 203 in response to triggering of a target event, where the target event includes at least one of the following:

[0210] Receive the device identification of the cloud computing device assigned by the cloud management device;

[0211] Receive the device identifier of the changed cloud computing device sent by the cloud management device;

[0212] Detect that the current time reaches the target time period.

[0213] In a possible example, when the target client receives the device identifier of the cloud computing device allocated by the cloud management device, the target client may send a first acquisition request to each cloud computing device allocated by the cloud management device. For example, when creating a new user computing cluster, a first acquisition request may be sent to the cloud computing devices in the newly created user computing cluster. For another example, when expanding an existing user computing cluster, a first acquisition request may also be sent to the newly added cloud computing devices.

[0214] In a possible example, the changed cloud computing device may be a faulty cloud computing device. For example, if device A fails, the cloud management device may return the device identifier of device A and the device identifier of its replacement device B to the target client. The target client may send a first acquisition request to replacement device B and may also delete the LA group identifier information of device A from the local.

[0215] In a possible example, the target client may also periodically update the LA group identifier information of the cloud computing devices. For example, at regular intervals of the target time period, a first acquisition request is sent to the cloud computing devices regularly; for example, the LA group identifier information of each cloud computing device may be detected regularly, and a first acquisition request is sent to the cloud computing devices whose LA group identifier information has expired or has not been acquired before regularly.

[0216] Step 205: The target client receives the LA group identifier information returned by each cloud computing device associated with the target client.

[0217] The target client may receive the LA group identifier information returned by each cloud computing device associated with the target client based on the SSH protocol. Based on this, the target client can obtain the distribution of the LA groups of the associated cloud computing devices.

[0218] Among them, in the sending and receiving processes of steps 204 and 205, since the LA group identifier information is the information after the cloud management device performs transformation processing on the identifier of the LA group and is the information after desensitization processing on the original identifier of the LA group, the security of the cloud network information during network transmission is greatly improved, and at the same time, the requirement for the client to obtain the LA group identifier information in a timely manner is met.

[0219] Step 206: The target client sorts each cloud computing device associated with the target client based on the LA group identifier information, so as to use each cloud computing device associated with the target client for calculation in the sorted order.

[0220] The target client can arrange each cloud computing device in ascending order according to the LA group identification information, for example, according to the implicit indication value of the LA group identification, so that the cloud computing devices in the same LA group are concentrated and adjacent. Based on this, each cloud computing device can be used in turn in the sorted order to calculate the relevant business. For example, for the relevant business using the Ring-Allreduce algorithm, each computing node will be arranged in a ring, and in the calculation process of the relevant business, data interaction is required between the computing nodes arranged in the ring, such as the transmission of intermediate calculation data; then each cloud computing device can be used in the sorted order, based on which, it is helpful to obtain the cloud computing devices of the same LA group that are arranged together to perform the calculation of the Ring-Allreduce algorithm, that is, to make each computing node arranged in a ring in the same LA group, thereby greatly reducing the network delay in the calculation process of the Ring-Allreduce algorithm, greatly saving traffic consumption, and ensuring the efficiency of the calculation process.

[0221] like Figure 7 As shown, each cloud computing device can obtain the desensitized LA group identification information through the Metadata interface and return the LA group identification information to the client. The client sorts the IP addresses of the cloud computing devices according to the LA group identification information, and the IP addresses of the cloud computing devices with the same LA group identification information are arranged together.

[0222] like Figure 8 As shown, rectangles with the same fill pattern represent devices associated with the same client. Figure 8 In the present invention, GPU servers associated with the same client are distributed in the same LA group. However, as time goes by, the resources associated with the client are adjusted, such as capacity expansion, capacity reduction, or machine migration caused by machine failure, and the distribution of GPU servers of the same client in various LA groups gradually becomes fragmented. Specifically, GPU servers of the same client may be distributed in different LA groups. Based on this, the communication between GPU servers of the same client needs to cross the LA group, which may cause network delays. Through the embodiments of the present application, the target client can arrange the GPU devices according to the LA group identification information, so that the cloud computing devices of the same LA group are concentrated together. Based on this, when the client uses the cloud computing devices for calculations in this order, the RDMA communication between cloud computing devices across LA groups can be greatly reduced, the traffic detour can be minimized, the network delay can be reduced, and the traffic consumption can be reduced.

[0223] The data processing method provided by this application. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Among the cloud computing devices associated with the same LA group, communication is carried out based on the LA of that LA group. Among the cloud computing devices associated with different LA groups, communication is carried out based on the LA of the corresponding LA group and the associated LC. The target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client. And, based on the LA group identification information, each cloud computing device is sorted, and each cloud computing device is used for calculation in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices of the same LA group for calculation, greatly reducing the communication across LA groups between cloud computing devices, minimizing traffic detours, reducing network latency, and reducing traffic consumption.

[0224] The data processing method of this application relates to technologies such as cloud computing and cloud storage in cloud technology.

[0225] It can be understood that cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.

[0226] As a basic capability provider of cloud computing, a cloud computing resource pool (abbreviated as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtual machines containing operating systems), storage devices, and network devices.

[0227] According to the logical function division, a PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and an SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Or the SaaS can be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0228] It can be understood that cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that, through functions such as cluster applications, grid technology, and distributed storage file systems, aggregates a large number of various types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and jointly provide data storage and business access functions to the outside world.

[0229] Figure 10 It is a schematic structural diagram of a data processing system provided by an embodiment of the present application. As Figure 10 shown, the system includes: a target client 1001 of a computing cluster, a cloud management device 1002, and any cloud computing device 1003 associated with the target client;

[0230] Among them, the computing cluster includes multiple aggregation devices LC, multiple access device LA groups associated with each LC, and multiple cloud computing devices associated with each LA group; each LA group includes multiple LAs, and the cloud computing devices associated with the same LA group communicate based on the LAs of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LAs of the corresponding LA groups and the associated LCs;

[0231] The cloud management device 1002 is configured to, when receiving a second acquisition request sent by any cloud computing device, determine the identifier of the LA group associated with the any cloud computing device based on the association relationship between the device identifiers of each cloud computing device and the identifiers of the LA groups pre-configured;

[0232] The cloud management device 1002 is configured to determine LA group identifier information based on the identifier of the LA group associated with the any cloud computing device, and return the LA group identifier information to the any cloud computing device;

[0233] The target client 1001 is configured to send a first acquisition request to each cloud computing device associated with the target client, and the first acquisition request is used to request to acquire the LA group identifier information corresponding to the cloud computing device;

[0234] The any cloud computing device 1003 is configured to receive the first acquisition request sent by the target client, and based on the first acquisition request, return the LA group identifier information of the any cloud computing device to the target client, and the LA group identifier information is determined based on the identifier of the LA group associated with the cloud computing device;

[0235] The target client 1001 is configured to receive the LA group identifier information returned by each cloud computing device associated with the target client;

[0236] The target client 1001 is used to sort each cloud computing device associated with the target client based on the LA group identification information, and calculate using each cloud computing device associated with the target client in the sorted order.

[0237] The data processing method provided by this application. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each cloud computing device associated with the same LA group communicates based on the LA of that LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC. The target client sends a first acquisition request to the associated cloud computing devices to obtain the LA group identification information of each cloud computing device associated with the target client. And, sort each cloud computing device based on the LA group identification information, and calculate using each cloud computing device in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to calculate using the cloud computing devices of the same LA group, greatly reducing the communication between cloud computing devices across LA groups, minimizing traffic detours, reducing network latency, and reducing traffic consumption.

[0238] Figure 11 It is a schematic structural diagram of a data processing device provided by an embodiment of this application. This device is applied to the target client, and the target client is the application client of the computing cluster. The computing cluster includes multiple aggregation devices LCs, multiple access device LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. Each LA group includes multiple LAs. Each cloud computing device associated with the same LA group communicates based on the LA of that LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC. As Figure 11 shown, this device includes:

[0239] The first sending module 1101 is used to send a first acquisition request to each cloud computing device associated with the target client, and this first acquisition request is used to request to obtain the LA group identification information corresponding to the cloud computing device.

[0240] The first receiving module 1102 is used to receive the LA group identification information returned by each cloud computing device associated with the target client, and this LA group identification information is determined based on the identification of the LA group associated with the cloud computing device.

[0241] The sorting module 1103 is used to sort each cloud computing device associated with the target client based on the LA group identification information, and calculate using each cloud computing device associated with the target client in the sorted order.

[0242] In a possible implementation, the data processing device of the cloud management device applied to the computing cluster includes an LA group identification information determination module. For any cloud computing device, when the LA group identification information determination module is used to determine the LA group identification information corresponding to the cloud computing device, it is specifically used for:

[0243] Perform transformation processing on the identification of the LA group associated with the cloud computing device to obtain an implicit indication value of the identification of the LA group, and use the implicit indication value as the LA group identification information corresponding to the cloud computing device.

[0244] In a possible implementation, the device further includes:

[0245] A second sending module, configured to send an allocation request to the cloud management device of the computing cluster, where the allocation request is used to request to allocate a cloud computing device for a target client;

[0246] A third receiving module, configured to receive the device identifiers of each cloud computing device returned by the cloud management device, and obtain each cloud computing device associated with the target client based on the device identifiers of each cloud computing device;

[0247] Among them, the data processing device applied to the cloud management device includes a cloud computing device determination module. When the cloud computing device determination module is used to determine each cloud computing device associated with the target client, it includes:

[0248] A first determination unit, configured to determine the number of idle devices of the cloud computing devices in the idle state associated with each LA group;

[0249] An allocation unit, configured to determine the allocation order of each LA group based on the number of idle devices of each LA group, and sequentially select cloud computing devices from the cloud computing devices associated with at least one LA group according to the allocation order of each LA group, and allocate the cloud computing devices to the target client.

[0250] In a possible implementation, the allocation request is a new request, and the new request is used to request to create a user computing cluster for the target client:

[0251] The allocation unit is used for:

[0252] Based on the number of idle devices of each LA group, perform a descending order arrangement on each LA group to obtain a first LA group sequence;

[0253] Take the arrangement order of each LA group in the first LA group sequence as the allocation order, sequentially select cloud computing devices from the cloud computing devices associated with at least one LA group, and construct a user computing cluster for the target client based on the selected cloud computing devices.

[0254] In a possible implementation, when the allocation unit selects cloud computing devices from at least one cloud computing device associated with an LA group in the order of arrangement of each LA group in the first LA group sequence, it is specifically configured to:

[0255] If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are sorted in ascending order based on the number of clients corresponding to each LA group in the at least two LA groups, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group, to obtain a second LA group sequence;

[0256] Cloud computing devices are sequentially selected from at least one cloud computing device associated with an LA group in the order of arrangement of each LA group in the second LA group sequence.

[0257] In a possible implementation, when the allocation unit sorts each LA group in descending order based on the number of idle devices of each LA group to obtain a first LA group sequence, it is specifically configured to:

[0258] Search for idle LA groups in each LA group, where each cloud computing device associated with the idle LA group is in an idle state;

[0259] If there are idle LA groups in each LA group, cloud computing devices are selected from the cloud computing devices associated with the idle LA groups, and a user computing cluster is constructed for the target client based on the selected cloud computing devices;

[0260] If there are no idle LA groups in each LA group, each LA group is sorted in descending order based on the number of idle devices of each LA group to obtain a first LA group sequence.

[0261] In a possible implementation, the allocation request is an expansion request, and the expansion request is used to request to add cloud computing devices to the user computing cluster of the target client:

[0262] The allocation unit is configured to:

[0263] Each associated LA group is sorted in ascending order based on the number of idle devices of the associated LA groups in each LA group, where each cloud computing device associated with the associated LA group includes the cloud computing devices already allocated to the target client, to obtain a third LA group sequence;

[0264] Cloud computing devices are sequentially selected from at least one associated LA group and allocated to the target client in the order of arrangement of each associated LA group in the third LA group sequence.

[0265] The data processing method provided by this application. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. The cloud computing devices associated with the same LA group communicate based on the LA of that LA group, and the cloud computing devices associated with different LA groups communicate based on the LA of the corresponding LA group and the associated LC. The target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client. Moreover, the cloud computing devices are sorted based on the LA group identification information, and the cloud computing devices are used for calculation in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices of the same LA group for calculation, greatly reducing the communication across LA groups between cloud computing devices, minimizing traffic detours, reducing network latency, and reducing traffic consumption.

[0266] Figure 12 It is a schematic structural diagram of a data processing device provided by an embodiment of this application. This device is applied to any cloud computing device associated with the target client in the computing cluster. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0267] Among them, each LA group includes multiple LAs. The cloud computing devices associated with the same LA group communicate based on the LA of that LA group, and the cloud computing devices associated with different LA groups communicate based on the LA of the corresponding LA group and the associated LC;

[0268] As Figure 12 shown, this device includes:

[0269] A second receiving module 1201, configured to receive a first acquisition request sent by the target client, where the first acquisition request is used to request to obtain the LA group identification information corresponding to this any cloud computing device;

[0270] A first returning module 1202, configured to, based on the first acquisition request, return the LA group identification information of this any cloud computing device to the target client. The LA group identification information is determined based on the identification of the LA group associated with the cloud computing device, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client in the sorted order for calculation.

[0271] In a possible implementation manner, this device further includes:

[0272] A third sending module, configured to send a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to obtain the LA group identification information of this any cloud computing device;

[0273] A fourth receiving module, configured to receive the LA group identification information returned by the cloud management device based on the second acquisition request.

[0274] The data processing method provided in this application, the computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group; the cloud computing devices associated with the same LA group communicate based on the LA of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LA of the corresponding LA group and the associated LC; the target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client; and, sort each cloud computing device based on the LA group identification information, so as to use each cloud computing device for calculation in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices of the same LA group for calculation, greatly reducing the communication across LA groups between cloud computing devices, minimizing traffic detours, reducing network latency, and reducing traffic consumption.

[0275] Figure 13 It is a schematic structural diagram of a data processing device provided in an embodiment of this application. This application provides that the device is applied to a cloud management device of a computing cluster, and the computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group;

[0276] Among them, each LA group includes multiple LAs, and the cloud computing devices associated with the same LA group communicate based on the LA of the LA group, and the cloud computing devices associated with different LA groups communicate based on the LA of the corresponding LA group and the associated LC;

[0277] As Figure 13 shown, the device includes:

[0278] A first determination module 1301, configured to determine the identification of the LA group associated with any one cloud computing device based on the association relationship between the device identification of each cloud computing device and the identification of the LA group pre-configured when receiving a second acquisition request sent by any one cloud computing device;

[0279] A second determination module 1302, configured to determine LA group identification information based on the identification of the LA group associated with any one cloud computing device, and return the LA group identification information to any one cloud computing device, so that any one cloud computing device performs the following steps:

[0280] When receiving a first acquisition request sent by an associated target client segment, return the LA group identification information of any one of the cloud computing devices to the target client, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client in the sorted order for computing.

[0281] In a possible implementation manner, when determining the LA group identification information based on the identification of the LA group associated with any one of the cloud computing devices, the second determination module is specifically configured to:

[0282] Perform a transformation process on the identification of the LA group associated with the cloud computing device to obtain an implicit indication value of the identification of the LA group, and use the implicit indication value as the LA group identification information corresponding to the cloud computing device.

[0283] In a possible implementation manner, the apparatus further includes:

[0284] An idle device number determination module, configured to determine the number of idle devices of the cloud computing devices in an idle state associated with each LA group when receiving an allocation request from a target client, where the allocation request is used to request to allocate cloud computing devices to the target client;

[0285] An allocation module, configured to determine the allocation order of each LA group based on the number of idle devices of each LA group, and sequentially select cloud computing devices from the cloud computing devices associated with at least one LA group according to the allocation order of each LA group and allocate them to the target client;

[0286] A second return module, configured to return the device identifiers of the allocated cloud computing devices to the target client.

[0287] The apparatus according to the embodiments of the present application can execute the method provided by the embodiments of the present application, and the implementation principle is similar. The actions performed by each module in the apparatus according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function descriptions of each module of the apparatus, reference can specifically be made to the descriptions in the corresponding method shown above, and details are not described herein again.

[0288] The data processing method provided by this application. The computing cluster includes multiple LCs, multiple LA groups associated with each LC, and multiple cloud computing devices associated with each LA group. The cloud computing devices associated with the same LA group communicate based on the LA of that LA group, and the cloud computing devices associated with different LA groups communicate based on the LA of the corresponding LA group and the associated LC. The target client sends a first acquisition request to the associated cloud computing device to obtain the LA group identification information of each cloud computing device associated with the target client. And, based on the LA group identification information, the cloud computing devices are sorted, and the cloud computing devices are used for calculation in the sorted order. Since the cloud computing devices of the same LA group are concentrated together after sorting, it is convenient to use the cloud computing devices of the same LA group for calculation, greatly reducing the communication across LA groups between cloud computing devices, minimizing traffic detours, reducing network latency, and reducing traffic consumption.

[0289] Figure 14 This is a schematic structural diagram of an electronic device provided in an embodiment of this application. As Figure 14 shown, the electronic device includes: a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the above method.

[0290] In an alternative embodiment, an electronic device is provided. As Figure 14 shown, Figure 14 the electronic device 1400 shown includes: a processor 1401 and a memory 1403. Among them, the processor 1401 and the memory 1403 are connected, such as connected through a bus 1402. Optionally, the electronic device 1400 may further include a transceiver 1404. The transceiver 1404 may be used for data interaction between this electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 1404 is not limited to one, and the structure of this electronic device 1400 does not constitute a limitation to the embodiments of this application.

[0291] The processor 1401 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 1401 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0292] The bus 1402 may include a path for transmitting information between the above components. The bus 1402 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 1402 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 14 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0293] The memory 1403 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0294] The memory 1403 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 1401 to execute. The processor 1401 is used to execute the computer program stored in the memory 1403 to implement the steps shown in the foregoing method embodiments.

[0295] Among them, the electronic device includes but is not limited to: servers, terminals, cloud computing devices, etc.

[0296] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0297] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0298] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0299] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "the" used here may also include the plural forms. The terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by this technology field.

[0300] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.

[0301] It should be understood that although the flowcharts in the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0302] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method is applied to a target client, which is an application client of a computing cluster; the computing cluster includes a plurality of convergence devices LC, a plurality of access device LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The method comprises: Sending a first acquisition request to each cloud computing device associated with the target client, wherein the first acquisition request is used to request to obtain LA group identification information corresponding to the cloud computing device; Receiving LA group identification information returned by each cloud computing device associated with the target client, where the LA group identification information returned by each cloud computing device is used to indicate the identification of the LA group associated with the cloud computing device, and the LA group identification information of cloud computing devices in different LA groups is different; The cloud computing devices associated with the target client are sorted based on the LA group identification information, so as to use the cloud computing devices associated with the target client for calculation in the sorted order.

2. The method according to claim 1, characterized in that For any cloud computing device, the LA group identification information corresponding to the cloud computing device is obtained by the following steps: The identifier of the LA group associated with the cloud computing device is transformed to obtain an implicit indication value of the identifier of the LA group, and the implicit indication value is used as the LA group identification information corresponding to the cloud computing device.

3. The method according to claim 1 or 2, characterized in that: For any cloud computing device, the LA group identification information corresponding to the cloud computing device is obtained by the cloud computing device in the following manner: Sending a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to obtain the LA group identification information of the cloud computing device; Receive LA group identification information returned by the cloud management device based on the second acquisition request, wherein the LA group identification information is determined by the cloud management device based on an identification of the LA group associated with the cloud computing device.

4. The method according to claim 1, characterized in that The method further comprises: Sending an allocation request to a cloud management device of the computing cluster, wherein the allocation request is used to request allocation of a cloud computing device to a target client; Receiving device identifications of various cloud computing devices returned by the cloud management device, and obtaining various cloud computing devices associated with the target client based on the device identifications of various cloud computing devices; The cloud computing devices associated with the target client are determined by the cloud management device through the following steps: Determine the number of idle devices of the cloud computing devices in an idle state associated with each LA group; Based on the number of idle devices in each LA group, the allocation order of each LA group is determined, and according to the allocation order of each LA group, cloud computing devices are selected from the cloud computing devices associated with at least one LA group and allocated to the target client.

5. The method according to claim 4, characterized in that The allocation request is a new creation request, and the new creation request is used to request to create a new user computing cluster for the target client: The method of determining the allocation order of each LA group based on the number of idle devices in each LA group, and selecting a cloud computing device from cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client includes: Based on the number of idle devices in each LA group, the LA groups are arranged in descending order to obtain a first LA group sequence; Taking the arrangement order of each LA group in the first LA group sequence as the allocation order, cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, and a user computing cluster is constructed for the target client based on the selected cloud computing devices.

6. The method according to claim 5, characterized in that The step of selecting cloud computing devices from at least one LA group-associated cloud computing device in sequence based on the arrangement order of the LA groups in the first LA group sequence as the allocation order includes: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are arranged in ascending order based on the number of clients corresponding to each LA group in the at least two LA groups to obtain a second LA group sequence, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group; The arrangement order of the LA groups in the second LA group sequence is used as the allocation order, and cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn.

7. The method according to claim 5 or 6, characterized in that: The method of arranging each LA group in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence includes: Find an idle LA group in each LA group, where each cloud computing device associated with the idle LA group is in an idle state; If there is an idle LA group in each LA group, a cloud computing device is selected from the cloud computing devices associated with the idle LA group, and a user computing cluster is constructed for the target client based on the selected cloud computing device; If there is no idle LA group in each LA group, the LA groups are arranged in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence.

8. The method according to claim 4, characterized in that The allocation request is a capacity expansion request, and the capacity expansion request is used to request to add a cloud computing device to the user computing cluster of the target client: The method of determining the allocation order of each LA group based on the number of idle devices in each LA group, and selecting a cloud computing device from cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client includes: Arrange each associated LA group in ascending order based on the number of idle devices of the associated LA group in each LA group to obtain a third LA group sequence, wherein each cloud computing device associated with the associated LA group includes the cloud computing device allocated to the target client; The arrangement order of each associated LA group in the third LA group sequence is used as the allocation order, and cloud computing devices are selected from at least one associated LA group in turn and allocated to the target client.

9. The method according to claim 1, characterized in that: The first acquisition request is sent by the target client to each cloud computing device associated with the target client in response to triggering of a target event, and the target event includes at least one of the following: Receive a device identifier of a cloud computing device assigned by a cloud management device of the computing cluster; Receiving a device identification of a changed cloud computing device sent by a cloud management device of the computing cluster; It is detected that the current time reaches the target time period.

10. A data processing method, characterized in that: The method is applied to any cloud computing device associated with a target client in a computing cluster, the computing cluster comprising a plurality of LCs, a plurality of LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The method comprises: Receive a first acquisition request sent by a target client, where the first acquisition request is used to request to obtain LA group identification information corresponding to any one of the cloud computing devices; Based on the first acquisition request, the LA group identification information of any cloud computing device is returned to the target client, and the LA group identification information returned by each cloud computing device is used to indicate the identification of the LA group associated with the cloud computing device. The LA group identification information of cloud computing devices in different LA groups is different, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client for calculation in the sorted order.

11. The method according to claim 10, characterized in that Before receiving the first acquisition request sent by the target client, the method further includes: Sending a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to obtain the LA group identification information of any cloud computing device; Receive LA group identification information returned by the cloud management device based on the second acquisition request.

12. The method according to claim 11, characterized in that The LA group identification information corresponding to the cloud computing device is obtained by the cloud management device through the following steps: The identifier of the LA group associated with the cloud computing device is transformed to obtain an implicit indication value of the identifier of the LA group, and the implicit indication value is used as the LA group identification information corresponding to the cloud computing device.

13. A data processing method, characterized in that: The method is applied to a cloud management device of a computing cluster, the computing cluster comprising a plurality of LCs, a plurality of LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The method comprises: When receiving a second acquisition request sent by any cloud computing device, determining the identifier of the LA group associated with any cloud computing device based on the pre-configured association relationship between the device identifiers of each cloud computing device and the identifier of the LA group; Based on the identifier of the LA group associated with any one of the cloud computing devices, determine LA group identifier information, and return the LA group identifier information to any one of the cloud computing devices, so that any one of the cloud computing devices performs the following steps: Upon receiving a first acquisition request sent by the associated target client, the LA group identification information of any cloud computing device is returned to the target client, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client for calculation in the sorted order; the LA group identification information is used to indicate the identification of the LA group associated with any cloud computing device, and the LA group identification information of cloud computing devices in different LA groups is different.

14. The method according to claim 13, characterized in that The method further comprises: When receiving an allocation request from a target client, determining the number of idle devices of the idle cloud computing devices associated with each LA group, wherein the allocation request is used to request allocation of a cloud computing device to the target client; Based on the number of idle devices in each LA group, determine the allocation order of each LA group, and select cloud computing devices from the cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client; The device identification of each allocated cloud computing device is returned to the target client.

15. The method according to claim 14, characterized in that The allocation request is a new creation request, and the new creation request is used to request to create a new user computing cluster for the target client: The method of determining the allocation order of each LA group based on the number of idle devices in each LA group, and selecting a cloud computing device from cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client includes: Based on the number of idle devices in each LA group, the LA groups are arranged in descending order to obtain a first LA group sequence; Taking the arrangement order of each LA group in the first LA group sequence as the allocation order, cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, and a user computing cluster is constructed for the target client based on the selected cloud computing devices.

16. The method according to claim 15, characterized in that The step of selecting cloud computing devices from at least one LA group-associated cloud computing device in sequence based on the arrangement order of the LA groups in the first LA group sequence as the allocation order includes: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are arranged in ascending order based on the number of clients corresponding to each LA group in the at least two LA groups to obtain a second LA group sequence, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group; The arrangement order of the LA groups in the second LA group sequence is used as the allocation order, and cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn.

17. The method according to claim 15 or 16, characterized in that The method of arranging each LA group in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence includes: Find an idle LA group in each LA group, where each cloud computing device associated with the idle LA group is in an idle state; If there is an idle LA group in each LA group, a cloud computing device is selected from the cloud computing devices associated with the idle LA group, and a user computing cluster is constructed for the target client based on the selected cloud computing device; If there is no idle LA group in each LA group, the LA groups are arranged in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence.

18. The method according to claim 14, characterized in that The allocation request is a capacity expansion request, and the capacity expansion request is used to request to add a cloud computing device to the user computing cluster of the target client: The method of determining the allocation order of each LA group based on the number of idle devices in each LA group, and selecting a cloud computing device from cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client includes: Arrange each associated LA group in ascending order based on the number of idle devices of the associated LA group in each LA group to obtain a third LA group sequence, wherein each cloud computing device associated with the associated LA group includes the cloud computing device allocated to the target client; The arrangement order of each associated LA group in the third LA group sequence is used as the allocation order, and cloud computing devices are selected from at least one associated LA group in turn and allocated to the target client.

19. The method according to claim 13, characterized in that When receiving a second acquisition request sent by any cloud computing device, determining the identifier of the LA group associated with any cloud computing device based on the association relationship between the pre-configured device identifiers of each cloud computing device and the identifier of the LA group, the method further includes: The device identification of each cloud computing device and the identification of the associated LA group are collected, and the association relationship between the device identification of each cloud computing device and the identification of the LA group is stored.

20. A data processing device, characterized in that: The device is applied to a target client, which is an application client of a computing cluster; the computing cluster includes a plurality of convergence devices LC, a plurality of access device LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The device comprises: A first sending module, configured to send a first acquisition request to each cloud computing device associated with the target client, wherein the first acquisition request is used to request to acquire LA group identification information corresponding to the cloud computing device; A first receiving module is used to receive LA group identification information returned by each cloud computing device associated with the target client, wherein the LA group identification information returned by each cloud computing device is used to indicate the identification of the LA group associated with the cloud computing device, and the LA group identification information of cloud computing devices in different LA groups is different; The sorting module is used to sort the cloud computing devices associated with the target client based on the LA group identification information, so as to use the cloud computing devices associated with the target client for calculation in the sorted order.

21. The device according to claim 20, characterized in that For any cloud computing device, the LA group identification information corresponding to the cloud computing device is obtained by the following steps: The identifier of the LA group associated with the cloud computing device is transformed to obtain an implicit indication value of the identifier of the LA group, and the implicit indication value is used as the LA group identification information corresponding to the cloud computing device.

22. The device according to claim 20 or 21, characterized in that For any cloud computing device, the LA group identification information corresponding to the cloud computing device is obtained by the cloud computing device in the following manner: Sending a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to obtain the LA group identification information of the cloud computing device; Receive LA group identification information returned by the cloud management device based on the second acquisition request, wherein the LA group identification information is determined by the cloud management device based on an identification of the LA group associated with the cloud computing device.

23. The device according to claim 20, characterized in that The device also includes: A second sending module is used to send an allocation request to the cloud management device of the computing cluster, wherein the allocation request is used to request allocation of a cloud computing device to a target client; A third receiving module is used to receive the device identification of each cloud computing device returned by the cloud management device, and obtain each cloud computing device associated with the target client based on the device identification of each cloud computing device; The cloud computing devices associated with the target client are determined by the cloud management device through the following steps: Determine the number of idle devices of the cloud computing devices in an idle state associated with each LA group; Based on the number of idle devices in each LA group, the allocation order of each LA group is determined, and according to the allocation order of each LA group, cloud computing devices are selected from the cloud computing devices associated with at least one LA group and allocated to the target client.

24. The device according to claim 23, characterized in that The allocation request is a new creation request, and the new creation request is used to request to create a new user computing cluster for the target client: The cloud computing device is allocated to the target client by the cloud management device through the following steps: Based on the number of idle devices in each LA group, the LA groups are arranged in descending order to obtain a first LA group sequence; Taking the arrangement order of each LA group in the first LA group sequence as the allocation order, cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, and a user computing cluster is constructed for the target client based on the selected cloud computing devices.

25. The device according to claim 24, characterized in that The cloud computing device is selected by the cloud management device from at least one LA group associated cloud computing device through the following steps: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are arranged in ascending order based on the number of clients corresponding to each LA group in the at least two LA groups to obtain a second LA group sequence, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group; The arrangement order of the LA groups in the second LA group sequence is used as the allocation order, and cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn.

26. The device according to claim 24 or 25, characterized in that The first LA group sequence is obtained by the cloud management device through the following steps: Find an idle LA group in each LA group, where each cloud computing device associated with the idle LA group is in an idle state; If there is an idle LA group in each LA group, a cloud computing device is selected from the cloud computing devices associated with the idle LA group, and a user computing cluster is constructed for the target client based on the selected cloud computing device; If there is no idle LA group in each LA group, the LA groups are arranged in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence.

27. The device according to claim 23, characterized in that The allocation request is a capacity expansion request, and the capacity expansion request is used to request to add a cloud computing device to the user computing cluster of the target client: The cloud computing device is selected by the cloud management device from at least one LA group associated cloud computing device through the following steps: Arrange each associated LA group in ascending order based on the number of idle devices of the associated LA group in each LA group to obtain a third LA group sequence, wherein each cloud computing device associated with the associated LA group includes the cloud computing device allocated to the target client; The arrangement order of each associated LA group in the third LA group sequence is used as the allocation order, and cloud computing devices are selected from at least one associated LA group in turn and allocated to the target client.

28. The device according to claim 20, characterized in that The first acquisition request is sent by the target client to each cloud computing device associated with the target client in response to triggering of a target event, and the target event includes at least one of the following: Receive a device identifier of a cloud computing device assigned by a cloud management device of the computing cluster; Receiving a device identification of a changed cloud computing device sent by a cloud management device of the computing cluster; It is detected that the current time reaches the target time period.

29. A data processing device, characterized in that: The apparatus is applied to any cloud computing device associated with a target client in a computing cluster, wherein the computing cluster includes a plurality of LCs, a plurality of LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The device comprises: A second receiving module is used to receive a first acquisition request sent by a target client, where the first acquisition request is used to request to obtain LA group identification information corresponding to any cloud computing device; A first return module is used to return the LA group identification information of any cloud computing device to the target client based on the first acquisition request, the LA group identification information returned by each cloud computing device is used to indicate the identification of the LA group associated with the cloud computing device, and the LA group identification information of cloud computing devices in different LA groups is different, so that the target client can sort the cloud computing devices associated with the target client based on the LA group identification information, and use the cloud computing devices associated with the target client for calculation in the sorted order.

30. The device according to claim 29, characterized in that The device also includes: A third sending module is used to send a second acquisition request to the cloud management device of the computing cluster, where the second acquisition request is used to request to obtain the LA group identification information of any cloud computing device; The fourth receiving module is used to receive the LA group identification information returned by the cloud management device based on the second acquisition request.

31. The device according to claim 30, characterized in that The LA group identification information corresponding to the cloud computing device is obtained by the cloud management device through the following steps: The identifier of the LA group associated with the cloud computing device is transformed to obtain an implicit indication value of the identifier of the LA group, and the implicit indication value is used as the LA group identification information corresponding to the cloud computing device.

32. A data processing device, characterized in that: The device is applied to a cloud management device of a computing cluster, wherein the computing cluster includes a plurality of LCs, a plurality of LA groups associated with each LC, and a plurality of cloud computing devices associated with each LA group; Each LA group includes multiple LAs, and each cloud computing device associated with the same LA group communicates based on the LA of the LA group, and each cloud computing device associated with different LA groups communicates based on the LA of the corresponding LA group and the associated LC; The device comprises: A first determination module is configured to determine, when receiving a second acquisition request sent by any cloud computing device, the identifier of the LA group associated with any cloud computing device based on the association relationship between the pre-configured device identifiers of each cloud computing device and the identifier of the LA group; The second determining module is used to determine LA group identification information based on the identification of the LA group associated with any one of the cloud computing devices, and return the LA group identification information to any one of the cloud computing devices, so that any one of the cloud computing devices performs the following steps: Upon receiving a first acquisition request sent by the associated target client, the LA group identification information of any cloud computing device is returned to the target client, so that the target client sorts the cloud computing devices associated with the target client based on the LA group identification information, and uses the cloud computing devices associated with the target client for calculation in the sorted order; the LA group identification information is used to indicate the identification of the LA group associated with any cloud computing device, and the LA group identification information of cloud computing devices in different LA groups is different.

33. The device according to claim 32, characterized in that The device also includes: An idle device number determination module, configured to determine the number of idle cloud computing devices associated with each LA group when receiving an allocation request from a target client, wherein the allocation request is used to request allocation of a cloud computing device to the target client; An allocation module, configured to determine an allocation order of each LA group based on the number of idle devices in each LA group, and select a cloud computing device from cloud computing devices associated with at least one LA group in turn according to the allocation order of each LA group to allocate to the target client; The second returning module is used to return the allocated device identification of each cloud computing device to the target client.

34. The device according to claim 33, characterized in that The allocation request is a new creation request, and the new creation request is used to request to create a new user computing cluster for the target client: The allocation module determines the allocation order of each LA group based on the number of idle devices in each LA group, and selects cloud computing devices from at least one LA group-associated cloud computing device to allocate to the target client in accordance with the allocation order of each LA group, specifically for: Based on the number of idle devices in each LA group, the LA groups are arranged in descending order to obtain a first LA group sequence; Taking the arrangement order of each LA group in the first LA group sequence as the allocation order, cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn, and a user computing cluster is constructed for the target client based on the selected cloud computing devices.

35. The device according to claim 34, characterized in that The allocation module selects cloud computing devices from at least one LA group-associated cloud computing device in sequence based on the arrangement order of the LA groups in the first LA group sequence as the allocation order, specifically for: If there are at least two LA groups with the same number of idle devices in the first LA group sequence, the at least two LA groups in the first LA group sequence are arranged in ascending order based on the number of clients corresponding to each LA group in the at least two LA groups to obtain a second LA group sequence, where the number of clients refers to the number of clients associated with each cloud computing device associated with the LA group; The arrangement order of the LA groups in the second LA group sequence is used as the allocation order, and cloud computing devices are selected from the cloud computing devices associated with at least one LA group in turn.

36. The device according to claim 34 or 35, characterized in that The allocation module is specifically used to arrange each LA group in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence: Find an idle LA group in each LA group, where each cloud computing device associated with the idle LA group is in an idle state; If there is an idle LA group in each LA group, a cloud computing device is selected from the cloud computing devices associated with the idle LA group, and a user computing cluster is constructed for the target client based on the selected cloud computing device; If there is no idle LA group in each LA group, the LA groups are arranged in descending order based on the number of idle devices in each LA group to obtain a first LA group sequence.

37. The device according to claim 33, characterized in that The allocation request is a capacity expansion request, and the capacity expansion request is used to request to add a cloud computing device to the user computing cluster of the target client: The allocation module determines the allocation order of each LA group based on the number of idle devices in each LA group, and selects cloud computing devices from at least one LA group-associated cloud computing device to allocate to the target client in accordance with the allocation order of each LA group, specifically for: Arrange each associated LA group in ascending order based on the number of idle devices of the associated LA group in each LA group to obtain a third LA group sequence, wherein each cloud computing device associated with the associated LA group includes the cloud computing device allocated to the target client; The arrangement order of each associated LA group in the third LA group sequence is used as the allocation order, and cloud computing devices are selected from at least one associated LA group in turn and allocated to the target client.

38. The device according to claim 37, characterized in that The first determining module is further used for: The device identification of each cloud computing device and the identification of the associated LA group are collected, and the association relationship between the device identification of each cloud computing device and the identification of the LA group is stored.

39. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 19.

40. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 19 are implemented.

41. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 19 are implemented.

Citation Information

Patent Citations

  • Method, node manager and system for load balancing in cloud computing system

    CN102624916A

  • Task execution method and device on virtual machine, storage medium and electronic equipment

    CN116501487A

  • Server allocation method and device, electronic equipment and computer readable storage medium

    CN118200321A