Cloud computing service system, cloud computing service providing method, storage medium and product

By adopting the combination of network domain isolation and scheduling systems in the cloud computing service system, the problem of complex and cost of cloud computing service system deployment is solved, efficient utilization and flexible management of resources are achieved, and the computing needs of users with high complexity and large data volume are met.

CN120343030AActive Publication Date: 2025-07-18ALIBABA CLOUD COMPUTING CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510551602.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing cloud computing service system is deployed in complex ways and has high cost, which is difficult to meet users' ever-increasing task processing needs. It is difficult to form a scale for resource dispersion, expanding is difficult, and resource waste is serious.

Method used

The network domain isolation method is adopted to connect multiple control clusters to different network domains of the same computing cluster, forward data through the first network switch system, and introduce a scheduling system for flexible switching of servers and resource configuration to avoid frequent relocation and redundant construction.

Benefits of technology

It reduces the deployment complexity and cost of cloud computing service systems, improves resource availability and allocation flexibility, avoids resource waste, and realizes large-scale and efficient computing services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343030A_ABST
    Figure CN120343030A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cloud computing service system, a cloud computing service providing method, a storage medium and a product. The system comprises a computing cluster, a first network switch system and a plurality of control clusters, the computing cluster comprises a plurality of servers; a first network provided by the first network switch system is divided into a plurality of first network domains; the first network switch system is used for accessing the plurality of control clusters to different first network domains based on the first network domain configuration information, and accessing at least one server corresponding to the plurality of control clusters to the respective corresponding first network domains, and a data forwarding service is provided between the control cluster and the server which are accessed to the same first network domain. According to the technical scheme provided by the embodiment of the invention, the complexity and deployment cost of a cloud computing server system deployment mode are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to a cloud computing service system, a cloud computing service providing method, a storage medium, and a product. Background Art

[0002] With the development of computer technologies, users' demands for data processing speed and computing power are also continuously increasing. Traditional single-machine servers are difficult to meet such demands. Therefore, building a computing cluster composed of one or more high-performance servers has become a solution to address such demands. However, since building such a computing cluster is costly and the wiring is complex, for ordinary users, renting a cloud-based service system (e.g., a cloud computing service system) to perform complex task processing has become an efficient and economical solution.

[0003] Currently, a cloud computing service system usually includes a computing cluster for performing complex task processing and a control cluster for managing and accessing the computing cluster. Multiple such cloud computing service systems can be built to provide cloud computing services for different users. However, this deployment method is relatively complex and costly, and thus urgently needs to be improved. Summary of the Invention

[0004] Embodiments of the present application provide a cloud computing service system and method, a storage medium, and a product, which are used to solve the problems of complex deployment method and high cost of existing cloud computing service systems.

[0005] In a first aspect, a cloud computing service system provided in an embodiment of the present application includes: a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; The first network switch system is configured to, based on first network domain configuration information, connect the multiple control clusters to different first network domains, connect at least one server corresponding to each of the multiple control clusters to the respective corresponding first network domains, and provide data forwarding services between the control clusters and the servers connected to the same first network domain.

[0006] In a second aspect, a cloud computing service providing method provided in an embodiment of the present application is applied to the first network switch system of a cloud computing service system. The cloud computing service system further includes a computing cluster and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the method includes: Based on the first network domain configuration information, connect the multiple control clusters to different first network domains, and connect at least one server corresponding to each of the multiple control clusters to the corresponding first network domain respectively; Provide data forwarding services between the control clusters and servers connected to the same first network domain.

[0007] In a third aspect, an embodiment of the present application provides a cloud computing service providing method, which is applied to a second network switch system of a cloud computing service system. The cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to the corresponding first network domain respectively; and data forwarding is supported between the control clusters and servers connected to the same first network domain. The method includes: Based on the second network domain configuration information, connect at least one server corresponding to the same control cluster or the same user to the second network domain corresponding to the at least one server; the second network domain is obtained by dividing the second network provided by the second network switch system; Provide data forwarding services between different servers connected to the same second network domain.

[0008] In a fourth aspect, an embodiment of the present application provides a cloud computing service providing method, which is applied to a scheduling system of a cloud computing service system. The cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to the corresponding first network domain respectively; data forwarding is supported between the control clusters and servers connected to the same first network domain. The method includes: Detect a switching operation for a target server, and determine the target control cluster requesting the switch; Based on the target control cluster, send a switching request to the first network switch system, so that the first network switch system responds to the switching request, determines the target control cluster, and switches the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0009] In a fifth aspect, the present application provides a computer-readable storage medium storing a computer program, and the computer program is executed by a computer to implement the cloud computing service providing method according to any one of the second to fourth aspects.

[0010] In a sixth aspect, the present application provides a computer program product storing a computer program, which when executed by a computer implements the cloud computing service providing method according to any one of the second to fourth aspects.

[0011] In the cloud computing service system provided by the solution of the embodiment of the present application, it includes a computing cluster composed of multiple servers, a first network switch system, and multiple control clusters; the first network provided by the first network switch system is divided into multiple first network domains; the first network switch system is used to connect multiple control clusters to different first network domains based on the first network domain configuration information, and connect at least one server corresponding to each of the multiple control clusters to their respective corresponding first network domains, so as to provide data forwarding services between the control clusters and the servers connected to the same first network domain. The embodiment of the present application can implement providing cloud computing services for multiple control clusters simultaneously by deploying only one computing cluster, greatly reducing the complexity and deployment cost of the deployment method. And the cloud service provider deploys its limited number of servers in one computing cluster, and through the method of network domain isolation, enables multiple control clusters to share the same computing cluster, which can not only realize dispersing server resources to provide cloud computing services for multiple control clusters simultaneously, but also realize that one control cluster exclusively uses all the servers in the computing cluster to provide users with large-scale, high-complexity, and high-data-volume computing services, avoiding the problem that it is difficult to form a scale due to the dispersion of the system and thus unable to meet the continuously increasing task processing requirements. In addition, when a new user or a processing task appears, the present solution only needs to redeploy one control cluster and then configure it into a new network domain of the existing cloud computing service system. Since the construction cost and the difficulty of environment setup of the control cluster are much lower than those of the computing cluster, the solution of the embodiment of the present application greatly reduces the construction cost and difficulty compared with redeploying a set of cloud computing service system including a computing cluster and a control cluster currently. In addition, when the task complexity or data volume of the user corresponding to any control cluster increases, only the number of servers connected to the network domain corresponding to the control cluster needs to be adjusted, and this process does not require expanding the computing cluster, reducing the adjustment cost and difficulty.

[0012] In addition, for the situation where any control cluster fails or is upgraded, resulting in the idleness of its corresponding servers, the present solution can also configure the idle servers for other control clusters to use through the configuration of the network domain, avoiding waste of resources. And during this process, there is no need to perform operations such as relocating the servers and re-wiring, further improving the availability of server resources and the flexibility of resource allocation on the premise of low cost and high efficiency.

[0013] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. Brief Description of the Drawings

[0014] The drawings described herein are provided to further understand the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings: Figure 1 The structural schematic diagram of an embodiment of a cloud computing service system provided by the present application is shown.

[0015] Figure 2 The structural schematic diagram of another embodiment of a cloud computing service system provided by the present application is shown.

[0016] Figure 3 The schematic diagram of the effect of an embodiment of a scheduling system display interface provided by the present application is shown.

[0017] Figure 4 The structural schematic diagram of yet another embodiment of a cloud computing service system provided by the present application is shown.

[0018] Figure 5 The flowchart of an embodiment of a cloud computing service providing method provided by the present application is shown.

[0019] Figure 6 The flowchart of another embodiment of a cloud computing service providing method provided by the present application is shown.

[0020] Figure 7 The flowchart of yet another embodiment of a cloud computing service providing method provided by the present application is shown.

[0021] Figure 8 The structural diagram of a cloud computing service system in an actual scenario is shown.

[0022] Figure 9 The structural schematic diagram of an embodiment of a cloud computing service device provided by the present application is shown.

[0023] Figure 10 The structural schematic diagram of another embodiment of a cloud computing service device provided by the present application is shown.

[0024] Figure 11 The structural schematic diagram of yet another embodiment of a cloud computing service device provided by the present application is shown.

[0025] Figure 12 The structural schematic diagram of an embodiment of a computing device provided by the present application is shown. Detailed Description of the Embodiments

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0027] In some processes described in the specification, claims, and the above-mentioned drawings of this application, a number of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The operation numbers such as 501 and 502 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0028] The technical solutions of the embodiments of this application can be applied to scenarios where complex cloud computing services are provided for users, especially scenarios where cloud computing services are provided based on tasks of different complexities. For example, based on a cloud computing service system, it is possible to disperse server resources to provide cloud computing services for multiple users, or it is possible to integrate all server resources to provide a cloud computing service with powerful computing power for a single user.

[0029] Currently, a set of cloud computing service systems provided by cloud service providers to users usually includes a computing cluster for processing complex tasks and a control cluster for managing and accessing the computing cluster. That is to say, a control cluster corresponds to a fixed computing cluster, and the two together form a set of cloud computing service systems. In order to meet the rental needs of more users, it is necessary to build multiple sets of such cloud computing service systems. Therefore, the following problems will exist: First, the systems are dispersed. Because this method requires deploying a set of independent cloud computing service systems for each processing task of each user. Therefore, it is necessary to build multiple sets of independent cloud computing service systems according to the number of users or processing tasks. At this time, the cloud service provider needs to disperse its limited number of high-performance servers into multiple small-scale computing clusters. This will result in overly dispersed computing resources when providing cloud computing services and it is difficult to form a scale.

[0030] Second, the construction cost is high. When there are new users or processing tasks, it is necessary to repurchase all the hardware devices in the computing cluster and the control cluster and then build the environment, increasing the initial investment cost and subsequent operation and maintenance costs.

[0031] Thirdly, it is difficult to expand. As the complexity of user tasks and the amount of data increase, the small-scale computing clusters after dispersion are difficult to meet the needs of users. At this time, expanding the computing clusters in the cloud computing service system has problems such as high construction costs and great difficulties.

[0032] To solve the above three problems, the inventor thought that the problem of scattered server resources could be solved by allowing multiple control clusters to share the same computing cluster. However, there are still problems such as possible interaction conflicts between the servers corresponding to multiple control clusters and the inability to guarantee communication security. To solve this problem, the inventor proposed the technical solution of this application through a series of innovative thinking. The cloud computing service system includes: a computing cluster composed of multiple servers, a first network switch system, and multiple control clusters; the first network provided by the first network switch system is divided into multiple first network domains; the first network switch system is used to connect multiple control clusters to different first network domains based on the first network domain configuration information, and connect at least one server corresponding to each of the multiple control clusters to their respective corresponding first network domains, so as to provide data forwarding services between the control clusters and servers connected to the same first network domain. The embodiment of this application can realize providing cloud computing services for multiple control clusters simultaneously by deploying only one computing cluster, greatly reducing the complexity and deployment cost of the deployment method. And the cloud service provider deploys its limited number of servers in one computing cluster. Through the method of network domain isolation, multiple control clusters can share the same computing cluster, that is, it can realize dispersing server resources to provide cloud computing services for multiple control clusters at the same time, and it can also realize that one control cluster exclusively occupies all the servers in the computing cluster to provide users with large-scale, high-complexity, and high-data-volume computing services, avoiding the problem that it is difficult to form a scale due to system dispersion and thus unable to meet the continuously increasing task processing requirements. In addition, when a new user or processing task appears, this solution only needs to redeploy a control cluster and then configure it into a new network domain of the existing cloud computing service system. Since the construction cost and environment setup difficulty of the control cluster are much lower than those of the computing cluster, the solution of the embodiment of this application greatly reduces the construction cost and difficulty compared with redeploying a set of cloud computing service systems including computing clusters and control clusters currently. In addition, when the task complexity or data volume of the user corresponding to any control cluster increases, only the number of servers connected to the network domain corresponding to the control cluster needs to be adjusted. This process does not require expanding the computing cluster, reducing the adjustment cost and difficulty.

[0033] In addition, in the case where any control cluster fails or is upgraded, resulting in the idleness of its corresponding server, through the configuration of the network domain, the idle server can also be configured for use by other control clusters, thus avoiding waste of resources. Moreover, during this process, there is no need to perform operations such as relocating the server and re-wiring. On the premise of low cost and high efficiency, the availability of server resources and the flexibility of resource allocation are further improved.

[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0035] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, it can be used in the solutions described herein within the scope permitted by applicable laws and regulations in accordance with the requirements of the applicable laws and regulations of the country where it is located (for example, with the explicit consent of the user, giving the user a practical notice, etc.).

[0036] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0037] It should be noted that the technical solutions in the embodiments of the present application are applicable to a network virtual environment. The users described generally refer to "virtual users". Real users can register user accounts on the server side through registration to obtain user identities in the network environment. The same user account can be logged in to the server side through different types of user terminals, enabling the server side to identify the same user.

[0038] The interaction operations between the server side and the user can be implemented based on the user account. The corresponding data received or sent by the server side to the user is also implemented based on the user account. Actually, it is the user terminal corresponding to the user account that receives or sends the corresponding data to the server side. In addition, through the user account, communication and the like can also be achieved between users. Among them, the user can refer to an individual or an organization, such as an enterprise, etc. The present application does not make specific limitations on this.

[0039] It should be noted that in the case where the embodiments of the present application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of the present application include, but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, eye movement operations, etc.; among them, touch operations include, but are not limited to: click operations, double-click operations, long-press operations, slide operations, pinch operations, or mouse hover operations, etc. Slide operations include, but are not limited to: linear slides, curved slides, etc.

[0040] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below.

[0041] As Figure 1 shown, it is a schematic structural diagram of an embodiment of a cloud computing service system provided by the present application. The system 10 may include a computing cluster 101, a first network switch system 102, and multiple control clusters 103. The computing cluster 101 includes multiple servers; the first network provided by the first network switch system 102 is divided into multiple first network domains.

[0042] Among them, the computing cluster 101 may be a server cluster used to provide high-complexity or high-data-volume operations. For example, it may be a server cluster dedicated to providing necessary hardware support and technical guarantee for artificial intelligence application programs. The servers included therein are all high-performance servers. For example, it may include, but is not limited to: Graphics Processing Unit (GPU) servers and high-performance memories, etc. GPU servers are mainly used to process high-complexity or high-data-volume tasks. For example, they execute large model training tasks. High-performance memories are mainly used to store high-data-volume data. For example, they store the training samples and their training labels required for large model training tasks. It should be noted that a cloud computing service system in this embodiment only includes one computing cluster, but this computing cluster can deploy all the limited high-performance servers that the cloud service provider can use in this cloud computing cluster to ensure the formation of a large-scale server cluster to meet the task processing requirements when the data processing speed and computing power continue to increase in the future.

[0043] The first network switch system 102 can be a switch device used to establish network connections between each control cluster 103 and its corresponding servers in the computing cluster 101. The switch devices included in the first network switch system 102 can further include at least one access switch and at least one out-of-band switch. The first network provided by the first network switch system 102 can be a network connected by switch ports in a wired manner. For example, it can be an optical fiber network. The multiple first network domains divided by the first network can be multiple virtual local area networks (VLANs), such as VLAN1 - VLANn. Optionally, the first network switch system 102 can also be deployed within the computing cluster 101, and this is not limited.

[0044] The first network switch system 102 is used to access multiple control clusters 103 into different first network domains based on the first network domain configuration information, and to access at least one server corresponding to each of the multiple control clusters 103 into their respective corresponding first network domains, and to provide data forwarding services between the control clusters 103 and the servers that are accessed into the same first network domain.

[0045] Among them, the first network domain configuration information can be pre-configured relevant information indicating how to divide multiple control clusters and multiple servers into different first network domains. For example, it can include the servers corresponding to each control cluster 103 and the first network domain identifier corresponding to each control cluster 103. It should be noted that in this embodiment, the first network domain identifiers corresponding to different control clusters 103 are different. The servers corresponding to each control cluster 103 can be the servers that the control cluster 103 needs to access and call to execute the task to be processed. In this embodiment, the servers corresponding to each control cluster 103 can be set in advance according to the hardware resources required for the task to be processed by the control cluster 103.

[0046] Optionally, the first network domain configuration information can be pre-configured and loaded into the first network switch system 102 by R & D personnel when building the cloud computing service system 10; or a scheduling system can be built inside or outside the cloud computing service system, and this scheduling system is used to send the first network domain configuration information to the first network switch system 102. For example, the scheduling system can be a visual unified scheduling platform. When users need to access the control cluster 103 and the servers into the first network domain, they can configure according to actual needs on the scheduling platform, and the scheduling platform generates the first network domain configuration information and then sends it to the first network switch system 102. Compared with directly loading it into the first network switch system 102, using the scheduling platform to send the first network domain configuration information facilitates subsequent adjustment of the devices added to different first network domains according to user needs.

[0047] Optionally, in this embodiment, the first network switch system 102 may, for each control cluster 103, look up the corresponding first network domain identifier in the first network domain configuration information, and configure the first network domain identifier for the port through which the control cluster 103 accesses the first network switch system 102. Then, it looks up one or more servers corresponding to the control cluster 103 in the first network domain configuration information, and configures the first network domain identifier for the ports through which the found one or more servers access the first network switch system 102. At this time, since the control cluster 103 and its corresponding servers are configured with the same first network domain identifier on the ports of the first network switch system 102, the control cluster 103 and its corresponding servers are connected to the same first network domain.

[0048] Exemplarily, as Figure 1 shown, the first network domain configuration information may be: control cluster 1 corresponds to the first network domain 1, and its corresponding servers are server 1 and server 2; control cluster 2 corresponds to the first network domain 2, and its corresponding servers are server 3 and server 4; control cluster 3 corresponds to the first network domain 3, and its corresponding servers are server 5 and server 6; at this time, based on the first network domain configuration information, the first network domain 1 may be configured for the ports through which control cluster 1, server 1, and server 2 access the first network switch system 102; the first network domain 2 may be configured for the ports through which control cluster 2, server 3, and server 4 access the first network switch system 102; the first network domain 3 may be configured for the ports through which control cluster 3, server 5, and server 6 access the first network switch system 102; thus, control cluster 1, server 1, and server 2 are connected to the first network domain 1; control cluster 2, server 3, and server 4 are connected to the first network domain 2; control cluster 3, server 5, and server 6 are connected to the first network domain 3.

[0049] It should be noted that since different first network domain identifiers are configured for different control clusters 103 in this embodiment, the mutual isolation of the first network domains where different control clusters are located is achieved, that is, the first network provided by the first network switch system 102 is finely planned and isolated. Through this network isolation strategy, when multiple control clusters 103 share the same computing cluster 101, the situation where multiple control clusters 103 simultaneously occupy the same server is avoided, ensuring the independence and security of server resource allocation.

[0050] In some embodiments, to ensure the reliability and security of network isolation in the first network, this embodiment may divide the first network based on a virtual local area network (VLAN). That is, at this time, the multiple first network domains obtained by dividing the first network in the above embodiments may be multiple virtual local area networks. Specifically, in this embodiment, the computing cluster and the multiple control clusters are respectively connected to the ports in the first network switch system; the first network provided by the first network switch system is divided into multiple virtual local area networks; the first network switch system, based on the first network domain configuration information, accesses the multiple control clusters into different first network domains, and accesses at least one server corresponding to each of the multiple control clusters into their respective corresponding first network domains, including: based on the first network domain configuration information, determining the virtual local area network identifiers corresponding to the multiple control clusters respectively; for any one of the control clusters, configuring the virtual local area network identifier corresponding to the control cluster for the first port to which the control cluster is connected and the second ports to which at least one server corresponding to the control cluster is connected. It should be noted that the specific implementation methods of dividing the virtual local area network and configuring the virtual local area network identifier in this embodiment are similar to the methods of dividing the first network domain and configuring the first network domain identifier introduced in the above embodiments, and only need to replace the first network domain in the above embodiments with the virtual local area network and the first network domain identifier with the virtual network domain identifier, and will not be elaborated here.

[0051] Optionally, in this embodiment, after the first network switch system completes the first network domain configuration for each server in the multiple control clusters and the computing cluster, the control cluster can communicate with its corresponding server. For example, it can send data to its corresponding server for the server to execute complex tasks based on the received data. At this time, the first network switch system needs to provide a data forwarding service between the control clusters and the servers accessing the same first network domain. In some embodiments, it may be to obtain the first data requested to be sent from any one control cluster to any one server, and forward the first data to the server when the server and the control cluster access the same first network domain. Among them, the first data in this embodiment may be the data that the control service cluster wants to send to any one server that it currently needs to access.

[0052] Specifically, when any control cluster has a data interaction requirement with any server, it will send a request to the first network switch system. The request may include the first data to be sent, the port identifier on the first network switch system 102 to which the server receiving the first data is connected, and the address of the server, such as the Internet Protocol (IP) address of the server. After receiving the request, the first network switch system can locate the port receiving the request (i.e., the first port), which is the port on the first network switch system to which the control cluster sending the request is connected. Then, it can obtain the first network domain identifier configured for this port. Then, based on the port identifier included in the request, it locates the port on the first network switch system to which the server is connected (i.e., the second port). Then, it compares whether the first network domain identifiers configured for the first port and the second port in the first network switch system are the same. If they are the same, it indicates that the control cluster sending the request and the server receiving the first data are accessing the same first network domain. At this time, it can further forward the first data to the server based on the address of the server included in the request. Optionally, in this embodiment, the request sent by the control cluster may not include the port identifier on the first network switch system 102 to which the server receiving the first data is connected. In this case, the first network switch system 102 can further locate the second port based on the pre-configured relationship between the address of the server and the connected port. Or it can also locate the second port through other means, which is not limited here.

[0053] When this embodiment provides a data exchange service for the control cluster and the server, it ensures that the data forwarding service is provided only when the first network domains configured for the access ports of the control cluster sending the data and the server receiving the data on the first network switch system are the same, ensuring the security of data transmission.

[0054] In some embodiments, if the computing cluster and multiple control clusters of this embodiment are respectively connected to ports in the first network switch system; the first network provided by the first network switch system is divided into multiple virtual local area networks. At this time, the implementation method for the first network switch system 102 to provide data forwarding services between the control cluster 103 and the servers accessing the same first network domain can be: obtain the first data that any control cluster requests to send to any server through the first port, determine the second port to which the server is connected, and when the virtual local area network identifiers configured on the first port and the second port are the same, forward the first data to the server through the second port. It should be noted that the specific implementation method of this embodiment is similar to the method introduced in the above embodiment. Only need to replace the first network domain in the above embodiment with a virtual local area network and replace the first network domain identifier with a virtual network domain identifier. Details are not described here. This embodiment provides the forwarding service of the first data only when the virtual local area network identifiers configured on the ports where the control cluster and the server access the first network switch are the same, thus ensuring the reliability and security of data transmission.

[0055] Optionally, in the R & D and test stage, due to problems such as frequent requirement changes or technology iterations, it is necessary to update or adjust the control cluster. Currently, common cloud computing service systems are constructed with a control cluster and a computing cluster in a one-to-one ratio. Therefore, when a control cluster fails or is upgraded, the computing cluster in the cloud computing service system where it is located will be idle due to lack of management and scheduling support. That is to say, although the servers in the computing cluster are still available, due to the lack of effective access points or management interfaces, they cannot actually be directly used, resulting in waste of resources. Especially for users who rely on continuous computing, such interruptions will not only reduce production efficiency but also may affect the user experience and cause serious losses. Currently, to solve this problem, a standby control cluster and its corresponding standby computing cluster are usually pre-deployed. Among them, the standby control cluster is the same as the regular control cluster in the cloud computing service system, but compared with the regular computing cluster deployed in the cloud computing service system, the standby computing cluster does not deploy servers, but only reserves the positions for deploying servers. When a certain control cluster fails or is upgraded, the servers in its corresponding computing cluster are relocated to the reserved positions in the standby cluster corresponding to the standby control cluster. Thus, the problem of waste of server resources caused by the unavailability of the control cluster is avoided. However, this solution still has the following problems: First, cluster redundancy construction: In order to ensure that the backup control cluster can normally access the standby control cluster after the server relocation, in addition to not deploying servers, other hardware facilities in the standby control cluster also need to be normally deployed. For example, a high-performance network required for interaction between servers needs to be deployed. This leads to redundant construction of other facilities in the control cluster and the standby control cluster except for the servers.

[0056] Second, the risk of frequent relocation. This solution involves moving the physical devices. Since the servers in the computing cluster are usually expensive and highly sensitive server devices, such as GPU servers. Any improper operation may cause damage to the devices, and frequent relocation will definitely pose high risks.

[0057] Third, high relocation costs. The relocation process of the servers usually requires technicians with professional knowledge to participate to ensure safe and error-free operations. In addition, the relocation process will still cause the server resources to be unavailable for a period of time, still affecting the continuity of task processing.

[0058] In some embodiments, to solve the above problems, the first network switch system of this embodiment is further configured to, in response to a switching request for a target server, determine a target control cluster, and switch the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0059] Among them, the target server can be a server that needs to be switched to another control cluster. For example, if Figure 1 Control Cluster 1 needs to be upgraded, at this time, Server 1 and / or Server 2 in the first network domain 1 where it is located can be used as the target server. The target control cluster can be the new control cluster that the target server needs to switch to at this time.

[0060] Optionally, an implementable way in this embodiment can be that if the user needs to expand the corresponding server due to the increase in the complexity of the task to be processed, the control cluster where the user is located can be used as the target control cluster, and one or more servers with an idle working state in the computing cluster can be selected as the target server according to the user's own needs, and then a switching request for the target server is triggered to be sent to the first network switch system through a scheduling system inside or outside the cloud computing service system. At this time, the first network switch system responds to the switching request for the target server, and determines the target server and the target control cluster from the switching request. Another implementable way can be that when the first network switch system detects that a certain server it is connected to is in an idle state, in order to prevent resource waste, it automatically uses it as the target server, triggers a switching request, and then determines the target control cluster from multiple control clusters according to a preset policy (such as using the control cluster with the highest current computing demand as the target control cluster).

[0061] After the first network switch system determines the target server and the target control cluster, it can switch the first network domain identifier configured on the port (i.e., the second port) of the first network switch system to which the target server is connected to the first network domain identifier corresponding to the target control cluster, so as to complete the switching of the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0062] This embodiment supports switching idle servers from the current control cluster to other target control clusters with usage requirements, which can well solve the problem of resource idleness of server resources due to the failure or upgrade of their control clusters. In addition, during the process of realizing server switching in this embodiment, there is no need to build multiple sets of standby control clusters and standby computing clusters, avoiding problems such as redundant construction of clusters, high risk of frequent relocation, and high relocation costs.

[0063] In some embodiments, as Figure 2 shown, in order to facilitate the management of each server in the computing cluster by the cloud service provider or user, this embodiment can further deploy a scheduling system 104 in the cloud computing service system 10 that has a connection relationship with the first network switch system 102; the scheduling system 104 is used to detect a switching operation for a target server, determine the target control cluster for which the switching is requested, and send a switching request to the first network switch system based on the target control cluster.

[0064] Among them, the scheduling system 104 can be oriented to the cloud service provider or user, for the cloud service provider or user to view the device information of each server in the computing cluster, and initiate the unified scheduling platform for switching the target server from the current control cluster to the target control cluster.

[0065] Optionally, a display interface for the cloud service provider or user to trigger the switching of the server control cluster can be provided on the scheduling system 104. The cloud service provider or user can select the target server to be switched and the target control cluster to which the server needs to be switched through this display interface, and then trigger the switching operation. At this time, the scheduling system 104 can obtain the target server and the target control cluster selected by the cloud service provider or user, and then send a switching request to the first network switch system 102 based on the target server and the target control cluster, requesting the first network switch system 102 to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster. As for how to specifically execute the switching process of the first network domain where the target server is located has been introduced in detail in the above embodiments and will not be elaborated here. This embodiment introduces a scheduling system, which facilitates the cloud service provider or user to initiate a switching request for the target server according to requirements, improving the convenience and flexibility of switching the target server control cluster.

[0066] In some embodiments, the scheduling system of this embodiment is further used to provide a display interface and display the device information of multiple servers and the corresponding switching prompt information for multiple servers on the display interface. Exemplarily, Figure 3 shows a schematic diagram of the display interface provided by the scheduling system, as Figure 3As shown, the display interface can display device information such as the device type, device identifier, hostname, current cluster, and user of each server in the computing cluster. Among them, the device type can indicate whether the server is a GPU or a memory. The device identifier can be a unique identifier characterizing the device. For example, it can be a device number. The hostname can be the name created by the control cluster to which each server belongs. It should be noted that in this embodiment, the hostnames given by different control clusters to the same server may not be the same. The current cluster can be the control cluster to which each server currently belongs. To ensure data security, a server in this embodiment usually belongs to only one control cluster. The user can be the user party currently using the server. The display interface also includes an operation page, where the operations at least include a migration operation of switching the server from one control cluster to another. For example, the cloud service provider or the user can locate the target server from each server by selecting the box in front of the device type, and then trigger the migration button on the interface. At this time, a prompt box for filling in the target control cluster will pop up on the interface. The cloud service provider or the user can fill in the new control cluster to which they want to migrate, that is, the target control cluster, in this prompt box, and then trigger the confirmation instruction. Thus, the cloud service provider or the user completes the trigger operation for switching the target server.

[0067] Correspondingly, the scheduling system detects the switching operation for the target server and determines the target control cluster requested to be switched, including: detecting the switching operation triggered by the switching prompt information corresponding to the target server, and determining the target control cluster requested to be switched. Specifically, it can be to use the server selected by the cloud service provider or the user in the Figure 3 interface shown as the target server, and use the control cluster filled in the prompt box after triggering the migration button as the target control cluster.

[0068] The scheduling system of this embodiment provides a display interface to display the device information of multiple servers, facilitating the cloud service provider or the user to understand the status and detailed information of the servers in real time, thereby assisting in the maintenance of the servers. In addition, the display interface can also provide a trigger for switching the control cluster to which the server belongs, improving the flexibility and convenience of server switching.

[0069] In some embodiments, since the server involves multiple complex steps when switching between different first network domains, such as address planning, in-band and out-of-band switch configuration, address segment negotiation, and network connection switching. If these operations rely on manual completion by humans, errors are very likely to occur. Therefore, in this embodiment, the scheduling system can control the target server to exit the original control cluster based on the target control cluster, send a switching request to the first network switch system to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster, and control the target server to join the target control cluster. Specifically, the scheduling system can track and record in real time information such as the control cluster to which each server currently belongs and the ports connecting to the first network switch system. After the scheduling system detects a switching operation for the target server, based on the connection port of the target server in the first network switch system (i.e., the first port), the control cluster to which the target server currently belongs, and the target control cluster to which it needs to be switched, it automatically calls and executes relevant scripts to automatically complete the operation of controlling the target server to exit the original control cluster, sending a switching request to the first network switch system to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster, and controlling the target server to join the target control cluster.

[0070] In this embodiment, the scheduling system replaces manual operation to complete the operation of switching the target server from the original control cluster to the target control cluster, avoiding errors caused by manual configuration and improving the reliability of the target server switching process.

[0071] In some embodiments, the application scenario of switching the target server from one control cluster to another is not limited to the scenarios of control cluster upgrade or failure introduced in the above embodiments. It can also be applicable to the scenario where any cloud service provider or user needs to expand its corresponding server due to increased task complexity or data volume, searches for a target server that meets its requirements in the computing cluster, and requests to switch it to the control cluster where it is located. For example, a certain user currently needs to perform a large model training task. Since this task has high requirements for server hardware and the servers it can currently access and schedule cannot meet this requirement, it is necessary to add other servers in the computing cluster to its own control cluster to expand the scale of the servers it can access, and then complete the large model training task. At this time, the user can also select the target server through the display interface provided by the scheduling system and then trigger a switching request to switch the target server to the control cluster where it is located.

[0072] For this scenario, it is possible that the target server that user 1 wants to switch to is currently being used by user 2. To avoid directly triggering the switching operation of the target server, which may cause user 2 to be unable to continue using the target server and thus interrupt their tasks, when the scheduling system of this embodiment executes a switching request to the first network switch system based on the target control cluster, it can specifically be used to: send an approval request to the client of the current user of the target server, and the approval request is used to request through the client whether the current user confirms the switching of the target server; correspondingly, if a confirmation instruction feedback by the client of the current user is received, the scheduling system sends a switching request to the first network switch system based on the target control cluster.

[0073] Specifically, the scheduling system can find the client identifier of the current user of the target server based on the display interface, and generate an approval request according to the relevant information of the target server (for example, the identifier of the target server), the switching time, etc. Based on the found client identifier, the generated approval request is sent to the client of the current user of the target server. This approval request can be used to request the current user to evaluate the impact of switching the target server on themselves through the client, so as to give a feedback instruction on whether to agree to the switch. For example, if the current user no longer needs to use the target server, then switching the target server will not affect themselves, and at this time the current user can feedback a confirmation instruction to the scheduling system through their client; if the current user is using the target server and the switch will affect the execution of their current tasks, at this time the current user can feedback a rejection instruction to the scheduling system through their client. If the scheduling system receives a confirmation instruction feedback by the client of the current user, it will send a switching request to the first network switch system based on the target control cluster. If a rejection instruction feedback by the client of the user is received, it will feedback the rejection message of the current user through the display interface and suggest communicating with the current user of the target server.

[0074] The solution of this embodiment solicits the opinions of the current user of the target server before switching the target server, avoids interrupting the tasks of the current user due to the switching of the target server, and thus ensures the stability and reliability of the entire cloud computing service system to provide services.

[0075] In some embodiments, when a server is re-switched to a new control cluster, the new control cluster usually needs to reinstall and configure it. This process not only takes time but also has a relatively high risk. Once a mistake occurs, it may lead to data loss and even threaten the stability and security of the overall system. To solve this problem, as Figure 2As shown, the scheduling system of this embodiment can also have a connection relationship with multiple control clusters; the scheduling system is also used to obtain the host name of the target server from the target control cluster, and establish a correspondence between the host name and the target server, and based on the correspondence, generate initialization data, and send the initialization data to the target control cluster; the target control cluster is also used to perform initialization operations on the target server based on the initialization data.

[0076] Specifically, when the target server switches to the target control cluster, the target control cluster will uniformly manage the newly switched target server, including but not limited to: reinstalling, installing drivers, setting hostnames, and a series of other operations, so as to prepare resources for the target control cluster to schedule the target server later. Among them, the initialization operations such as reinstalling, installing drivers, and verifying whether the software configuration of the target server is consistent with the hardware need to rely on the topological mapping relationship between the hostname and the network in the computing cluster (i.e., the second network provided by the subsequent second network switch) (such as which second network domain the target server belongs to after the second network is divided) to be implemented. Since the scheduling system of this embodiment has established a communication connection with multiple control clusters, the scheduling system can exchange the host name of the target server with the target control cluster after the target control cluster sets a new host name for the target server, and then establish a corresponding relationship between the host name and the unique identifier of the target server. The corresponding relationship is used to locate the corresponding target server based on the host name when performing subsequent initialization operations. The scheduling system will generate initialization data based on the generated correspondence. The initialization data can be an automatically generated and updated configuration file, and then send the initialization data to the target control cluster. The target control cluster will automatically complete the relevant operations of initializing the target server based on the initialization data, for example, including but not limited to: reinstalling, installing drivers, and verifying whether the software configuration of the target server is consistent with the hardware.

[0077] This embodiment opens the communication connection between the dispatching system and the control cluster, obtains the host name set by the control cluster for the target server, and then generates initialization data for the target control cluster to automatically complete the initialization operation on the target server. Compared with manual operation, this solution automatically completes the initialization operation, reducing the time and risk of the initialization operation.

[0078] In some embodiments, Figure 4As shown, the computing cluster in this embodiment further includes a second network switch system; multiple servers are connected through a second network provided by the second network switch system; the second network switch system is configured to, based on second network domain configuration information, connect at least one server corresponding to the same control cluster or the same user to at least one second network domain corresponding to the server, and provide data forwarding services between different servers connected to the same second network domain; the second network domain is obtained by partitioning the second network provided by the second network switch system.

[0079] Optionally, since the servers in the computing cluster are all high-performance servers and the tasks they process are all data with relatively high complexity or large data volume, in order to ensure the normal operation of the servers, the second network switch system is usually a high-performance network switch. Correspondingly, the second network provided by the high-performance network switch is also usually a high-performance network (i.e., a high-performance network). The second network domain configuration information can be pre-configured indication information for indicating how to partition multiple servers into different second network domains. The second network switch system in this embodiment can partition the second network based on the second network domain configuration information, combined with the number of control clusters or the number of users corresponding to each control cluster, to obtain multiple second network domains, and connect the servers corresponding to each control cluster or the servers corresponding to each user of each control cluster to the same second network domain, so as to isolate one or more servers corresponding to different control clusters or different users from each other in the second network. For example, Figure 4 As shown, if the second network is partitioned according to the number of control clusters, at this time, based on the second network domain configuration information, server 1 and server 2 can be connected to the second network domain 1 corresponding to control cluster 1, server 3 and server 4 can be connected to the second network domain 2 corresponding to control cluster 2, and server 5 and server 6 can be connected to the second network domain 3 corresponding to control cluster 3.

[0080] Optionally, after the second network switch system in this embodiment connects at least one server belonging to the same control cluster or corresponding to the same user to the same second network domain, the servers in the same second network domain can communicate with each other. For example, a memory-type server can transmit data to a GPU processor-type server. Specifically, when any server needs to transmit data to other servers, it can generate a data transmission request based on the data to be transmitted and the destination address for receiving the data (such as the address of the server that will receive the data to be transmitted), and send the request to the second network switch system. The second network switch system responds to the data transmission request, determines the source address of the server that sent the request and the destination address for receiving the data, and then determines whether the source address and the destination address comply with the transmission rules based on the Access Control List (ACL), that is, whether they belong to the same second network domain. If so, it means that the transmission rules are complied with, and at this time, the data to be transmitted is forwarded to the server corresponding to the destination address. Among them, the access control list can be a predefined rule for data forwarding or discarding based on the second network, and it is usually defined based on information such as the source address, destination address, port number, and protocol type that allow data forwarding.

[0081] The solution of this embodiment can avoid configuration conflicts in the second network switch system. In addition, in order to ensure the network independence between different control clusters, this solution not only needs to partition the network domain provided by the first network switch system to achieve network isolation, but also needs to partition the network domain provided by the second network switch system to achieve network partition isolation. This dual isolation mechanism significantly increases the complexity and security of network management in the cloud computing service system.

[0082] In some embodiments, in order to further improve the security of the second network isolation in this embodiment, EVPC (Express Virtual Private Cloud) can be used as an isolation means. Specifically, the second network switch system is specifically used for: based on the second network domain configuration information, partitioning at least one server belonging to the same control cluster or corresponding to the same user into at least one network segment corresponding to the server, and providing data forwarding services between different servers partitioned into the same network segment; it should be noted that the second network domain in the above embodiment can be a network segment obtained by partitioning the address range corresponding to the second network at this time; different control clusters or different users correspond to different network segments.

[0083] Specifically, in this embodiment, the addresses corresponding to the second network can be divided into multiple network segments based on the number of control clusters in the cloud computing service system or the number of users. Alternatively, several additional network segments can be reserved on the basis of meeting the number of control clusters or the number of users, so that when new control clusters are added to the cloud computing server system subsequently, corresponding network segments can be allocated for them based on the reserved network segments. Then, based on the second configuration, one or more servers corresponding to each control cluster are divided into the network segment corresponding to the control cluster, so as to isolate the second network between the servers of the same control cluster.

[0084] Correspondingly, when providing data forwarding services between different servers divided into the same network segment at this time, it may be to obtain the second data requested by the first server to be sent to the second server, and based on the access control list, determine that the destination address of the second server and the source address of the first server belong to the same network segment, and forward the second data to the second server.

[0085] Among them, the first server in this embodiment is the server that needs to send data, the second server is the recipient of the data, and the second data is the data to be transmitted this time. The access control list at this time can be a rule that defines the judgment of the source address and the destination address belonging to the same network segment. At this time, when the first server in this embodiment needs to transmit the second data to the second server, a data transmission request can be generated based on the second data and the target address of the second server and sent to the second network switch system. The second network switch system responds to the data transmission request, determines the source address of the first server that sends the request and the destination address of the second server that receives the data, and then determines whether the source address and the destination address conform to the transmission rule, that is, whether they belong to the same network segment, based on the access control list (Access Control List, ACL). If so, it means that the transmission rule is met, and at this time, the data to be transmitted is forwarded to the server corresponding to the destination address.

[0086] Optionally, based on the above embodiments, since the servers in the computing cluster interact based on the second network domain where they are located, in order to ensure that the target server can still interact with other servers in the target control cluster normally after switching to the new control cluster, the control machine system in this embodiment is not only used to control the first network switch system to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster. The scheduling system is also used to control the second network switch system to determine the target control cluster in response to the switching request for the target server and switch the target server from the original control cluster to the second network domain corresponding to the target control cluster. Among them, the process of the second network switch system determining the target control cluster in response to the switching request for the target server is similar to the process of the first network switch system determining the target control cluster in response to the switching request for the target server, which will not be elaborated here. When switching the target server from the original control cluster to the second network domain corresponding to the target control cluster, the second network domain corresponding to the target control cluster or the user of the target server can be determined, and the rules in the access control list can be adjusted to use the address of the target server as the address allowed for interaction in the second network domain, so as to realize switching the target server from the original control cluster to the second network domain corresponding to the target control cluster. It should be noted that the second network domain can be a network segment obtained by dividing the second network.

[0087] It should be noted that in the embodiments of the present application Figures 1 to 4 The number of servers included in the computing cluster 101, the number of control clusters 103, the number of first network domains into which the first network is divided, and the number of second network domains into which the second network is divided shown are only examples. For example, the number of servers included in the computing cluster 101 and the specific types of the servers can be set according to actual needs. The number of control clusters 103 can be set according to the number of users or the needs of users renting the cloud computing service system; the number of first network domains and second network domains can be determined according to the number of control clusters, etc. None of these are limited.

[0088] Figure 5 is a flowchart of an embodiment of a cloud computing service providing method provided by the present application. The technical solution of this embodiment can be executed by the first network switch system of the cloud computing service system. Among them, the cloud computing service system further includes a computing cluster and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the structure and working principle of the cloud computing service system are described in Figure 1 the relevant embodiments shown and will not be elaborated here.

[0089] Figure 5 As shown, the cloud computing service providing method may include the following steps: 501: Based on the first network domain configuration information, connect multiple control clusters to different first network domains, and connect at least one server corresponding to each of the multiple control clusters to the respective first network domains.

[0090] 502: Provide data forwarding services between the control clusters and servers connected to the same first network domain.

[0091] In some embodiments, the method further includes: in response to a handover request for a target server, determine a target control cluster, and switch the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0092] Optionally, the computing cluster and the multiple control clusters are respectively connected to ports in the first network switch system; the first network provided by the first network switch system is divided into multiple virtual local area networks; that is, the multiple first network domains in the above embodiments are multiple virtual local area networks. Correspondingly, based on the first network domain configuration information, connecting multiple control clusters to different first network domains, and connecting at least one server corresponding to each of the multiple control clusters to the respective first network domains includes: based on the first network domain configuration information, determine the virtual local area network identifiers corresponding to each of the multiple control clusters; for any one control cluster, configure the virtual local area network identifier corresponding to the control cluster for the first port to which the control cluster is connected and the second port to which at least one server connected thereto is connected.

[0093] In some embodiments, providing data forwarding services between the control clusters and servers connected to the same first network domain includes: obtaining first data requested to be sent from any one control cluster to any one server, and when the server and the control cluster are connected to the same first network domain, forwarding the first data to the server.

[0094] In some embodiments, if the first network domain is a virtual local area network, obtaining first data requested to be sent from any one control cluster to any one server, and when the server and the control cluster are connected to the same first network domain, forwarding the first data to the server includes: obtaining, through a first port, first data requested to be sent from any one control cluster to any one server, determining the second port to which the server is connected, and when the virtual local area network identifiers configured for the first port and the second port are the same, forwarding the first data to the server through the second port.

[0095] Figure 5 The specific implementation manners and technical effects of the cloud computing service providing method shown above have been correspondingly described in the relevant embodiments of the above cloud computing service system, and will not be elaborated herein.

[0096] Figure 6 The flowchart of an embodiment of a method for providing a cloud computing service provided by this application. The technical solution of this embodiment can be executed by the second network switch system of the cloud computing service system. Among them, the cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to the corresponding first network domain; and data forwarding is supported between the control clusters and servers connected to the same first network domain. The structure and working principle of the cloud computing service system have been described in the relevant embodiments shown in Figure 3 and will not be elaborated here.

[0097] Figure 6 As shown, the method for providing the cloud computing service may include the following steps: 601: Based on the second network domain configuration information, at least one server corresponding to the same control cluster or the same user is connected to the second network domain corresponding to the at least one server; the second network domain is obtained by dividing the second network provided by the second network switch system.

[0098] 602: Provide data forwarding services between different servers connected to the same second network domain.

[0099] In some embodiments, the specific implementation manners of 601 and 602 may be: based on the second network domain configuration information, at least one server corresponding to the same control cluster or the same user is divided into the network segment corresponding to the at least one server, and data forwarding services are provided between different servers divided into the same network segment; the network segment is obtained by dividing the address range corresponding to the second network; different control clusters or different users correspond to different network segments.

[0100] In some embodiments, different second network domains correspond to different network segments. Correspondingly, providing data forwarding services between different servers connected to the same second network domain includes: obtaining the second data sent by the first server to the second server, and based on the access control list, determining that the destination address of the second server and the source address of the first server belong to the same network segment, and forwarding the second data to the second server.

[0101] In some embodiments, this method is also used to respond to a switching request for a target server, determine the target control cluster, and switch the target server from the original control cluster to the second network domain corresponding to the target control cluster.

[0102] Figure 6 The specific implementation manners and technical effects of the cloud computing service providing method shown above have been described in the relevant embodiments of the above cloud computing service system, and will not be elaborated here.

[0103] Figure 7 FIG. is a flowchart of an embodiment of a cloud computing service providing method provided by the present application. The technical solution of this embodiment can be executed by the scheduling system of the cloud computing service system. Among them, the cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to the first network domain corresponding to it; data forwarding is supported between the control clusters and servers connected to the same first network domain; the structure and working principle of the cloud computing service system are in Figure 2 the relevant embodiments shown above have been described accordingly, and will not be elaborated here.

[0104] Figure 7 As shown, the cloud computing service providing method may include the following steps: 701: Detect a switching operation for a target server, and determine a target control cluster that requests the switch.

[0105] 702: Based on the target control cluster, send a switching request to the first network switch system, so that the first network switch system responds to the switching request, determines the target control cluster, and switches the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0106] In some embodiments, sending a switching request to the first network switch system based on the target control cluster includes: controlling the target server to exit the original control cluster based on the target control cluster, sending a switching request to the first network switch system to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster, and controlling the target server to join the target control cluster.

[0107] In some embodiments, sending a switching request to the first network switch system based on the target control cluster includes: sending an approval request to the client of the current user of the target server, where the approval request is used to request the current user through the client whether to confirm the switch of the target server; if a confirmation instruction feedback by the client of the user is received, then based on the target control cluster, send a switching request to the first network switch system.

[0108] In some embodiments, it further includes obtaining the host name of the target server from the target control cluster, establishing the correspondence between the host name and the target server, generating initialization data based on the correspondence, and sending the initialization data to the target control cluster for the target control cluster to perform initialization operations on the target server based on the initialization data.

[0109] In some embodiments, it further includes providing a display interface and displaying device information of multiple servers and switching prompt information corresponding to the multiple servers respectively on the display interface; Correspondingly, it further includes detecting a switching operation for the target server and determining the target control cluster requesting the switch, including: detecting a switching operation triggered by the switching prompt information corresponding to the target server and determining the target control cluster requesting the switch.

[0110] In some embodiments, it further includes sending the first network domain configuration information to the first network switch system.

[0111] Figure 7 The specific implementation manners and technical effects of the cloud computing service providing method shown above have been correspondingly described in the relevant embodiments of the above cloud computing service system, and will not be elaborated here.

[0112] In some specific implementation scenarios, the computing cluster of the present application may be an artificial intelligence (AI) service cluster, including multiple storage servers and multiple computing servers configured with graphics processing units (GPUs), that is, GPU servers; the first network switch system is specifically used to connect at least one computing server and / or at least one storage server corresponding to each of the multiple control clusters to their respective first network domains, and provide data forwarding services between the control clusters and the computing servers, and between the control clusters and the storage servers connected to the same first network domain. Optionally, the storage server may be a high-performance memory.

[0113] For ease of understanding, Figure 8 shows the structural diagram of a cloud computing service system in an actual scenario. As Figure 8As shown, the cloud product clusters A801 - C803 in the figure serve as multiple control clusters, the centralized AI cluster 804 serves as a computing cluster, the access switch serves as the first network switch system, the high - performance network switch serves as the second network switch system; the unified scheduling platform 805 serves as the scheduling system. Among them, the access switch and the high - performance network switch can be deployed together in the centralized AI cluster 804, and the multiple services included in the centralized AI cluster 804 include multiple high - performance memories and GPU servers. Optionally, the high - performance network in this embodiment can be a high - speed network constructed through RoCE v2 (Remote Direct Memory Access based on Converged Ethernet) high - speed interconnection technology.

[0114] In this scenario, the cloud product clusters A801 - C803, GPU servers 1 - n, and high - performance memories 1 - n are respectively connected to the ports of the access switch; the network provided by the access switch is divided into multiple VLANs, namely VLAN1 - VLAN3. Among them, the cloud product cluster A801 corresponds to VLAN1, the cloud product cluster B802 corresponds to VLAN2, and the cloud product cluster C803 corresponds to VLAN3. Therefore, configure VLAN1 for the connection interfaces of the cloud product cluster A801 and its corresponding GUP server 1 and high - performance memory 1 on the access switch, configure VLAN2 for the connection interfaces of the cloud product cluster B802 and its corresponding GUP server 2 and high - performance memory 2 on the access switch, and configure VLAN3 for the connection interfaces of the cloud product cluster C803 and its corresponding GUP server 3 and high - performance memory 3 on the access switch. Thus, each cloud product cluster and its corresponding GPU server and high - performance memory are isolated from the physical network based on VLAN1 - VLAN3. To ensure that in the centralized cluster 804, the GUP servers and high - performance memories of different cloud product clusters or different users are isolated from each other in the high - performance network, the high - performance network can be divided into multiple network segments, namely high - performance network segment 1 - high - performance network segment 2. Divide the GUP server 1 and high - performance memory 1 that belong to the cloud product cluster A801 into high - performance network segment 1; divide the GUP server 2 and high - performance memory 2 that belong to the cloud product cluster B802 into high - performance network segment 2; divide the GUP server 3 and high - performance memory 3 that belong to the cloud product cluster C803 into high - performance network segment 3. The above - mentioned dual isolation mechanism significantly increases the complexity of network management.

[0115] Based on Figure 8The cloud computing service system shown can achieve data transmission between the cloud product clusters 801-803 and the centralized AI cluster 804. Taking the large model training task of the cloud product cluster A801 as an example, the cloud product cluster A801 needs to send the large model data to the GPU server 1 and send the training sample data required for this training to the high-performance memory 1. When the GPU server 1 performs training based on the received large model data, it also needs to obtain the training samples from the high-performance memory 1. The above data transmission process can be as follows: The access switch obtains the large model data sent by the cloud product cluster A801 to the GPU server 1 and the training samples sent to the high-performance memory 1 through the access port of the cloud product cluster A801 on the access switch. Then, it is determined that the access port of the GPU server 1 configured on the access switch is VLAN1, the access port of the high-performance memory 1 configured on the access switch is also VLAN1, and the access port of the cloud product cluster A801 configured on the access switch is still VLAN1. At this time, it can be determined that the VLAN identifiers of the cloud product cluster A801, the GPU server 1, and the high-performance memory 1 are the same. The large model data can be forwarded to the GPU server 1 based on the IP address of the GPU server 1 through the corresponding port, and the training samples can be forwarded to the high-performance memory 1 based on the IP address of the high-performance memory 1. When the GPU server 1 executes the training task, it needs the high-performance memory 1 to transmit the training samples to the GPU server 1. At this time, the high-performance memory 1 will send the training samples and the address of the GPU server 1 that receives the training samples, that is, the destination address, to the high-performance network switch. The high-performance network switch will determine whether the destination address corresponding to the GPU server 1 belongs to the same network segment as the source address of the high-performance memory 1 that sends the training data based on the access control list. If so, the training samples will be forwarded to the GPU server 1 based on the destination address.

[0116] In this cloud computing service system, if product cluster A801 needs to be upgraded, the corresponding GPU server 1 and high-performance memory 1 will be in an idle state at this time. If a user A of cloud product cluster B802 views the current state of GPU server 1 as idle through the display interface of the unified scheduling platform 805 and wants to use GPU server 1, user A can trigger a switching operation based on the switching prompt information of GPU server 1 on the display interface of the unified scheduling platform 805 and fill in the target cloud product cluster to which GPU server 1 is to be switched, that is, cloud product cluster B802. At this time, the unified scheduling platform 805 detects the switching operation corresponding to GPU server 1, determines that the target server to be switched is GPU server 1, and then determines whether there is a user for GPU server 1 currently. If there is, it sends an approval request to the client of the current user of GPU server 1 to request whether the current user confirms the switching of GPU server 1 through this client; if the current user agrees to the switching, it will send a confirmation instruction to the unified scheduling platform 805 through its client. After receiving the confirmation instruction, the unified scheduling platform 805 will determine that GPU server 1 needs to be switched to cloud product cluster B802, and then the access switch will control GPU server 1 to exit cloud product cluster A801 through an automatically executed script file, send a switching request to the access switch to switch GPU server 1 from VLAN1 corresponding to cloud product cluster A801 to VLAN2 corresponding to cloud product cluster B802, and control GPU server 1 to join cloud product cluster B802. At the same time, the unified scheduling platform 805 is also used to control the high-performance network switch to switch GPU server 1 from high-performance network segment 1 corresponding to cloud product cluster A801 to high-performance network segment 2 corresponding to cloud product cluster B802. At this time, only the network access switching of GPU server 1 is completed. Cloud product cluster B802 will set a host name for the newly added GPU server 1. The unified scheduling platform 805 will obtain the host name of GPU server 1 from cloud product cluster B802 and establish a correspondence between the host name and GPU server 1. Based on this correspondence, initialization data will be generated and the initialization data will be sent to cloud product cluster B802; cloud product cluster B802 will automatically perform an initialization operation based on this initialization data. Only after the initialization operation is performed on GPU server 1 is the operation of switching to cloud product cluster B802 truly completed.

[0117] In this example, for the situation where any cloud product failure or upgrade causes the corresponding GPU server or high-performance memory to be idle, it is possible to configure the idle GPU server or high-performance memory for use by other cloud product clusters by performing operations such as VLAN switching and high-performance network segment switching on the target GPU server or high-performance memory, as well as subsequent initialization, etc., thus avoiding waste of resources. In addition, for large-scale tasks such as large model training, this embodiment can use a similar method as above to switch all GPU servers and high-performance memories to a single cloud product cluster, so as to integrate all server resources to provide large-scale task services for a cloud product cluster. During the switching process of the servers of this embodiment between different cloud product clusters, there is no need to perform operations such as relocating the servers and re-wiring. On the premise of low cost and high efficiency, it further improves the availability of resources and the flexibility of resource allocation.

[0118] The detailed implementation manners and beneficial effects of each step in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.

[0119] As Figure 9 shown, it is a schematic structural diagram of an embodiment of a cloud computing service device provided by the present application, which can be applied to the first network switch system of a cloud computing service system. The cloud computing service system further includes a computing cluster and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains. Among them, the structure and working principle of the cloud computing service system have been correspondingly described in the Figure 1 related embodiments shown, and will not be elaborated herein.

[0120] The device may include the following several modules: The first network isolation module 901 is used to connect the multiple control clusters to different first network domains based on the first network domain configuration information, and connect at least one server corresponding to each of the multiple control clusters to their respective corresponding first network domains; The first data forwarding module 902 is used to provide data forwarding services between the control clusters and servers connected to the same first network domain.

[0121] Figure 9 The cloud computing service device shown can be used to execute Figure 5 the cloud computing service providing method shown, and its implementation principle and technical effects will not be elaborated. Among them, Figure 9 one or more modules of the cloud computing service device shown can constitute Figure 1The first network switch system in the system architecture shown. For the cloud computing service device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0122] As Figure 10 shown, a structural schematic diagram of an embodiment of a cloud computing service device provided by the present application can be applied to the second network switch system of a cloud computing service system. The cloud computing service system further includes a computing cluster, a first network switch system, and a plurality of control clusters; the computing cluster includes a plurality of servers; the first network provided by the first network switch system is divided into a plurality of first network domains; the plurality of control clusters are connected to different first network domains, and at least one server corresponding to each of the plurality of control clusters is connected to the corresponding first network domain; and data forwarding is supported between the control clusters and the servers connected to the same first network domain. Among them, the structure and working principle of the cloud computing service system have been described in the relevant embodiments shown Figure 3 and will not be elaborated herein.

[0123] The device may include the following modules: A second network isolation module 1001, configured to connect at least one server corresponding to the same control cluster or the same user to the second network domain corresponding to the at least one server based on second network domain configuration information; the second network domain is obtained by dividing the second network provided by the second network switch system; A second data forwarding module 1002, configured to provide a data forwarding service between different servers connected to the same second network domain.

[0124] Figure 10 The cloud computing service device shown can be used to execute Figure 6 the cloud computing service providing method shown, and its implementation principle and technical effects will not be elaborated. Among them, Figure 10 one or more modules of the cloud computing service device shown can constitute Figure 3 the second network switch system in the system architecture shown. For the cloud computing service device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0125] As Figure 11As shown in the figure, it is a schematic structural diagram of an embodiment of a cloud computing service device provided by the present application, which can be applied to the scheduling system of a cloud computing service system. The cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to the first network domain corresponding to it; data forwarding is supported between the control clusters and servers connected to the same first network domain. Among them, the structure and working principle of the cloud computing service system are described in the relevant embodiments shown in Figure 2 and will not be elaborated here.

[0126] The device may include the following modules: The cluster determination module 1101 is used to detect a switching operation for a target server and determine the target control cluster requesting the switch. The sending module 1102 is used to send a switching request to the first network switch system based on the target control cluster, so that the first network switch system can determine the target control cluster in response to the switching request and switch the target server from the original control cluster to the first network domain corresponding to the target control cluster.

[0127] Figure 11 The cloud computing service device shown in the figure can be used to execute Figure 7 the cloud computing service providing method shown in the figure, and its implementation principle and technical effects will not be elaborated. Among them, Figure 11 one or more modules of the cloud computing service device shown in the figure can constitute Figure 2 the scheduling system in the system architecture shown in the figure. For the cloud computing service device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0128] Figure 12 It is a schematic structural diagram of an embodiment of a computing device provided by the present application. As Figure 12 shown, in practice, the computing device may include: a storage component 1201 and a processing component 1202.

[0129] The storage component 1201 is used to store computer programs and can be configured to store various other data to support operations on the computing device. Examples of these data include instructions for any application program or method operating on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0130] The processing component 1202, coupled to the storage component 1201, is configured to execute the computer program in the storage component 1201 to implement the cloud computing service providing method as shown in Figures 5 - 7 the following.

[0131] Furthermore, as shown in Figure 12 the following, the computing device may further include other components such as a communication component 1203, a display component 1204, a power supply component 1205, an audio component 1206, etc. Figure 12 Only some components are schematically shown herein, and it does not mean that the computing device only includes the Figure 12 components shown. Additionally, Figure 12 the components within the dashed box in the following are optional components, not mandatory components, and may vary depending on the product form of the computing device. The computing device in this embodiment may be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT (Internet of Things) device, or may also be a server device such as a conventional server, a cloud server, or a server array. If the computing device in this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include the Figure 12 components within the dashed box in the following; if the computing device in this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include the Figure 12 components within the dashed box in the following.

[0132] The above-mentioned processing component includes one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for executing the above method.

[0133] The above storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0134] The above communication component is configured to facilitate communication, either wired or wireless, between the device where the communication component is located and other devices. The device where the communication component is located can access a communication standard-based wireless network, such as a mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0135] The above display component may include a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation.

[0136] The above power supply component provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0137] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0138] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method embodiments. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions that, when executed by a processor, enable the processor to implement the steps in the above method embodiments. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiments.

[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0140] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0141] Finally, it should be noted that the above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A cloud computing service system, characterized in that, Including: A computing cluster, a first network switch system, and multiple control clusters; The computing cluster includes multiple servers; The first network provided by the first network switch system is divided into multiple first network domains; The first network switch system is used to access the multiple control clusters into different first network domains based on the first network domain configuration information, and access at least one server corresponding to each of the multiple control clusters into the first network domain corresponding to each of them, and provide data forwarding services between the control clusters and servers accessed into the same first network domain.

2. The system according to claim 1, wherein The first network switch system is further used to, in response to a switching request for a target server, determine a target control cluster, and switch the target server from the original control cluster to the first network domain corresponding to the target control cluster.

3. The system according to claim 2, wherein Further including: A scheduling system having a connection relationship with the first network switch system; The scheduling system is used to detect a switching operation for a target server, determine the target control cluster requesting the switch, and send a switching request to the first network switch system based on the target control cluster.

4. The system according to claim 3, wherein When the scheduling system sends a switching request to the first network switch system based on the target control cluster, it specifically uses: Based on the target control cluster, control the target server to exit the original control cluster, send a switching request to the first network switch system to switch the target server from the original control cluster to the first network domain corresponding to the target control cluster, and control the target server to join the target control cluster.

5. The system according to claim 3, characterized in that, When the scheduling system sends a switching request to the first network switch system based on the target control cluster, it specifically uses: Send an approval request to the client of the current user of the target server, and the approval request is used to request through the client whether the current user confirms to switch the target server; If receiving the confirmation indication fed back by the client of the current user, then send a switching request to the first network switch system based on the target control cluster.

6. The system according to claim 3, wherein The scheduling system has a connection relationship with multiple control clusters; The scheduling system is further used to obtain the host name of the target server from the target control cluster, establish the corresponding relationship between the host name and the target server, generate initialization data based on the corresponding relationship, and send the initialization data to the target control cluster; The target control cluster is further used to perform an initialization operation on the target server based on the initialization data.

7. The system according to claim 3, wherein The scheduling system is further used to provide a display interface, and display the device information of multiple servers and the switching prompt information corresponding to each of the multiple servers on the display interface; When the scheduling system detects a switching operation for a target server and determines the target control cluster requesting the switch, it specifically uses: Detect a switching operation triggered by the switching prompt information corresponding to the target server, and determine the target control cluster requesting the switch.

8. The system according to claim 1, wherein The computing cluster and the multiple control clusters are respectively connected to ports in the first network switch system; the first network domain is a virtual local area network; When the first network switch system accesses the multiple control clusters into different first network domains and accesses at least one server corresponding to each of the multiple control clusters into the respective corresponding first network domains based on the first network domain configuration information, it specifically is used for: Based on the first network domain configuration information, determine the virtual local area network identifiers corresponding to the multiple control clusters respectively; for any one control cluster, configure the virtual local area network identifier corresponding to the control cluster for the first port to which the control cluster is connected and the second ports to which at least one server connected thereto is connected.

9. The system according to claim 1, wherein When the first network switch system provides data forwarding services between the control clusters and servers accessing the same first network domain, it specifically is used for: Obtain the first data that any one control cluster requests to send to any one server, and when the server and the control cluster access the same first network domain, forward the first data to the server.

10. The system according to claim 9, wherein The first network domain is a virtual local area network; When the first network switch system obtains the first data that any one control cluster requests to send to any one server and, when the server and the control cluster access the same first network domain, forwards the first data to the server, it specifically is used for: Obtain the first data that any one control cluster requests to send to any one server through the first port, determine the second port to which the server is connected, and when the virtual local area network identifiers configured for the first port and the second port are the same, forward the first data to the server through the second port.

11. The system according to claim 1, wherein The computing cluster further includes a second network switch system; the multiple servers establish connections through the second network provided by the second network switch system; The second network switch system is used for, based on the second network domain configuration information, accessing at least one server corresponding to the same control cluster or the same user into the second network domain corresponding to the at least one server, and providing data forwarding services between different servers accessing the same second network domain; The second network domain is obtained by partitioning the second network provided by the second network switch system.

12. The system according to claim 11, wherein The second network domain is a network segment obtained by partitioning the address range corresponding to the second network; different control clusters or different users correspond to different network segments; The second network switch system is specifically used for: based on the second network domain configuration information, partitioning at least one server corresponding to the same control cluster or the same user into the network segment corresponding to the at least one server, and providing data forwarding services between different servers partitioned into the same network segment.

13. The system according to claim 12, wherein Different second network domains correspond to different network segments. When the second network switch system provides data forwarding services between different servers accessing the same second network domain, it specifically is used for: Obtain the second data sent from the first server to the second server, and based on the access control list, when it is determined that the destination address of the second server and the source address of the first server belong to the same network segment, forward the second data to the second server.

14. The system according to claim 11, wherein: The scheduling system is further configured to control the second network switch system. In response to a handover request for a target server, determine a target control cluster, and switch the target server from the original control cluster to the second network domain corresponding to the target control cluster.

15. The system according to any one of claims 1-14, characterized in that, The computing cluster is an artificial intelligence service cluster, including multiple storage servers and multiple computing servers configured with graphics processors; The first network switch system is specifically configured to connect at least one computing server and / or at least one storage server corresponding to each of the multiple control clusters to their respective corresponding first network domains, and provide data forwarding services between the control clusters and computing servers, and between the control clusters and storage servers that are connected to the same first network domain.

16. A method for providing cloud computing services, characterized in that, Applied to the first network switch system in a cloud computing service system, the cloud computing service system further includes a computing cluster and multiple control clusters; the computing cluster includes multiple servers; The first network provided by the first network switch system is divided into multiple first network domains; The method includes: Based on the first network configuration information, connect the multiple control clusters to different first network domains, and connect at least one server corresponding to each of the multiple control clusters to their respective corresponding first network domains; Provide data forwarding services between the control clusters and servers that are connected to the same first network domain.

17. A method for providing cloud computing services, characterized in that Applied to the second network switch system in a cloud computing service system, the cloud computing service system further includes a computing cluster, a first network switch system, and multiple control clusters; the computing cluster includes multiple servers; the first network provided by the first network switch system is divided into multiple first network domains; the multiple control clusters are connected to different first network domains, and at least one server corresponding to each of the multiple control clusters is connected to their respective corresponding first network domains; And data forwarding is supported between the control clusters and servers that are connected to the same first network domain; The method includes: Based on the second network configuration information, connect at least one server corresponding to the same control cluster or the same user to the second network domain corresponding to the at least one server; the second network domain is obtained by dividing the second network provided by the second network switch system; Provide data forwarding services between different servers that are connected to the same second network domain.

18. A method for providing cloud computing services, characterized in that, A scheduling system applied to a cloud computing service system, where the cloud computing service system further includes a computing cluster, a first network switch system, and a plurality of control clusters; the computing cluster includes a plurality of servers; the first network provided by the first network switch system is divided into a plurality of first network domains; the plurality of control clusters are connected to different first network domains, and at least one server corresponding to each of the plurality of control clusters is connected to the corresponding first network domain. Data forwarding is supported between the control clusters and servers connected to the same first network domain. The method includes: Detecting a switching operation for a target server and determining a target control cluster that requests the switch. Based on the target control cluster, sending a switching request to the first network switch system, so that the first network switch system responds to the switching request, determines the target control cluster, and switches the target server from the original control cluster to the first network domain corresponding to the target control cluster.

19. A computer-readable storage medium, characterized in that, Stored thereon is a computer program, which when executed by a processing component, implements the cloud computing service providing method according to any one of claims 16-18.

20. A computer program product, characterized in that, Including a computer program or instruction, which when executed by a processing component, implements the cloud computing service providing method according to any one of claims 16-18.

Citation Information

Patent Citations

  • Cloud computing distributed service cluster system and method of using the system

    CN105812488A

  • Resource access method, device and system under server-free architecture and storage medium

    CN112019475A

  • Multi-cluster service system, service access and information configuration method, equipment and medium

    CN114726827A

  • Network resource management method and device, equipment and storage medium

    CN119109928A

  • Multi-cluster service system, service access method, information configuration method, device, and medium

    WO2023185938A1