Expansion method of distributed system and related device
By calling the deployment information interface and parsing deployment information in the distributed reinforcement learning framework, the configuration parsing files are generated, and the automated deployment of custom components and processes is achieved, which solves the problem that existing frameworks cannot be expanded and improves scaling efficiency and reliability.
Patent Information
- Application Number
- CN202410020386.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
The existing distributed reinforcement learning framework cannot support the scaling of custom components and processes, resulting in inefficiency and poor scalability in complex distributed operation scenarios.
By calling the deployment information interface, the deployment information of the target component is analyzed to generate a configuration parsing file, and the target components are configured in the container management cluster based on the workload information, supporting the automated deployment and management of custom components and processes.
Improves the scaling efficiency and reliability of distributed systems, reduces manual configuration time and errors, ensures consistency and repeatability in different environments and clusters, and improves deployment and configuration efficiency.
Smart Images

Figure CN120256012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distributed reinforcement learning, and particularly to an extension method for a distributed system and related devices. Background Art
[0002] In reinforcement learning, in order to achieve more effective and complex learning and decision-making, a reinforcement learning system usually includes multiple different types of components. Taking an artificial intelligence-based game generation scenario as an example, in a reinforcement learning system for generating games, some components are used to run the core logic of the game, some components are used to predict the actions of virtual objects in the game, and some components are used to perform model training. In distributed reinforcement learning, these components operate in a distributed manner.
[0003] However, with the in-depth research on distributed reinforcement learning, more and more complex distributed operation scenarios have emerged. However, the existing distributed reinforcement learning frameworks cannot support the extension of custom components. If specific modifications are made to the existing frameworks for each complex distributed operation scenario, it is time-consuming and laborious, and the efficiency is low, and the subsequent scalability is also very poor. Summary of the Invention
[0004] The embodiments of this application provide an extension method for a distributed system and related devices. By reading and parsing the deployment information of a target component and configuring the target component according to the workload information in the target component configuration parsing file obtained by parsing, the extension of custom components in the distributed reinforcement learning framework is realized, and the target component is automatically deployed and managed through a container management cluster, improving the efficiency and reliability of the extension.
[0005] One aspect of this application provides an extension method for a distributed system, including:
[0006] Invoking a deployment information interface to query the deployment information of a target component, where the deployment information of the target component is used to represent the information for configuring the target component in a container management cluster, and the deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task;
[0007] Parsing the deployment information of the target component to generate a target component configuration parsing file, where the target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is obtained by parsing the target component attribute information;
[0008] Configuring the target component in the container management cluster according to the workload information in the target component configuration parsing file.
[0009] Another aspect of the present application provides an expansion device for a distributed system, including: a deployment information query module for a target component, a deployment information parsing module for the target component, and a target component configuration module; specifically:
[0010] The deployment information query module for the target component is used to call the deployment information interface to query the deployment information of the target component. Among them, the deployment information of the target component is used to characterize the information for configuring the target component in the container management cluster. The deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task;
[0011] The deployment information parsing module for the target component is used to parse the deployment information of the target component to generate a target component configuration parsing file. Among them, the target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is obtained by parsing the target component attribute information;
[0012] The target component configuration module is used to configure the target component in the container management cluster according to the workload information in the target component configuration parsing file.
[0013] In another implementation manner of the embodiment of the present application, the target component configuration module is further used for:
[0014] Read the target component configuration parsing file to obtain M container configuration information, where the M container configuration information is used to characterize the configuration information corresponding to configuring M containers in the target component, M≥1;
[0015] Configure the M containers corresponding to the target component in the container management cluster according to the M container configuration information.
[0016] In another implementation manner of the embodiment of the present application, the expansion device of the distributed system further includes: a first target process configuration module; specifically, the first target process configuration module is used for:
[0017] Read the target component configuration parsing file to obtain N first target process configuration information, where the N first target process configuration information is used to characterize the configuration information corresponding to configuring N first target processes in M containers. The configuration information corresponding to the N first target processes corresponds to N container identification information. The container identification information is used to indicate the container for configuring the first target process. At least one first target process is configured in each container, and the N first target process configuration information carries the configured container identification information, N≥M;
[0018] Configure the N first target processes in the M containers according to the N container identification information.
[0019] In another implementation manner of the embodiment of the present application, the expansion device of the distributed system further includes: a second target process configuration module; specifically, the second target process configuration module is used for:
[0020] Call the deployment information interface to query the deployment information of the second target process, where the deployment information of the second target process carries specified component identification information, and the deployment information of the second target process is used to represent the information for configuring the second target process;
[0021] Parse the deployment information of the second target process to generate a second target process configuration parsing file;
[0022] Deploy the second target process in the specified component corresponding to the specified component identification information in the container management cluster according to the second target process configuration parsing file.
[0023] In another implementation manner of the embodiment of the present application, the second target process configuration module is further used for:
[0024] Read the second target process configuration parsing file to obtain K service port configuration information, where the K service port configuration information is used to represent the port configuration information for configuring K servers in the second target process, and K≥1;
[0025] Configure the K servers in the second target process according to the K service port configuration information.
[0026] In another implementation manner of the embodiment of the present application, the second target process configuration module is further used for:
[0027] Obtain remote server call information, where the remote server call information is used to represent the information for connecting to the remote server;
[0028] Obtain the operation permission of the remote server according to the remote server call information;
[0029] After obtaining the operation permission of the remote server, control the remote server to parse the deployment information of the second target process to generate a second target process configuration parsing file.
[0030] In another implementation manner of the embodiment of the present application, the expansion device of the distributed system further includes: a client configuration module; specifically, the client configuration module is used for:
[0031] Read the target component configuration parsing file to obtain every P client configuration information, where the P client configuration information is used to represent the port configuration information for configuring P clients in N first target processes, and the P client configuration information corresponds to P first target process identification information, and P≥1;
[0032] Configure the P clients in the corresponding first target process according to the P client configuration information and the P first target process identification information.
[0033] In another implementation manner of the embodiment of the present application, the extension device of the distributed system further includes: a process communication configuration module; specifically, the process communication configuration module is used for:
[0034] Call the deployment information interface to query the process communication relationship deployment information, where the process communication relationship deployment information includes the first target process identification information of the first target process and the second target process identification information of the second target process, and the process communication relationship deployment information is used to indicate the establishment of the communication relationship between the first target process and the second target process;
[0035] Parse the process communication relationship deployment information to generate a process communication relationship configuration parsing file;
[0036] Establish a communication connection between the first target process and the second target process according to the process communication relationship configuration parsing file.
[0037] In another implementation manner of the embodiment of the present application, the process communication configuration module is further used for:
[0038] Obtain the first target process port information corresponding to the first target process according to the first target process identification information, and obtain the second target process port information corresponding to the second target process according to the second target process identification information;
[0039] Generate a first target process node according to the first target process port information, and generate a second target process node according to the second target process port information;
[0040] Establish a communication connection between the first target process node and the second target process node according to the process communication relationship configuration parsing file.
[0041] In another implementation manner of the embodiment of the present application, the process communication configuration module is further used for:
[0042] Obtain the port information of the target client in the first target process according to the first target process identification information, and obtain the port information of the target server in the second target process according to the second target process identification information,
[0043] Generate a target client node according to the port information of the target client, and generate a target server node according to the port information of the target server;
[0044] Establish a communication connection between the target client node and the target server node according to the process communication relationship configuration parsing file.
[0045] In another implementation manner of the embodiments of the present application, the extension device of the distributed system further includes: a container instance generation module; specifically, the container instance generation module is configured to:
[0046] Generate an image file of the target component according to the target component configuration parsing file;
[0047] Generate M container instances corresponding to M containers according to the image file, and implement the target service corresponding to the target component through the M container instances.
[0048] On the other hand, the present application provides a computer device, including:
[0049] A memory, a transceiver, a processor, and a bus system;
[0050] Wherein, the memory is used to store programs;
[0051] The processor is used to execute the programs in the memory, including executing the methods in the above aspects;
[0052] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate with each other.
[0053] On the other hand, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the methods in the above aspects.
[0054] On the other hand, the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.
[0055] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0056] The present application provides an extension method for a distributed system and related devices. The method includes: First, call the deployment information interface to query the deployment information of the target component. The deployment information of the target component is used to represent the information for configuring the target component in the container management cluster. The deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task. Then, parse the deployment information of the target component to generate a target component configuration parsing file. The target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is parsed according to the target component attribute information. Finally, configure the target component in the container management cluster according to the workload information in the target component configuration parsing file. The extension method for the distributed system provided by the embodiments of the present application automatically generates a target component configuration parsing file by calling the deployment information interface and parsing the deployment information, and configures the target component in the container management cluster according to the workload information, reducing the time and errors of manual configuration and improving the efficiency of deployment and configuration. Description of the Drawings
[0057] Figure 1 It is a schematic architecture diagram of an extension system for a distributed system provided by an embodiment of the present application;
[0058] Figure 2 It is a flowchart of an extension method for a distributed system provided by an embodiment of the present application;
[0059] Figure 3 It is a flowchart of an extension method for a distributed system provided by another embodiment of the present application;
[0060] Figure 4 It is a flowchart of an extension method for a distributed system provided by another embodiment of the present application;
[0061] Figure 5 It is a structural diagram of a target component provided by an embodiment of the present application;
[0062] Figure 6 It is a flowchart of an extension method for a distributed system provided by another embodiment of the present application;
[0063] Figure 7 It is a flowchart of an extension method for a distributed system provided by another embodiment of the present application;
[0064] Figure 8 It is a flowchart of an extension method for a distributed system provided by another embodiment of the present application;
[0065] Figure 9 It is a structural diagram of a specified component provided by an embodiment of the present application;
[0066] Figure 10 Flow chart of the extension method for the distributed system provided by another embodiment of the present application;
[0067] Figure 11 Flow chart of the extension method for the distributed system provided by another embodiment of the present application;
[0068] Figure 12 Flow chart of the extension method for the distributed system provided by another embodiment of the present application;
[0069] Figure 13 Flow chart of the extension method for the distributed system provided by another embodiment of the present application;
[0070] Figure 14 Schematic diagram of the process relationship configuration provided by a certain embodiment of the present application;
[0071] Figure 15 Schematic diagram of the process relationship configuration process provided by a certain embodiment of the present application;
[0072] Figure 16 Schematic diagram of the process relationship configuration process provided by a certain embodiment of the present application;
[0073] Figure 17 Schematic diagram of the process relationship configuration process provided by a certain embodiment of the present application;
[0074] Figure 18 Flow chart of the extension method for the distributed system provided by yet another embodiment of the present application;
[0075] Figure 19 Architecture diagram of the system for implementing the extension method of the distributed system provided by a certain embodiment of the present application;
[0076] Figure 20 Framework diagram of the distributed deep learning system provided by a certain embodiment of the present application;
[0077] Figure 21 Framework diagram of the distributed deep learning system after extension provided by a certain embodiment of the present application;
[0078] Figure 22 Schematic diagram of the structure of the extension device for the distributed system provided by a certain embodiment of the present application;
[0079] Figure 23 Schematic diagram of the structure of the extension device for the distributed system provided by another embodiment of the present application;
[0080] Figure 24 Schematic diagram of the structure of the extension device for the distributed system provided by another embodiment of the present application;
[0081] Figure 25 Schematic structural diagram of an expansion device for a distributed system provided by another embodiment of the present application;
[0082] Figure 26 Schematic structural diagram of an expansion device for a distributed system provided by another embodiment of the present application;
[0083] Figure 27 Schematic structural diagram of an expansion device for a distributed system provided by still another embodiment of the present application;
[0084] Figure 28 Schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0085] The embodiment of the present application provides an expansion method for a distributed system. By reading and parsing the deployment information of a target component, and configuring the target component according to the workload information in the target component configuration parsing file obtained by parsing, the expansion of custom components in a distributed reinforcement learning framework is realized. The target component is automatically deployed and managed through a container management cluster, improving the efficiency and reliability of the expansion.
[0086] In the specification, claims and above-mentioned drawings of the present application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, terms such as "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0087] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of that module or unit.
[0088] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0089] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0090] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0091] To facilitate the understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application are explained here first:
[0092] Reinforcement Learning: It is a branch of machine learning that mainly studies how an agent can achieve specific goals through continuous trial and error and learning. In reinforcement learning, the agent receives reward or punishment signals from the environment to learn how to select the optimal action strategy to maximize long-term rewards. The basic principle of reinforcement learning is based on dynamic programming and Markov decision processes. In dynamic programming, the agent traverses the state space to select the optimal action strategy to maximize future rewards. In the Markov decision process, the states and actions of the agent are randomly generated by a Markov chain, and the agent needs to select the optimal action strategy according to the probability distribution of the current state and action. Reinforcement learning has applications in many fields, such as robot control, game players, autonomous driving cars, etc. In these applications, reinforcement learning can help the agent autonomously learn and make decisions in an unknown or dynamic environment to achieve specific goals.
[0093] Distributed Reinforcement Learning: It is an implementation method of reinforcement learning that distributes the training tasks across multiple computing nodes to improve the training efficiency and effect. In distributed reinforcement learning, multiple agents interact in a distributed environment and achieve common goals through cooperation and competition. In traditional reinforcement learning algorithms, the interaction between agents is usually achieved by sharing states. However, in a distributed environment, the states may be distributed among different agents, so an effective method is needed to coordinate and share state information. The application scenarios of distributed reinforcement learning include multi-agent cooperation, distributed robot control, distributed decision-making, etc. In these applications, distributed reinforcement learning can help agents autonomously learn and make decisions in a distributed environment to achieve better performance and efficiency.
[0094] Distributed Reinforcement Learning Components: They together constitute the basic framework of distributed reinforcement learning, enabling it to handle large-scale and complex problems and achieve more efficient and scalable learning. Common distributed reinforcement learning components include:
[0095] 1) Agents Component: The Agents component is the entity that executes actions and interacts with the environment. In distributed reinforcement learning (DRL), the Agents component can be distributed across different computing nodes, and each Agents component is responsible for executing a part of the learning task.
[0096] 2) Learner Component: In distributed reinforcement learning (DRL), the Learner component is the entity responsible for executing the learning task. It is usually responsible for collecting experience data from the environment and using this data to update the agent's policy or model.
[0097] 3) Actor Component: In distributed reinforcement learning (DRL), the Actor component is the entity responsible for executing actions and interacting with the environment. It is usually an agent that can select appropriate actions based on the current environmental state and policy.
[0098] These components collaborate to complete the training of a reinforcement learning. In a distributed scenario, each component runs in multiple containers.
[0099] User Code: The user-defined logic executed in the Agent component, which is part of the Agent component process and is used to customize the processing of game state data.
[0100] Kubernetes (K8s): An open-source container orchestration platform used for automating the deployment, scaling, and management of containerized applications. It provides a scalable, elastic, and highly available platform for running distributed systems such as microservice architectures. Kubernetes uses containerization technology to isolate applications and their dependencies, thus enabling flexible deployment and management of containers. Specifically, multiple containers are combined to form a cluster, and then K8s manages this cluster, including cluster creation, scaling up, scaling down, deletion, etc.
[0101] Container: A lightweight virtualization technology that can package a program and its dependencies together for running in different environments. Specifically, a container is a technology for packaging and deploying applications. It packages an application and its dependencies in a lightweight, portable, and self-contained file called a container image. The container image can run in different environments without recompiling or modifying the code.
[0102] Pod: In Kubernetes, a Pod is a combination of one or more containers, as well as a collection of resources and specifications related to these containers. It is the smallest deployable unit in Kubernetes and is also the basis for other Kubernetes objects (such as ReplicaSet, Deployment, etc.). A pod can contain one or more containers.
[0103] Workload: In Kubernetes, Workload refers to the workload, which means the applications or services running in the Kubernetes cluster. Workloads in Kubernetes can include various types of applications, such as web applications, databases, message queues, caches, etc. In Kubernetes, a workload is usually composed of one or more Pods, and these Pods jointly execute the tasks of the application. Kubernetes provides various types of workload objects, such as Deployments, ReplicaSets, StatefulSets, DaemonSets, etc., and these objects can be used to manage and deploy different types of workloads. Each component corresponds to a workload during operation.
[0104] Component process: It represents the component program code after running. In computer science, a component process refers to an independent program unit that performs a specific task. It is a type of process in the operating system and usually collaborates with other component processes to achieve more complex system functions. Component processes are usually developed as part of a larger software system. They can be libraries, services, drivers, or other reusable code modules. Each component process has its own address space, execution context, and resource allocation. Component processes interact through inter-process communication (IPC) mechanisms, such as message passing, shared memory, or pipes. This communication allows component processes to work together to achieve the overall goals of the system. In a distributed system, component processes can run on different computers and communicate and collaborate through the network. This architecture of distributed component processes enables the system to scale better and be fault-tolerant.
[0105] Custom component: A component that is not built into the framework. In a distributed scenario, each custom component also runs in multiple containers.
[0106] Custom process: A running process created by the execution of non-framework code. In a distributed scenario, a custom process can run in the container belonging to the framework component or in the process belonging to the custom component.
[0107] In reinforcement learning, there are multiple different types of components, such as some that run the core logic of the game, some that do action prediction, and some that do model training. In distributed reinforcement learning, these components will run in multiple containers in a distributed manner, and each container executes the corresponding process.
[0108] However, with the in-depth research on distributed reinforcement learning, more and more complex distributed operation scenarios have emerged. Developers need to create custom components and run custom processes in tasks, and conduct data communication between different processes. For example, in a shooting game, depth maps or height maps are required for complex calculations, which are run as independent processes in existing component containers or in independent components. At the same time, data communication also occurs between the processes in the existing components and the custom processes to transfer the input and output data of the calculations.
[0109] However, the existing distributed reinforcement learning frameworks cannot support such custom extensions. For example, in the distributed framework Avatar, developers have no way to create custom components and can only select different built-in components of the framework according to the needs of the scenario to cooperate in completing the operation of the task. At the same time, the processes running in the components are also specified by the framework itself, and developers cannot create custom processes. In the distributed framework Ray, although the distributed framework Ray supports developers to create custom processes, it does not support developers to deploy and create custom components. If specific modifications are made to the existing frameworks for each complex distributed operation scenario, it is time-consuming and laborious, and the efficiency is low, and the subsequent scalability is also very poor.
[0110] To address the above problems, the embodiments of the present application provide an extension method for a distributed system and related devices. The method includes: First, call the deployment information interface to query the deployment information of the target component, where the deployment information of the target component is used to represent the information for configuring the target component in the container management cluster, and the deployment information of the target component includes target component attribute information. Then, parse the deployment information of the target component to generate a target component configuration parsing file, where the target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is obtained by parsing the target component attribute information. Finally, configure the target component in the container management cluster according to the workload information in the target component configuration parsing file. The extension method for the distributed system provided by the embodiments of the present application enables developers to define the deployment information of components through parameter configuration without complex modification of the existing framework. By calling the deployment information interface and parsing the deployment information, a target component configuration parsing file is automatically generated, and the target component is configured in the container management cluster according to the workload information, reducing the time and errors of manual configuration and improving the efficiency of deployment and configuration.
[0111] For ease of understanding, please refer to Figure 1 , Figure 1 which is the application environment diagram of the extension method for the distributed system in the embodiments of the present application, as shown in Figure 1As shown in the figure, the method for expanding a distributed system in an embodiment of the present application is applied to an expansion system of the distributed system. The expansion system of the distributed system includes: a server and a terminal device; wherein, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this here.
[0112] The server first calls the deployment information interface to query the deployment information of the target component. Among them, the deployment information of the target component is used to represent the information for configuring the target component in the container management cluster, and the deployment information of the target component includes the target component attribute information; then, the server parses the deployment information of the target component to generate a target component configuration parsing file, where the target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is parsed according to the target component attribute information; finally, the server configures the target component in the container management cluster according to the workload information in the target component configuration parsing file.
[0113] Next, the method for expanding the distributed system in the present application will be introduced from the perspective of the server. Please refer to Figure 2 , the method for expanding the distributed system provided by the embodiment of the present application includes: step S110 to step S130. Specifically:
[0114] S110. Call the deployment information interface to query the deployment information of the target component.
[0115] Among them, the deployment information of the target component is used to represent the information for configuring the target component in the container management cluster, the deployment information of the target component includes the target component attribute information, and the target component is used to run a distributed reinforcement learning task.
[0116] It is understandable that when developers add custom target components to a container management cluster (K8s cluster), they need to fill in the deployment information of the custom target components. For example, developers can fill in the deployment information of the custom target components through the client and pull the deployment information of the target components by calling the deployment information interface. The target components are used to run distributed reinforcement learning tasks. For example, in the game field, the target components can run the core logic of the game, or be used to predict the actions of virtual objects in the game, or be used for reinforcement learning tasks such as training game models. The deployment information refers to the information for configuring the target components in the container management cluster. The deployment information specifically describes the characteristics and configuration requirements of the target components. The deployment information includes the attribute information of the target components, etc. Through the deployment information of the target components, the deployment requirements of the target components in the container management cluster can be understood, such as the required resources, network configuration, storage requirements, etc. Specifically, the deployment information includes the type of the workload of the target components, the supported system of the target component image (such as the Linux system), the number of pods in the target component, the number of containers in each pod, the startup command and startup command parameters of the containers, the address of the target component image, the number of CPU cores of the containers in the target component, the container memory size, and so on. The deployment information of the target components filled in by developers can be in the format of an information table.
[0117] By calling the deployment information interface, query and obtain the deployment information of the target components filled in by developers. The queried deployment information of the target components will be used to guide and configure the actual deployment process of the target components in the container management cluster. The deployment information can help ensure that the target components can run correctly and be integrated with other components and services. By obtaining the deployment information of the target components, proper configuration and deployment can be carried out in the container management cluster. These deployment information include the attribute information of the target components, which is crucial for the successful deployment and operation of the target components.
[0118] S120. Parse the deployment information of the target components to generate a target component configuration parsing file.
[0119] Among them, the target component configuration parsing file includes the workload information of the target components, and the workload information of the target components is parsed from the target component attribute information.
[0120] It is understandable that after querying and obtaining the deployment information of the target component filled in by the developer, it is necessary to parse the deployment information of the target component to generate a target component configuration parsing file. Specifically, it is necessary to parse the attribute information of the target component in the deployment information to obtain the workload information (workload) of the target component. By parsing the deployment information of the target component, the information related to the workload is extracted. These workload information include the computing resources required by the target component (such as the number of CPU cores and the size of memory), network resources (such as port mapping and network bandwidth), and other related configuration parameters. The generated target component configuration parsing file will contain these workload information, so that subsequent according to the workload information in the target component configuration parsing file, the target component can be correspondingly configured and deployed in the container management cluster. By parsing the deployment information and extracting the parameters related to the workload, it is ensured that the target component is correctly configured and deployed in the container management cluster.
[0121] S130. Configure the target component in the container management cluster according to the workload information in the target component configuration parsing file.
[0122] It is understandable that according to the workload information in the parsing file, the target component is added and configured in the container management cluster (K8s cluster). Specifically, the configuration operations in the container management cluster are guided by the workload information in the target component configuration parsing file. Specifically, according to the workload information, appropriate resources and containers are created in the container management cluster to meet the requirements of the target component.
[0123] When configuring the target component, the custom target component can be configured from aspects such as resource allocation, network configuration, storage configuration, and container configuration. Specifically, in the process of resource allocation, according to the computing resource requirements in the workload information, an appropriate number of CPU cores, memory and other computing resources are allocated in the container management cluster to ensure that the target component has sufficient computing power to run the distributed reinforcement learning task. In the process of network configuration, according to the network resource requirements in the workload information, network connections are created and port mapping is set in the container management cluster to ensure that the target component can communicate with other components or the external network. In the process of storage configuration, according to the storage requirements in the workload information, appropriate storage resources are configured in the container management cluster, such as creating persistent volumes and mounting file systems, to meet the data storage requirements of the target component. In the process of container configuration, according to other configuration parameters in the workload information, the startup command, environment variables, image address, etc. of the container are set in the container management cluster to ensure that the target component can correctly run the distributed reinforcement learning task.
[0124] By configuring according to the workload information in the target component configuration parsing file, an environment matching the requirements of the target component is created in the container management cluster. In this way, the target component can run correctly in the container management cluster and be integrated with other components and services. According to the workload information of the target component, corresponding configuration operations are performed in the container management cluster to ensure that the target component can run normally in the cluster and meet its requirements. By parsing the workload information in the file and applying it to the cluster configuration, an automated and standardized component deployment process is achieved.
[0125] The extension method of the distributed system provided by the embodiment of the present application supports the deployment of custom components in a Kubernetes (K8s) cluster. Developers can define the deployment information of the components in the way of parameter configuration without complex transformation of the existing framework. By calling the deployment information interface and parsing the deployment information, a configuration parsing file of the target component is automatically generated, and the target component is configured in the container management cluster according to the workload information, realizing the automated deployment of the target component, reducing the time and errors of manual configuration, reducing the possibility of manual intervention and errors, and improving the efficiency of deployment and configuration; by using the configuration file to describe the requirements and configuration of the target component, the consistency and repeatability during deployment in different environments and clusters are ensured, and the reliability and stability of the deployment are improved;
[0126] By querying and parsing the deployment information of the target component, different extension requirements can be flexibly adapted; the automated deployment process can save time and effort, improving efficiency and productivity; resource allocation and configuration according to the workload information can better manage computing resources, network resources and storage resources to ensure that the target component obtains appropriate resource support in the container management cluster.
[0127] In the Figure 2 corresponding optional embodiment of the extension method of the distributed system provided by the embodiment of the present application, please refer to Figure 3 , step S130 further includes sub-steps S131 to S132.
[0128] Specifically:
[0129] S131. Read the target component configuration parsing file to obtain M container configuration information.
[0130] Among them, the M container configuration information is used to represent the configuration information corresponding to configuring M containers in the target component, and M≥1.
[0131] It is understandable that by reading the target component configuration parsing file, M container configuration information is obtained. The container configuration information includes the name of the container, resource requirements (such as CPU, memory, etc.), image address, startup command, environment variables, etc. Reading the target component configuration parsing file is a key operation. The configuration parsing file is a file containing all the container configuration information in the target component. By reading the configuration parsing file, the detailed configuration of the container can be obtained, including the name of the container, resource requirements (such as CPU, memory, etc.), image address, startup command, environment variables, etc. The obtained M container configuration information is used to characterize the specific configuration details of the M containers configured in the target component. M represents the number of containers, and M is greater than or equal to 1, meaning that the target component can be configured with at least one container. These container configuration information are necessary for the target component to run correctly in the container management cluster. Each container has its unique configuration requirements, and these information ensure that the container can correctly set the environment, obtain the required resources, and execute the corresponding tasks as expected when starting. By obtaining these container configuration information, in subsequent steps, these containers can be created, started, and managed in the container management cluster according to the requirements of the target component to achieve the functional and performance requirements of the target component.
[0132] S132. Configure the M containers corresponding to the target component in the container management cluster according to the M container configuration information.
[0133] It is understandable that the process of configuring containers according to the container configuration information includes: creating containers, allocating resources, pulling images, setting environment variables, setting connections and dependency management, executing startup commands, configuring networks, and configuring log collection mechanisms. Specifically:
[0134] Create containers: Create corresponding container instances in the container management cluster according to the container configuration information (such as name and resource requirements). Each container instance is initialized according to its configuration information, including pulling the corresponding image, setting environment variables, executing startup commands, etc.
[0135] Allocate resources: Allocate appropriate computing resources (such as the number of CPU cores, memory size, etc.) for each container according to the resource requirements of the container to meet the performance requirements of the target component.
[0136] Pull images: Pull the corresponding images from the appropriate image repository according to the image addresses of the containers. The images contain the application programs and related dependencies of the containers, ensuring that the containers can run normally when starting.
[0137] Set environment variables: Set appropriate environment variables according to the container configuration information to meet the running requirements of the application programs. These environment variables can include database connection strings, configuration file paths, etc.
[0138] Connection and Dependency Management: If there are connections or dependencies between containers in the target component, corresponding configurations need to be made in the container management cluster to ensure that the containers can communicate and cooperate correctly.
[0139] Execute the startup command: According to the startup command of the container, execute the corresponding commands or scripts in the container to start the application or service. The startup command can include starting a daemon process, executing a script, etc.
[0140] Configure the network: According to the requirements of the target component, configure appropriate network connections for the containers, including creating networks, setting port mappings, etc., to ensure that the containers can communicate with other components or the external network.
[0141] Configure the log collection mechanism: Configure the log collection mechanism of the container to observe the running status of the container in real time and collect log information for troubleshooting and performance optimization.
[0142] The extension method of the distributed system provided by the embodiments of the present application enables each container in the target component to be finely configured and deployed according to its configuration information by performing container configuration on the target component, ensuring that the target component can run normally in the container management cluster and meet the expected function and performance requirements; by reading the target component configuration parsing file to obtain the container configuration information and automatically configuring and starting the containers in the container management cluster according to this information, the automatic deployment of the target component is realized, reducing the possibility of manual intervention and errors; using the configuration parsing file to describe the container requirements and configurations of the target component ensures consistency and repeatability when deploying in different environments and clusters, improving the reliability and stability of the deployment; the configuration parsing file can be customized and extended according to different requirements to adapt to different deployment scenarios and environments, providing flexibility and scalability; the automated container configuration and deployment process can save time and effort, improving efficiency and productivity; resource allocation and configuration according to the container configuration information can better manage computing resources, network resources, and storage resources, ensuring that the target component obtains appropriate resource support in the container management cluster; it helps to improve the deployment and management efficiency of the distributed system, reduce errors, and improve reliability.
[0143] In the Figure 3 corresponding optional embodiment of the extension method of the distributed system provided by the embodiment of the present application, please refer to Figure 4 , after sub-step S132, there are also sub-steps S133 to S134.
[0144] Specifically:
[0145] S133. Read the target component configuration parsing file to obtain N pieces of first target process configuration information.
[0146] Among them, the N first target process configuration information is used to represent the configuration information corresponding to the N first target processes configured in the M containers. The configuration information corresponding to the N first target processes corresponds to N container identification information. The container identification information is used to indicate the container in which the first target process is configured. At least one first target process is configured in each container. The N first target process configuration information carries the configured container identification information, and N≥M.
[0147] It can be understood that the main purpose of this step is to read the configuration information related to the first target process from the target component configuration parsing file. The target component configuration parsing file is a file containing the detailed configuration information of the target component. By parsing this file, the configuration parameters related to the first target process can be extracted. The obtained N first target process configuration information is used to represent the specific configuration details of the N first target processes configured in the M containers. Specifically, by reading the target component configuration parsing file, N first target process configuration information is obtained. N represents the number of first target processes. Each first target process configuration information contains specific configuration parameters related to the process, such as process name, start command, resource requirements, etc. Each first target process has its corresponding configuration information, and these configuration information will guide how to configure the first target process in the container. The configuration information may include the name of the process, start command, resource requirements (such as CPU, memory, etc.), environment variables, file paths, etc. The configuration information of each first target process is associated with the container identification information. These first target process configuration information is associated with the container identification information. The container identification information is used to indicate in which container each first target process should be configured. Each container has a unique identification information, such as the name of the container, identifier or other information that uniquely identifies the container. At least one first target process is configured in each container, which means that a container can run multiple first target processes at the same time. By associating the configuration information of the first target process with the container identification information, it can be ensured that each process is correctly configured into the corresponding container. Each container provides an independent execution environment, including file system, network, resource limits, etc. By using containers, process isolation and resource limits can be achieved, improving the portability and security of application programs. The number of containers can be adjusted according to requirements to meet the performance requirements and resource limits of the target component. By configuring the first target processes in multiple containers, load balancing and horizontal scaling can be achieved. By reading the target component configuration parsing file, obtaining the configuration information of each first target process, and associating each process with the container identification information, to guide the correct configuration of these processes into the corresponding containers. At least one first target process is configured in each container, and the number of first target processes is at least equal to the number of containers. These configuration information carry the container identification information for correct deployment and configuration in the container management cluster.
[0148] S134. Configure N first target processes in M containers according to N container identification information.
[0149] It can be understood that the main purpose of this step is to configure the first target process into the corresponding container according to the container identification information. The container identification information is used to uniquely identify each container, which can be the name, identifier or other relevant information of the container. Through the container identification information, it can be determined which specific container each first target process should be configured into. Configure N first target processes in M containers according to the container identification information. The configuration process includes operations such as creating a process, setting an execution command, and allocating resources to ensure that each process runs normally in the container. At least one first target process is configured in each container. In step S133, the configuration information of each first target process and the corresponding container identification information have been obtained. Therefore, in this step, the process can be allocated to the corresponding container according to this information. By executing step S134, the first target process is associated with the container, and corresponding configuration operations are performed in the container, so that the target component can run correctly in the container management cluster.
[0150] For ease of understanding, please refer to Figure 5 , Figure 5 which is the structural diagram of the target component provided by the embodiment of the present application. When adding a custom target component to the container management cluster, it is necessary to create a workload corresponding to the target component during the startup process. Therefore, the developer needs to fill in the deployment information of the custom target component, and this deployment information includes the workload information required to configure the workload. If the workload of the custom target component needs to automatically start some built-in custom first target processes after configuration, the developer also needs to fill it in the deployment information. As Figure 5 shown, the workload of the custom target component can include at least one pod, each pod includes a container, that is, the number of pods in the workload of the target component is equal to the number of containers, and at least one custom first target process can be started in each container. Further, each first target process can include at least one accessed server, and each server can provide at least one server port for external access.
[0151] The extension method of the distributed system provided by the embodiment of the present application, by configuring the first target process in the container of the target component, is a key step in the configuration process of the target component. It involves extracting the configuration information of the first target process from the configuration parsing file and allocating these processes to the container for configuration, so as to realize the flexible deployment and management of the processes of the target component in the container management cluster; by reading the configuration parsing file of the target component, obtaining the configuration information of the first target process, and associating it with the container identification information, it is possible to realize automatic process deployment and container configuration, reduce the possibility of manual intervention and errors, and improve the efficiency and accuracy of deployment; at least one first target process is configured in each container, and the computing resources can be reasonably allocated and utilized according to the resource requirements and limitations of the container, avoiding resource waste and conflicts; the number of containers and the number of first target processes can be adjusted according to needs to meet different load and performance requirements, realizing load balancing and horizontal expansion; each first target process runs in an independent container, providing process isolation and resource limitation, improving the security and reliability of the application; associating the first target process with the container identification information can facilitate the management, log collection and troubleshooting of the process, improving the maintainability and management efficiency of the application; the extension method of the distributed system provided by the embodiment of the present application improves the deployment efficiency, resource management ability, scalability and flexibility, isolation and security, and maintainability, which helps to improve the performance and reliability of the distributed system.
[0152] In the Figure 4 corresponding optional embodiment of the extension method of the distributed system provided by the embodiment of the present application, please refer to Figure 6 , the extension method of the distributed system further includes steps S210 to S230.
[0153] Specifically:
[0154] S210. Invoke the deployment information interface to query the deployment information of the second target process.
[0155] Among them, the deployment information of the second target process carries the specified component identification information, and the deployment information of the second target process is used to characterize the information for configuring the second target process.
[0156] It can be understood that by calling the deployment information interface, the configuration information related to the second target process is queried and obtained. The deployment information interface is a functional module that provides the function of querying and obtaining deployment information, and it can interact with other components in the distributed system. The deployment information of the second target process queried includes the name of the process, the startup command, resource requirements (such as CPU, memory, etc.), environment variables, file paths, etc. The deployment information contains the detailed information for configuring and deploying the second target process, and these information include the startup parameters of the process, resource allocation, network configuration, logging, etc. The purpose of the deployment information is to guide the system on how to correctly create, start, and manage the second target process in the container management cluster. The deployment information of the second target process carries the specified component identification information, and the specified component identification information is used to uniquely identify the component to which the second target process belongs. By carrying the specified component identification information, when deploying and configuring the second target process, it can be ensured that it is correctly associated with the corresponding component or module.
[0157] S220. Parse the deployment information of the second target process to generate a second target process configuration parsing file.
[0158] It can be understood that the queried deployment information of the second target process is parsed and processed. By parsing the deployment information, the key configuration parameters and information can be extracted and saved into the second target process configuration parsing file. The configuration parsing file is a file used to record and save the configuration information of the second target process, and it can be used for subsequent deployment and management operations. The result of the parsing includes the specific configuration details of the second target process, such as the startup command, environment variables, file paths, etc. The parsing process is carried out according to the format and structure of the deployment information, and involves operations such as parsing configuration files, extracting key information, and processing parameters. The generated second target process configuration parsing file will contain the parsed configuration information.
[0159] S230. According to the second target process configuration parsing file, deploy the second target process in the specified component corresponding to the specified component identification information in the container management cluster.
[0160] It is understandable that according to the parsed second target process configuration information, the second target process is deployed in the container management cluster. By using container technology, the application program and its dependencies can be packaged into independent containers to achieve isolation and resource limitation. According to the specified component identification information, it is determined in which specific component the second target process is to be deployed. The specified component identification information is used to uniquely identify a specific component in the container management cluster. By parsing the component identification information carried in the file, it can be determined in which specific component the second target process should be deployed, ensuring that the second target process is correctly deployed to the corresponding component environment. The specified component can be an original component in the container management cluster or a custom target component.
[0161] The extension method of the distributed system provided by the embodiments of the present application can achieve flexible management and extension of the target process in the distributed system by calling the deployment information interface, parsing the deployment information, generating a configuration parsing file, and deploying the second target process in the specified component in the container management cluster, ensuring the correct configuration and deployment of the target process, and improving the reliability and maintainability of the system; by calling the deployment information interface and parsing the deployment information, generating a configuration parsing file, and deploying the second target process in the container management cluster according to this file, an automated deployment process is realized, reducing the possibility of manual intervention and errors; the automated deployment process can improve the efficiency and accuracy of deployment, reducing the deployment time and labor costs; by deploying the second target process through the container management cluster, the computing resources can be better managed and utilized, improving the utilization rate and efficiency of resources; the container-based deployment method provides better scalability and flexibility, and can quickly expand or adjust the deployment scale and configuration of the second target process according to needs; container technology provides isolation of processes and limitation of resources, improving the security and reliability of application programs.
[0162] In the Figure 6 corresponding optional embodiment of the extension method of the distributed system provided by the embodiment of the present application, please refer to Figure 7 , step S230 further includes sub-steps S231 to S232.
[0163] Specifically:
[0164] S231. Read the second target process configuration parsing file to obtain K service port configuration information.
[0165] Among them, the K service port configuration information is used to characterize the port configuration information corresponding to K servers configured in the second target process, and K≥1.
[0166] It can be understood that by reading the configuration parsing file, configuration information related to the second target process can be obtained, and the configuration information related to the second target process includes service port configuration information. Specifically, the port configuration information related to each server in the second target process is provided in the configuration parsing file, and each server has a corresponding port configuration information for specifying the port used by the server. By configuring these service ports, it can be ensured that the servers in the second target process can correctly obtain and receive requests from the client.
[0167] S232. Configure the K servers in the second target process according to the K service port configuration information.
[0168] It can be understood that through the obtained K service port configuration information, corresponding server configurations are performed in the second target process. Specifically, according to the port configuration information provided in the configuration parsing file, the correct ports are set for the K servers in the second target process. By configuring the ports of the servers, it can be ensured that the servers can obtain and receive requests from the client on the specified ports.
[0169] The method for expanding a distributed system provided by the embodiments of the present application, by reading the service port configuration information in the configuration parsing file and performing server configuration in the second target process according to this information, can ensure that the server can correctly receive and receive requests from the client; by reading the service port configuration information from the configuration parsing file and configuring the server in the second target process according to this information, can ensure that the server can correctly obtain and receive requests from the client, improving the accuracy and efficiency of deployment; saving the service port configuration information in the configuration parsing file facilitates subsequent maintenance and management. When it is necessary to modify the service port configuration, only the configuration parsing file needs to be updated, without modifying the code or redeploying; by configuring the service port, access to specific ports can be restricted, improving the security of the system.
[0170] In the Figure 6 corresponding optional embodiment of the method for expanding a distributed system provided by the embodiment of the present application, please refer to Figure 8 , step S220 further includes sub-steps S221 to S223.
[0171] Specifically:
[0172] S221. Obtain remote server call information.
[0173] Among them, the remote server call information is used to characterize the information for connecting to the remote server.
[0174] It can be understood that by obtaining the remote server call information, parameters such as the address, port, username, and password of the remote server to be connected can be determined, and this information will be used to establish a secure connection (SSH connection) with the remote server. Specifically, the remote server call information contains the key parameters required to connect to the remote server, including the IP address, SSH port, username, and password of the remote server. By providing this information, it can ensure that the connection with the remote server is correctly established.
[0175] S222. Obtain the operation permission of the remote server according to the remote server call information.
[0176] It can be understood that according to the remote server call information obtained in step S221, use the SSH protocol to establish a connection with the remote server. After the connection is established, the identity can be verified by providing the correct credentials (such as username and password). Once the identity verification is successful, the operation permission of the remote server can be obtained.
[0177] SSH (Secure Shell) is a secure remote login protocol used to connect to a remote server over a network. By using the SSH protocol, an encrypted connection can be established to ensure the security of communication. - In step S222, according to the parameters such as the IP address, port, and username provided by the remote server call information, use an SSH client tool (such as OpenSSH) to establish a connection with the remote server.
[0178] S223. After obtaining the operation permission of the remote server, control the remote server to parse the deployment information of the second target process and generate a second target process configuration parsing file.
[0179] It can be understood that the deployment information of the second target process is parsed on the remote server and the corresponding configuration parsing file is generated. Once the operation permission of the remote server is successfully obtained, appropriate commands and tools can be used to perform the parsing operation. The parsing process involves reading the configuration file of the second target process, extracting relevant parameters, and saving them to the second target process configuration parsing file. By executing the corresponding commands or scripts on the remote server, the deployment information of the second target process is parsed. The purpose of parsing is to generate a second target process configuration parsing file, which will contain the relevant configuration information of the second target process.
[0180] For easy understanding, please refer to Figure 9 . Figure 9It is a structural diagram of a specified component after adding a second target process to a specified component (agent component) provided by an embodiment of the present application. Adding a custom second target process to a component requires specifying the information of the component to be added and the relevant configuration information of the custom second target process. As Figure 9 shown, in addition to the built-in agent server process, the agent component also includes multiple custom second target processes. Each custom second target process may include multiple accessed servers, and each server can provide one or more ports that can be accessed from the outside. Each pod will start the custom second target process according to the same configuration.
[0181] The method for expanding a distributed system provided by an embodiment of the present application obtains the call information and operation permissions of a remote server, parses the deployment information of the second target process on the remote server, and generates a configuration parsing file, which helps to achieve remote management and configuration of specific components in the distributed system; by obtaining the remote server call information, it can accurately connect to the remote server for subsequent operations; using the SSH protocol to establish a connection with the remote server provides an encrypted communication channel to protect the security of data transmission; once the operation permissions of the remote server are obtained, various operations can be performed on the remote server, including parsing the deployment information of the second target process; by performing the parsing operation on the remote server, the computing resources of the remote server can be utilized to accelerate the parsing process and improve the overall deployment efficiency; allowing deployment on different remote servers provides greater flexibility and scalability to adapt to different deployment requirements and environments; the method for expanding a distributed system provided by an embodiment of the present application provides a safe, efficient, and flexible way to connect to a remote server and perform parsing operations on the remote server to support the expansion and management of the distributed system.
[0182] In the Figure 4 corresponding embodiment of the method for expanding a distributed system provided by an embodiment of the present application, in an optional embodiment, please refer to Figure 10 , after sub-step S134, there are also sub-steps S135 to S136.
[0183] Specifically:
[0184] S135. Read the target component configuration parsing file to obtain the configuration information of every P clients.
[0185] Among them, the configuration information of P clients is used to represent the port configuration information corresponding to P clients configured in N first target processes. The configuration information of P clients corresponds to P first target process identification information, and P≥1.
[0186] It can be understood that by reading the configuration parsing file, the configuration information related to each client can be obtained. The configuration information includes the port configuration corresponding to the client, etc. Among the N first target processes, P clients are configured, and each client has its corresponding port configuration information. These information will be used to identify and configure each client. The configuration information of the P clients corresponds to the P first target process identification information, and the configuration information of each client is associated with the specific identification information in the first target process to ensure correct configuration and mapping.
[0187] S136. Configure the P clients in the corresponding first target processes according to the P client configuration information and the P first target process identification information.
[0188] It can be understood that according to the obtained client configuration information and the first target process identification information, the clients are configured in the corresponding first target processes, and the information such as the port configuration of the clients is applied to the corresponding first target processes to ensure that the clients can correctly communicate and interact with the target processes.
[0189] The method for expanding a distributed system provided by the embodiments of the present application achieves the goal of configuring clients in the first target processes. This helps to improve the flexibility and scalability of the system, enabling multiple clients to be configured and managed in different target processes. At the same time, through the association of the configuration parsing file and the target process identification information, it can be ensured that the configuration of the clients matches the specific requirements of the target processes, improving the reliability and stability of the system; by reading the target component configuration parsing file, the port configuration information corresponding to each client can be obtained and associated with the first target process identification information; according to different client configuration information, each client can be configured in the corresponding first target process to meet the needs of different clients; it supports configuring multiple clients, enabling the system to adapt to different application scenarios and requirements; by combining the client configuration information with the first target process identification information, it can be ensured that the clients can correctly communicate and interact with the target processes; it provides the ability to configure and manage multiple clients, realizes the personalized configuration of the clients, improves the flexibility and scalability of the system, and ensures the correct communication and interaction between the clients and the target processes.
[0190] In the Figure 9 corresponding optional embodiment of the method for expanding a distributed system provided by the embodiments of the present application, please refer to Figure 11 , the method for expanding a distributed system further includes steps S310 to S330. Specifically:
[0191] S310. Call the deployment information interface to query the deployment information of the process communication relationship.
[0192] Among them, the process communication relationship deployment information includes the first target process identifier information of the first target process and the second target process identifier information of the second target process. The process communication relationship deployment information is used to indicate the establishment of a communication relationship between the first target process and the second target process.
[0193] It can be understood that by calling the deployment information interface, the process communication relationship deployment information is queried and obtained. The deployment information interface is an interface for querying and obtaining system deployment-related information. By calling this interface, the communication relationship deployment information between the first target process and the second target process can be obtained. The process communication relationship deployment information includes the identifier information of two target processes: the first target process identifier information and the second target process identifier information. The first target process identifier information is used to uniquely identify the first target process, and the second target process identifier information is used to uniquely identify the second target process. The main purpose of the process communication relationship deployment information is to indicate the establishment of a communication relationship between the first target process and the second target process.
[0194] S320. Analyze the process communication relationship deployment information to generate a process communication relationship configuration analysis file.
[0195] It can be understood that the obtained process communication relationship deployment information is analyzed. The analysis process is to convert the content in the deployment information into a form that can be used for subsequent processing. By analyzing, a process communication relationship configuration analysis file can be generated, and this file contains the communication relationship configuration information between the first target process and the second target process.
[0196] S330. Establish a communication connection between the first target process and the second target process according to the process communication relationship configuration analysis file.
[0197] It can be understood that according to the process communication relationship configuration analysis file, a communication connection between the first target process and the second target process is established. The process of establishing the communication connection involves network communication protocols and related operating system services. By using the configuration information in the analysis file, it can be ensured that the first target process and the second target process can communicate and exchange data correctly.
[0198] The extension method of the distributed system provided by the embodiment of the present application can achieve the goals of deploying the communication relationship between processes in the distributed system and establishing a communication connection, which helps to improve the flexibility and scalability of the system, enabling different processes to effectively collaborate and exchange data; meanwhile, through the use of parsing and configuration files, the communication relationship between processes can be conveniently managed and maintained, improving the reliability and stability of the system; by calling the deployment information interface, the communication relationship deployment information between the first target process and the second target process can be quickly obtained, including process identification information, etc. Parse the obtained process communication relationship deployment information and generate a process communication relationship configuration parsing file. According to this file, a communication connection between the first target process and the second target process can be automatically established. By establishing a process communication connection, communication and collaboration between more processes can be supported, enabling the system to better adapt to different application scenarios and requirements; using the configuration information in the process communication relationship configuration parsing file can ensure that the first target process and the second target process can correctly communicate and exchange data, improving the stability and reliability of the system; by querying and obtaining the process communication relationship deployment information in a simple and efficient manner, the automatic configuration of the process communication relationship is realized, improving the scalability and flexibility of the system, and ensuring the correctness and reliability of process communication.
[0199] In the Figure 10 corresponding optional embodiment of the extension method of the distributed system provided by the embodiment of the present application, please refer to Figure 12 , step S330 further includes sub-steps S331 to S333. Specifically:
[0200] S331. Obtain the first target process port information corresponding to the first target process according to the first target process identification information, and obtain the second target process port information corresponding to the second target process according to the second target process identification information.
[0201] It can be understood that according to the first target process identification information and the second target process identification information, the port information corresponding to each target process is obtained. The first target process identification information is used to uniquely identify the first target process, and the second target process identification information is used to uniquely identify the second target process. By using these identification information, the relevant process information database or other data sources can be queried to obtain the corresponding first target process port information and second target process port information.
[0202] S332. Generate a first target process node according to the first target process port information, and generate a second target process node according to the second target process port information.
[0203] It can be understood that, according to the obtained first target process port information and second target process port information, corresponding process nodes are generated. A process node is an abstract concept in a communication relationship graph or other models representing the relationships between processes. By generating the first target process node and the second target process node, the processes can be associated with their corresponding port information, so as to establish communication connections subsequently.
[0204] S333. Establish a communication connection between the first target process node and the second target process node according to the process communication relationship configuration parsing file.
[0205] It can be understood that, according to the process communication relationship configuration parsing file, a communication connection between the first target process node and the second target process node is established. The process communication relationship configuration parsing file contains detailed information on how processes communicate with each other. By parsing this file, information such as parameters, protocols, channels, etc. required for establishing the communication connection can be obtained. Using this information, a communication connection can be established between the first target process node and the second target process node, enabling them to send and receive data from each other.
[0206] The extension method of the distributed system provided by the embodiments of the present application obtains port information according to process identification information, generates corresponding process nodes, and finally establishes a communication connection between the process nodes according to the process communication relationship configuration parsing file. This helps to achieve flexible communication and collaboration between processes in the distributed system. At the same time, by parsing the configuration file, it is convenient to manage and maintain the communication relationships between processes, improving the scalability and maintainability of the system; by obtaining the corresponding port information according to the process identification information, it is convenient to know the ports used by each process for network communication and security management; according to the obtained port information, corresponding process nodes are automatically generated, associating the processes with their port information, facilitating the subsequent establishment and management of communication connections; by establishing a communication connection between the process nodes, it supports communication and collaboration between more processes, enabling the system to better adapt to different application scenarios and requirements; using the configuration information in the process communication relationship configuration parsing file can ensure that processes can communicate and exchange data correctly, improving the stability and reliability of the system; by obtaining and managing the port information of processes, automatic generation of process nodes is achieved, improving the scalability and flexibility of the system and ensuring the correctness and reliability of process communication.
[0207] In the Figure 10 corresponding optional embodiment of the extension method of the distributed system provided by the embodiments of the present application, please refer to Figure 13 , step S330 further includes sub-steps S3301 to S3303.
[0208] Specifically:
[0209] S3301. Obtain the port information of the target client in the first target process according to the first target process identification information, and obtain the port information of the target server in the second target process according to the second target process identification information.
[0210] It can be understood that according to the first target process identification information and the second target process identification information, obtain the port information of the target client and the target server in each target process. The first target process identification information is used to uniquely identify the first target process, and the second target process identification information is used to uniquely identify the second target process. By using these identification information, relevant process information databases or other data sources can be queried to obtain the corresponding target client port information and target server port information.
[0211] S3302. Generate a target client node according to the port information of the target client, and generate a target server node according to the port information of the target server.
[0212] It can be understood that according to the obtained target client port information and target server port information, generate corresponding nodes. The target client node and the target server node are abstract concepts in a communication relationship diagram or other models representing the relationships between processes. By generating these nodes, the processes can be associated with their port information for subsequent establishment of communication connections.
[0213] S3303. Establish a communication connection between the target client node and the target server node according to the process communication relationship configuration parsing file.
[0214] It can be understood that according to the process communication relationship configuration parsing file, establish a communication connection between the target client node and the target server node. The process communication relationship configuration parsing file contains detailed information on how processes communicate with each other. By parsing this file, information such as parameters, protocols, channels, etc. required for establishing a communication connection can be obtained. Using this information, a communication connection can be established between the target client node and the target server node so that they can send and receive data from each other.
[0215] For ease of understanding, please refer to Figure 14 . To support more flexible configuration of the communication relationship between the custom first target process and the second target process, further abstract the first target process and the second target process themselves and the parts related to the communication relationship therein, such as Figure 14 shown Figure 14It is a schematic diagram of process relationship configuration provided by an embodiment of the present application. First, the first target process is abstracted as the first target process node, and the second target process is abstracted as the second target process node. Secondly, the target client that requests to access the service in the first target process is abstracted as the target client node, and the target server that can provide access services in the second target process is abstracted as the target server node. It should be noted that each process node can include only a server or only a client, or both a server and a client, depending on the usage scenario; each client node can access multiple server nodes, and these accessed server nodes can be distributed on different process nodes; multiple process nodes can be included in the same component container.
[0216] For example, please refer to Figures 15 to 17 . Through the method provided by the embodiment of the present application, the original components (Agent component, Learner component, Actor component, Serving component) in the container management cluster are all abstracted as component nodes, and the second target processes customized in each original component are abstracted as second target process nodes. Taking the Agent component as an example, please refer to Figure 15 . The Agent component includes a proxy service process, and the proxy service process includes framework code and user code. First, a customized second target process is configured in the Agent component, and a server is configured in the second target process. Then, a customized target component is configured in the container management cluster, and a first target process is configured in the customized target component. Then, please refer to Figure 16 . The user code in the proxy service process in the Agent component is abstracted as the target client node (user code node), the second target process in the Agent component is abstracted as the second target process node, the target server in the second target process is abstracted as the target server node, and the target server in the customized target component is abstracted as the target server node. When the client or user code in a customized process needs to access the server of another customized process, the developer needs to create one or more clients in the customized process (the first target process or the second target process), and finally define the access relationship between the client and the server. Finally, the access relationship after the process communication configuration is as shown in Figure 17 . Specifically, the process communication configuration can be implemented through the ADK encapsulated in the CustomTopology class. The client is registered in the customized process through register_client, and the registered client is connected to the server of the process node through sdd_client_connection, that is, the access relationship is defined. Finally, the client connected to the server is obtained through build_client.
[0217] The extension method of the distributed system provided by the embodiment of the present application obtains port information according to the process identification information, generates corresponding nodes, and finally establishes communication connections between the nodes by parsing the configuration file according to the communication relationship, which helps to achieve flexible communication and collaboration between processes in the distributed system; at the same time, by parsing the configuration file, the communication relationship between processes can be conveniently managed and maintained, improving the scalability and maintainability of the system; by obtaining the port information of the target client and the target server according to the process identification information, the ports used by each process can be conveniently understood for network communication and security management; according to the obtained port information, the target client node and the target server node are automatically generated, associating the process with its port information, which is convenient for subsequent establishment and management of communication connections; by establishing a communication connection between the target client node and the target server node, communication and collaboration between more processes can be supported, enabling the system to better adapt to different application scenarios and requirements; using the configuration information in the process communication relationship configuration parsing file can ensure that processes can communicate and exchange data correctly, improving the stability and reliability of the system; by providing a flexible and efficient way to obtain and manage the port information of processes, the automatic generation of process nodes is realized, improving the scalability and flexibility of the system, and ensuring the correctness and reliability of process communication.
[0218] In the present application Figure 3 In an alternative embodiment of the extension method of the distributed system provided by the corresponding embodiment of the present application, please refer to Figure 18 , after step S132, steps S137 to S138 are further included. Specifically:
[0219] S137. Generate an image file of the target component according to the target component configuration parsing file.
[0220] It can be understood that, according to the obtained target component configuration parsing file, an image file of the target component is generated. The target component configuration parsing file contains various configuration information of the target component, such as code, dependency libraries, environment variables, etc. By parsing this file, the relevant configuration of the target component can be extracted and used to construct the image file of the target component.
[0221] S138. Generate M container instances corresponding to the M containers according to the image file, and implement the target service corresponding to the target component through the M container instances.
[0222] It is understandable that M container instances are generated according to the generated image file, and the target service corresponding to the target component is implemented through these container instances. The image file is a file that contains the running environment and related code of the target component and can be used to quickly create container instances. According to the number M of container instances to be generated as needed, M container instances are created using the image file. Each container instance is an independent running environment and can independently execute the related code and tasks of the target component. By deploying the code and configuration of the target component into these container instances, the target service corresponding to the target component can be implemented.
[0223] The method for expanding a distributed system provided by the embodiments of the present application generates an image file according to the configuration parsing file of the target component and uses the image file to create container instances to implement the target service. This helps to achieve containerized deployment and microservice architecture, improving the flexibility, scalability, and maintainability of the system; at the same time, through the isolation and resource limitation of container instances, the running environment of the target component can be better managed and controlled, improving the security and stability of the system; generating an image file according to the target component configuration parsing file and using the image file to create container instances can quickly deploy the target service corresponding to the target component, reducing the time and workload of manual configuration and deployment; the image file contains the running environment and related code of the target component, which can ensure that each container instance has the same running environment and configuration, thus ensuring the consistency of the target service in different container instances; by generating multiple container instances, the scale of the target service can be flexibly expanded as needed to meet different business requirements; each container instance is an independent running environment, which can avoid the situation where a single instance fails and affects the entire target service, improving the reliability and stability of the target service.
[0224] For the sake of easy understanding, the following will be combined with Figures 19 to 21 to introduce a method for expanding a distributed system applied in the game field. As Figure 19 shown, Figure 19It is an architecture diagram of a system for implementing an extended method of a distributed system provided by an embodiment of the present application. Among them, it includes a Job Engine module, a Component manager module, a Processor manager module, and a Processor Wrapper module. Specifically, the Job Engine module is mainly responsible for overall control and management, including reading and parsing the configurations of custom target components and custom target processes (the first target process and the second target process), communicating with the component manager module and the Processor manager module, issuing actual deployment tasks, and providing an interface to the Processor Wrapper module to query the deployment information of custom target components and custom target processes. The Component manager module is used to create, manage, and release component workloads through the container management cluster interface according to the custom component information configured by the user; at the same time, it is necessary to consider the port information required by the internal processes of the components and perform port configuration when creating the workload. The Processor manager module is used to execute custom target processes in specified components in a remote control manner according to the custom target process information configured by the user; at the same time, it is responsible for obtaining the running status of the custom processes in real time. The Processor Wrapper module is used to encapsulate user-oriented data communication, specifically including: providing a packaged data client SDK for developers to access the server in other processes in the custom process; providing an interface to customize the communication relationship between different processes; and interacting with the Job Engine module to obtain the underlying address information required for mutual access between processes.
[0225] As Figure 20 shown, before configuring custom target components and target processes, distributed deep learning tasks are composed of several built-in framework components, Figure 20 showing the components and processes included in a typical task. As Figure 21 shown, after introducing the custom topology, developers can add custom target components and custom target processes to the existing framework through configuration. Among them, the custom target component refers to a component in the framework other than the original Agent component, Learner component, Actor component, and Serving component. The custom process refers to a process other than the framework-built-in processes included in any component (including custom target components and original components).
[0226] The following will describe in detail the extended device of the distributed system in the present application. Please refer to Figure 22 . Figure 21FIG. 0 is a schematic diagram of an embodiment of the extension device 10 of the distributed system in the embodiment of the present application. The extension device 10 of the distributed system includes: a deployment information query module 110 for the target component, a deployment information parsing module 120 for the target component, and a target component configuration module 130; specifically:
[0227] The deployment information query module 110 for the target component is configured to call a deployment information interface to query the deployment information of the target component. The deployment information of the target component is used to represent the information for configuring the target component in the container management cluster. The deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task.
[0228] The deployment information parsing module 120 for the target component is configured to parse the deployment information of the target component to generate a target component configuration parsing file. The target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is parsed from the target component attribute information.
[0229] The target component configuration module 130 is configured to configure the target component in the container management cluster according to the workload information in the target component configuration parsing file.
[0230] For the extension device of the distributed system provided in the embodiment of the present application, developers can define the deployment information of components by means of parameter configuration without complex transformation of the existing framework. By calling the deployment information interface and parsing the deployment information, a target component configuration parsing file is automatically generated, and the target component is configured in the container management cluster according to the workload information, realizing the automatic deployment of the target component, reducing the time and errors of manual configuration, reducing the possibility of manual intervention and errors, and improving the efficiency of deployment and configuration; by using a configuration file to describe the requirements and configurations of the target component, the consistency and repeatability during deployment in different environments and clusters are ensured, and the reliability and stability of deployment are improved; by querying and parsing the deployment information of the target component, different extension requirements can be flexibly adapted; the automated deployment process can save time and effort, improving efficiency and productivity; resource allocation and configuration based on workload information can better manage computing resources, network resources, and storage resources to ensure that the target component obtains appropriate resource support in the container management cluster.
[0231] In an alternative embodiment of the extension device of the distributed system provided in the corresponding embodiment of the present application, the target component configuration module 130 is further configured to: Figure 22
[0232] Read the target component configuration parsing file to obtain M container configuration information, where the M container configuration information is used to characterize the configuration information corresponding to M containers configured in the target component, and M≥1;
[0233] Configure the M containers corresponding to the target component in the container management cluster according to the M container configuration information.
[0234] The extension device of the distributed system provided by the embodiment of the present application enables each container in the target component to be finely configured and deployed according to its configuration information by performing container configuration on the target component, ensuring that the target component can run normally in the container management cluster and meet the expected functional and performance requirements; by reading the target component configuration parsing file, obtaining the container configuration information, and automatically configuring and starting the containers in the container management cluster according to this information, the automatic deployment of the target component is realized, reducing the possibility of manual intervention and errors; using the configuration parsing file to describe the container requirements and configurations of the target component ensures consistency and repeatability when deploying in different environments and clusters, improving the reliability and stability of the deployment; the configuration parsing file can be customized and extended according to different requirements to adapt to different deployment scenarios and environments, providing flexibility and scalability; the automated container configuration and deployment process can save time and effort, improving efficiency and productivity; allocating and configuring resources according to the container configuration information can better manage computing resources, network resources, and storage resources, ensuring that the target component obtains appropriate resource support in the container management cluster; it helps to improve the deployment and management efficiency of the distributed system, reduce errors, and improve reliability.
[0235] In the Figure 22 corresponding optional embodiment of the extension device of the distributed system provided by the embodiment of the present application, please refer to Figure 23 , the extension device of the distributed system further includes: a first target process configuration module 140; specifically, the first target process configuration module 140 is used for:
[0236] Read the target component configuration parsing file to obtain N first target process configuration information, where the N first target process configuration information is used to characterize the configuration information corresponding to N first target processes configured in M containers, the configuration information corresponding to the N first target processes corresponds to N container identification information, the container identification information is used to indicate the container in which the first target process is configured, at least one first target process is configured in each container, and the N first target process configuration information carries the configured container identification information, and N≥M;
[0237] Configure the N first target processes in the M containers according to the N container identification information.
[0238] The expansion device of the distributed system provided by the embodiment of the present application, by configuring the first target process in the container of the target component, is a key step in the configuration process of the target component. It involves extracting the configuration information of the first target process from the configuration parsing file and allocating these processes to the container for configuration, so as to realize the flexible deployment and management of the processes of the target component in the container management cluster; by reading the configuration parsing file of the target component, obtaining the configuration information of the first target process, and associating it with the container identification information, it is possible to realize automated process deployment and container configuration, reduce the possibility of manual intervention and errors, and improve the efficiency and accuracy of deployment; at least one first target process is configured in each container, and the computing resources can be reasonably allocated and utilized according to the resource requirements and limitations of the container, avoiding resource waste and conflicts; the number of containers and the number of first target processes can be adjusted according to requirements to meet different load and performance requirements, realizing load balancing and horizontal expansion; each first target process runs in an independent container, providing process isolation and resource limitation, improving the security and reliability of the application; associating the first target process with the container identification information can facilitate the management, log collection and troubleshooting of the process, improving the maintainability and management efficiency of the application; the expansion device of the distributed system provided by the embodiment of the present application improves the deployment efficiency, resource management ability, scalability and flexibility, isolation and security, and maintainability, which helps to improve the performance and reliability of the distributed system.
[0239] In the Figure 23 corresponding optional embodiment of the expansion device of the distributed system provided by the embodiment of the present application, please refer to Figure 24 , the expansion device of the distributed system further includes: a second target process configuration module 150; specifically, the second target process configuration module 150 is used for:
[0240] Call the deployment information interface to query the deployment information of the second target process, where the deployment information of the second target process carries the specified component identification information, and the deployment information of the second target process is used to represent the information for configuring the second target process;
[0241] Parse the deployment information of the second target process to generate a second target process configuration parsing file;
[0242] Deploy the second target process in the specified component corresponding to the specified component identification information in the container management cluster according to the second target process configuration parsing file.
[0243] The extension device of the distributed system provided by the embodiment of the present application can realize flexible management and extension of the target process in the distributed system by calling the deployment information interface, parsing the deployment information, generating a configuration parsing file, and deploying the second target process in the specified component in the container management cluster, ensuring the correct configuration and deployment of the target process, and improving the reliability and maintainability of the system; by calling the deployment information interface and parsing the deployment information, generating a configuration parsing file, and deploying the second target process in the container management cluster according to this file, an automated deployment process is realized, reducing the possibility of manual intervention and errors; the automated deployment process can improve the efficiency and accuracy of deployment, reducing the deployment time and labor costs; by deploying the second target process through the container management cluster, the computing resources can be better managed and utilized, improving the utilization rate and efficiency of resources; the container-based deployment method provides better scalability and flexibility, and the deployment scale and configuration of the second target process can be quickly expanded or adjusted according to requirements; container technology provides process isolation and resource limitation, improving the security and reliability of the application program.
[0244] In the Figure 24 corresponding optional embodiment of the extension device of the distributed system provided by the embodiment of the present application, the second target process configuration module 150 is further configured to:
[0245] Read the second target process configuration parsing file to obtain K service port configuration information, where the K service port configuration information is used to represent the port configuration information corresponding to configuring K servers in the second target process, and K≥1;
[0246] Configure the K servers in the second target process according to the K service port configuration information.
[0247] The extension device of the distributed system provided by the embodiment of the present application can ensure that the server can correctly receive and receive requests from the client by reading the service port configuration information in the configuration parsing file and configuring the server in the second target process according to this information; by reading the service port configuration information from the configuration parsing file and configuring the server in the second target process according to this information, it can ensure that the server can correctly obtain and receive requests from the client, improving the accuracy and efficiency of deployment; saving the service port configuration information in the configuration parsing file facilitates subsequent maintenance and management. When the service port configuration needs to be modified, only the configuration parsing file needs to be updated, without modifying the code or redeploying; by configuring the service port, access to specific ports can be restricted, improving the security of the system.
[0248] In the Figure 24In an alternative embodiment of the extension device of the distributed system provided by the corresponding embodiment, the second target process configuration module 150 is further configured to:
[0249] Obtain remote server call information, where the remote server call information is used to represent information for connecting to a remote server;
[0250] Obtain the operation permission of the remote server according to the remote server call information;
[0251] After obtaining the operation permission of the remote server, control the remote server to parse the deployment information of the second target process to generate a second target process configuration parsing file.
[0252] The extension device of the distributed system provided by the embodiments of the present application, by obtaining the call information and operation permission of the remote server, parsing the deployment information of the second target process on the remote server, and generating a configuration parsing file, helps to achieve remote management and configuration of specific components in the distributed system; by obtaining the remote server call information, it can accurately connect to the remote server for subsequent operations; using the SSH protocol to establish a connection with the remote server provides an encrypted communication channel to protect the security of data transmission; once the operation permission of the remote server is obtained, various operations can be performed on the remote server, including parsing the deployment information of the second target process; by performing the parsing operation on the remote server, the computing resources of the remote server can be utilized to accelerate the parsing process and improve the overall deployment efficiency; allowing deployment on different remote servers provides greater flexibility and scalability to adapt to different deployment requirements and environments; the extension device of the distributed system provided by the embodiments of the present application provides a safe, efficient, and flexible way to connect to the remote server and perform parsing operations on the remote server to support the extension and management of the distributed system.
[0253] In the present application Figure 24 In an alternative embodiment of the extension device of the distributed system provided by the corresponding embodiment, please refer to Figure 25 , the extension device of the distributed system further includes: a client configuration module 160; specifically, the client configuration module 160 is configured to:
[0254] Read the target component configuration parsing file to obtain every P client configuration information, where the P client configuration information is used to represent the port configuration information of P clients configured in N first target processes, the P client configuration information corresponds to P first target process identification information, and P≥1;
[0255] Configure the P clients in the corresponding first target processes according to the P client configuration information and the P first target process identification information.
[0256] The extension device of the distributed system provided by the embodiment of the present application realizes the goal of configuring the client in the first target process. This helps to improve the flexibility and scalability of the system, enabling multiple clients to be configured and managed in different target processes. At the same time, by associating the configuration parsing file with the target process identification information, it can be ensured that the configuration of the client matches the specific requirements of the target process, improving the reliability and stability of the system; by reading the target component configuration parsing file, the port configuration information corresponding to each client can be obtained and associated with the first target process identification information; according to different client configuration information, each client can be configured in the corresponding first target process to meet the needs of different clients; it supports configuring multiple clients, enabling the system to adapt to different application scenarios and requirements; by combining the client configuration information with the first target process identification information, it can be ensured that the client can correctly communicate and interact with the target process; it provides the configuration management ability for multiple clients, realizes the personalized configuration of the client, improves the flexibility and scalability of the system, and ensures the correct communication and interaction between the client and the target process.
[0257] In the Figure 25 corresponding optional embodiment of the extension device of the distributed system provided by the embodiment of the present application, please refer to Figure 26 , the extension device of the distributed system further includes: a process communication configuration module 170; specifically, the process communication configuration module 170 is used for:
[0258] Call the deployment information interface to query the process communication relationship deployment information, where the process communication relationship deployment information includes the first target process identification information of the first target process and the second target process identification information of the second target process, and the process communication relationship deployment information is used to indicate the establishment of the communication relationship between the first target process and the second target process;
[0259] Parse the process communication relationship deployment information to generate a process communication relationship configuration parsing file;
[0260] Establish a communication connection between the first target process and the second target process according to the process communication relationship configuration parsing file.
[0261] The extension device of the distributed system provided by the embodiments of the present application can achieve the goals of deploying the communication relationships between processes in the distributed system and establishing communication connections, which helps improve the flexibility and scalability of the system, enabling different processes to effectively collaborate and exchange data. At the same time, through the use of parsing and configuration files, the communication relationships between processes can be conveniently managed and maintained, improving the reliability and stability of the system. By calling the deployment information interface, the communication relationship deployment information between the first target process and the second target process can be quickly obtained, including process identification information, etc. The obtained process communication relationship deployment information is parsed, and a process communication relationship configuration parsing file is generated. According to this file, a communication connection between the first target process and the second target process can be automatically established. By establishing the process communication connection, communication and collaboration between more processes can be supported, enabling the system to better adapt to different application scenarios and requirements. Using the configuration information in the process communication relationship configuration parsing file can ensure that the first target process and the second target process can correctly communicate and exchange data, improving the stability and reliability of the system. By querying and obtaining the process communication relationship deployment information in a simple and efficient manner, the automatic configuration of the process communication relationship is achieved, improving the scalability and flexibility of the system and ensuring the correctness and reliability of process communication.
[0262] In an optional embodiment of the extension device of the distributed system provided in the corresponding embodiment of the present application, the process communication configuration module 170 is further configured to: Figure 26 Obtain the first target process port information corresponding to the first target process according to the first target process identification information, and obtain the second target process port information corresponding to the second target process according to the second target process identification information;
[0263] Generate a first target process node according to the first target process port information, and generate a second target process node according to the second target process port information;
[0264] Establish a communication connection between the first target process node and the second target process node according to the process communication relationship configuration parsing file.
[0265]
[0266] The expansion device of the distributed system provided by the embodiment of the present application obtains port information according to process identification information, generates corresponding process nodes, and finally establishes communication connections between the process nodes by parsing the configuration file according to the communication relationship. This helps to achieve flexible communication and collaboration between processes in the distributed system. At the same time, by parsing the configuration file, the communication relationship between processes can be conveniently managed and maintained, improving the scalability and maintainability of the system; by obtaining the corresponding port information according to the process identification information, the ports used by each process can be conveniently understood for network communication and security management; according to the obtained port information, the corresponding process nodes are automatically generated, associating the processes with their port information, facilitating the subsequent establishment and management of communication connections; by establishing communication connections between the process nodes, communication and collaboration between more processes can be supported, enabling the system to better adapt to different application scenarios and requirements; using the configuration information in the process communication relationship configuration parsing file can ensure that processes can communicate and exchange data correctly, improving the stability and reliability of the system; by obtaining and managing the port information of the processes, the automatic generation of process nodes is realized, improving the scalability and flexibility of the system, and ensuring the correctness and reliability of process communication.
[0267] In the Figure 26 corresponding optional embodiment of the expansion device of the distributed system provided by the embodiment of the present application, the process communication configuration module 170 is further configured to:
[0268] Obtain the port information of the target client in the first target process according to the first target process identification information, and obtain the port information of the target server in the second target process according to the second target process identification information,
[0269] Generate a target client node according to the port information of the target client, and generate a target server node according to the port information of the target server;
[0270] Establish a communication connection between the target client node and the target server node according to the process communication relationship configuration parsing file.
[0271] The expansion device of the distributed system provided by the embodiment of the present application obtains port information according to the process identification information, generates corresponding nodes, and finally establishes communication connections between the nodes by parsing the configuration file according to the communication relationship, which helps to achieve flexible communication and collaboration between processes in the distributed system; at the same time, by parsing the configuration file, the communication relationship between processes can be conveniently managed and maintained, improving the scalability and maintainability of the system; by obtaining the port information of the target client and the target server according to the process identification information, the ports used by each process can be conveniently understood for network communication and security management; according to the obtained port information, target client nodes and target server nodes are automatically generated, associating the processes with their port information, facilitating the subsequent establishment and management of communication connections; by establishing communication connections between the target client nodes and the target server nodes, communication and collaboration between more processes can be supported, enabling the system to better adapt to different application scenarios and requirements; using the configuration information in the process communication relationship configuration parsing file can ensure that processes can communicate and exchange data correctly, improving the stability and reliability of the system; by providing a flexible and efficient way to obtain and manage the port information of processes, the automatic generation of process nodes is realized, improving the scalability and flexibility of the system, and ensuring the correctness and reliability of process communication.
[0272] In the Figure 22 corresponding optional embodiment of the expansion device of the distributed system provided by the embodiment of the present application, please refer to Figure 27 , the expansion device of the distributed system further includes: a container instance generation module 180; specifically, the container instance generation module 180 is used to:
[0273] Generate an image file of the target component according to the target component configuration parsing file;
[0274] Generate M container instances corresponding to M containers according to the image file, and implement the target service corresponding to the target component through the M container instances.
[0275] The expansion device of the distributed system provided by the embodiment of the present application generates an image file according to the configuration parsing file of the target component, and creates a container instance using the image file to implement the target service. This helps to achieve containerized deployment and microservice architecture, improving the flexibility, scalability, and maintainability of the system. At the same time, through the isolation and resource limitation of the container instance, the running environment of the target component can be better managed and controlled, improving the security and stability of the system; generating an image file according to the target component configuration parsing file and creating a container instance using the image file can quickly deploy the target service corresponding to the target component, reducing the time and workload of manual configuration and deployment; the image file contains the running environment and related code of the target component, which can ensure that each container instance has the same running environment and configuration, thus ensuring the consistency of the target service in different container instances; by generating multiple container instances, the scale of the target service can be flexibly expanded as needed to meet different business requirements; each container instance is an independent running environment, which can avoid the situation where a single instance fails and affects the entire target service, improving the reliability and stability of the target service.
[0276] Figure 28 FIG. 4 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 342 or data 344. Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.
[0277] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0278] The steps performed by the server in the above embodiments may be based on thisFigure 27 The server structure shown
[0279] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0280] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0281] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0282] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0283] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other media that can store program codes.
[0284] Above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An extension method for a distributed system, characterized in that Including: Call the deployment information interface to query the deployment information of the target component. The deployment information of the target component is used to characterize the information for configuring the target component in the container management cluster. The deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task. Parse the deployment information of the target component to generate a target component configuration parsing file. The target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is parsed from the target component attribute information. Configure the target component in the container management cluster according to the workload information in the target component configuration parsing file.
2. The method for expanding a distributed system according to claim 1, wherein The step of configuring the target component in the container management cluster according to the workload information in the target component configuration parsing file includes: Read the target component configuration parsing file to obtain M container configuration information. The M container configuration information is used to characterize the configuration information corresponding to configuring M containers in the target component, where M≥1. Configure the M containers corresponding to the target component in the container management cluster according to the M container configuration information.
3. The method for expanding a distributed system according to claim 2, wherein, After configuring the target component in the container management cluster according to the workload information in the target component configuration parsing file, it further includes: Read the target component configuration parsing file to obtain N first target process configuration information. The N first target process configuration information is used to characterize the configuration information corresponding to configuring N first target processes in the M containers. The configuration information corresponding to the N first target processes corresponds to N container identification information. The container identification information is used to indicate the container in which the first target process is configured. At least one of the first target processes is configured in each container, and the N first target process configuration information carries the configured container identification information, where N≥M. Configure the N first target processes in the M containers according to the N container identification information.
4. The method for expanding a distributed system according to claim 3, wherein, The method further includes: Call the deployment information interface to query the deployment information of the second target process. The deployment information of the second target process carries specified component identification information, and the deployment information of the second target process is used to characterize the information for configuring the second target process. Parse the deployment information of the second target process to generate a second target process configuration parsing file. Deploy the second target process in the specified component corresponding to the specified component identification information in the container management cluster according to the second target process configuration parsing file.
5. The method for expanding a distributed system according to claim 4, wherein The step of deploying the second target process in the specified component corresponding to the specified component identification information in the container management cluster according to the second target process configuration parsing file includes: Read the second target process configuration parsing file to obtain K service port configuration information, where the K service port configuration information is used to characterize the port configuration information corresponding to K servers configured in the second target process, and K≥1; Configure the K servers in the second target process according to the K service port configuration information.
6. The method for expanding a distributed system according to claim 4, wherein The parsing of the deployment information of the second target process to generate a second target process configuration parsing file includes: Obtain remote server call information, where the remote server call information is used to characterize information for connecting to a remote server; According to the remote server call information, obtain the operation permissions of the remote server; After obtaining the operation permissions of the remote server, control the remote server to parse the deployment information of the second target process to generate a second target process configuration parsing file.
7. The method for expanding a distributed system according to claim 3, wherein After configuring the N first target processes in the M containers according to the N container identification information, it further includes: Read the target component configuration parsing file to obtain every P client configuration information, where the P client configuration information is used to characterize the port configuration information corresponding to P clients configured in the N first target processes, and the P client configuration information corresponds to P first target process identification information, and P≥1; Configure the P clients in the corresponding first target processes according to the P client configuration information and the P first target process identification information.
8. The method for expanding a distributed system according to claim 7, wherein After configuring the P clients in the corresponding first target processes, it further includes: Call the deployment information interface to query the process communication relationship deployment information, where the process communication relationship deployment information includes the first target process identification information of the first target process and the second target process identification information of the second target process, and the process communication relationship deployment information is used to indicate establishing a communication relationship between the first target process and the second target process; Parse the process communication relationship deployment information to generate a process communication relationship configuration parsing file; Establish a communication connection between the first target process and the second target process according to the process communication relationship configuration parsing file.
9. The method for expanding a distributed system according to claim 8, wherein, The establishing of the communication connection between the first target process and the second target process according to the process communication relationship configuration parsing file includes: Obtain the first target process port information corresponding to the first target process according to the first target process identification information, and obtain the second target process port information corresponding to the second target process according to the second target process identification information; Generate a first target process node according to the first target process port information, and generate a second target process node according to the second target process port information; Establish a communication connection between the first target process node and the second target process node according to the process communication relationship configuration parsing file.
10. The method for expanding a distributed system according to claim 8, characterized in that, The establishing of the communication connection between the first target process and the second target process according to the process communication relationship configuration parsing file includes: Obtain the port information of the target client in the first target process according to the first target process identification information, and obtain the port information of the target server in the second target process according to the second target process identification information. Generate a target client node according to the port information of the target client, and generate a target server node according to the port information of the target server. Establish a communication connection between the target client node and the target server node according to the process communication relationship configuration parsing file.
11. The method for expanding a distributed system according to claim 2, wherein, After configuring the M containers corresponding to the target component in the container management cluster according to the M container configuration information, it includes: Generate an image file of the target component according to the target component configuration parsing file. Generate M container instances corresponding to the M containers according to the image file, and implement the target service corresponding to the target component through the M container instances.
12. An expansion device for a distributed system, characterized in that, It includes: A deployment information query module for the target component, which is used to call a deployment information interface to query the deployment information of the target component. The deployment information of the target component is used to characterize the information for configuring the target component in the container management cluster. The deployment information of the target component includes target component attribute information, and the target component is used to run a distributed reinforcement learning task. A deployment information parsing module for the target component, which is used to parse the deployment information of the target component to generate a target component configuration parsing file. The target component configuration parsing file includes the workload information of the target component, and the workload information of the target component is parsed from the target component attribute information. A target component configuration module, which is used to configure the target component in the container management cluster according to the workload information in the target component configuration parsing file.
13. A computer device, characterized in that, It includes: A memory, a transceiver, a processor, and a bus system. Among them, the memory is used to store programs. The processor is used to execute the programs in the memory, including executing the extended method of the distributed system according to any one of claims 1 to 11. The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
14. A computer-readable storage medium, comprising instructions, characterized in that, When it runs on a computer, it causes the computer to execute the extended method of the distributed system according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, This computer program is executed by the processor to perform the extended method of the distributed system according to any one of claims 1 to 11.