Micro-service capacity expansion and contraction and deployment method and device

By constructing a temporal directed acyclic graph and a spatiotemporal graph neural network model combined with reinforcement learning, the problems of cross-layer load propagation and unreasonable resource allocation in traditional methods are solved, and efficient resource management and fast QoS recovery are achieved in a microservice environment.

CN120849102APending Publication Date: 2025-10-28BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510925026.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional microservice auto-scaling methods fail to effectively consider cross-layer load propagation and the diversity of resource requirements, resulting in delayed scaling decisions and unreasonable resource allocation, which affects service quality and responsiveness. Furthermore, they fail to effectively address network fluctuations and resource availability differences between edge nodes.

Method used

By constructing a temporal directed acyclic graph, a spatiotemporal graph neural network model is used to extract cross-layer load propagation relationships. Combined with a reinforcement learning model, the scaling and deployment strategies of microservices are dynamically optimized, accurately capturing the cascading impact of upstream services on target services and the complex dependencies between microservices, thereby achieving dynamic resource adjustment.

Benefits of technology

It improves the collaborative efficiency of resource allocation and service quality assurance in the microservice environment, reduces service response time and resource waste, and improves QoS recovery speed and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849102A_ABST
    Figure CN120849102A_ABST
Patent Text Reader

Abstract

The invention provides a micro-service capacity expansion and contraction and deployment method and device. The method comprises the following steps: constructing a time sequence directed acyclic graph based on time sequence micro-service data; performing spatio-temporal feature extraction on the features of each node and each directed edge in the time sequence directed acyclic graph to obtain a spatio-temporal feature set sequence, and outputting an influence coefficient result for representing a cross-layer load propagation relationship based on the spatio-temporal feature set sequence; and determining a target capacity expansion and contraction and deployment strategy of each micro-service based on the enhanced state by adopting a reinforcement learning model, and executing copy capacity expansion and contraction and deployment operation of each micro-service according to each target capacity expansion and contraction and deployment strategy. According to the invention, the joint decision of micro-service automatic capacity expansion and contraction and deployment can be dynamically realized to reduce service response time and accelerate QoS recovery as optimization objectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microservices technology, and in particular to a method and apparatus for scaling up and deploying microservices. Background Art

[0002] With the rapid development of edge computing and microservice architecture, microservice deployment and automatic scaling technologies in edge environments have become a research hotspot. Edge computing significantly reduces data transmission latency and improves the real-time response capability of the system by deploying computing and storage resources close to the data source. It is particularly suitable for application scenarios with high requirements for timeliness and data processing speed, such as smart manufacturing, autonomous driving, and smart cities. Microservice architecture, on the other hand, breaks down monolithic applications into a series of independent small service units, each focusing on a specific functional module. It has good module decoupling and flexible expansion capabilities, thereby enhancing the maintainability, scalability, and fault tolerance of the system. The combination of edge computing and microservice architecture effectively promotes the flexible deployment and rapid response of edge applications, improving the user experience.

[0003] However, with the continuous expansion of edge applications and the heterogeneous and dynamic nature of edge environment resources, ensuring Quality of Service (QoS) has become a critical challenge. In scenarios with high load or sudden traffic surges, failure to adjust resource configuration in a timely manner can easily lead to increased response time or even service unavailability. Therefore, automatic scaling, as an important means of dynamically adjusting the number of microservice replicas, is widely used in QoS assurance. At the same time, a reasonable replica placement strategy is also crucial for ensuring stable service operation. In edge environments, due to frequent fluctuations in network conditions between nodes and significant differences in communication latency between different deployment locations, improper replica deployment locations can lead to QoS degradation even with timely scaling. Therefore, the rationality of the deployment strategy is also essential for QoS assurance in edge environments.

[0004] Traditional microservice auto-scaling methods typically rely on load monitoring of individual microservices, triggering scaling decisions through preset thresholds or simple predictive models. While these methods can address load changes to some extent, they neglect the impact of cross-layer load propagation. In reality, changes in the load of upstream microservices can propagate through the call chain, ultimately affecting the resource requirements and performance of downstream services. Ignoring this cross-layer load propagation can easily lead to delayed scaling decisions, resulting in network latency, service responsiveness, and QoS degradation. Secondly, traditional methods often assume that the same request frequency (QPS) corresponds to the same resource consumption, failing to consider the differences in resource requirements caused by the heterogeneity of request sources and data volumes. Requests from different upstream microservices may have drastically different resource consumption characteristics due to differences in their business logic. Without differentiation, this can easily lead to uneven resource allocation, impacting service performance. Furthermore, these methods ignore the dynamic nature of request data volume. Even with stable QPS, a surge in the amount of data carried by requests can cause a sharp increase in CPU and memory resource pressure. Therefore, scaling decisions based solely on request frequency are often inaccurate.

[0005] This approach, which ignores cross-layer load propagation and the diversity of resource requirements, can easily lead to delayed scaling decisions and unreasonable resource allocation, resulting in QoS degradation and resource overload or waste. Furthermore, traditional deployment methods do not consider network fluctuations and resource availability differences between edge nodes, nor the impact of network topology on communication latency between edge nodes, leading to unreasonable deployment and poor QoS. Summary of the Invention

[0006] In view of this, embodiments of the present invention provide a microservice scaling and deployment method and apparatus to eliminate or improve one or more defects existing in the prior art.

[0007] One aspect of the present invention provides a method for scaling up and deploying microservices, the method comprising the following steps: constructing a time-series directed acyclic graph based on time-series microservice data, wherein the time-series microservice data includes microservice data for multiple time periods obtained through periodic collection, the microservice data for each time period includes performance index data of each microservice in the microservice cluster and call relationships and call performance index data between each microservice, the time-series directed acyclic graph includes multiple directed acyclic graphs corresponding to multiple time periods, each directed acyclic graph includes multiple nodes representing their respective microservices and multiple directed edges representing call relationships between their respective microservices, the performance index data of each microservice serves as the feature of each node, and the call performance index data between each microservice serves as the feature of each directed edge; Spatiotemporal features are extracted from the features of each node and each directed edge in the temporal directed acyclic graph to obtain a spatiotemporal feature set sequence. Based on the spatiotemporal feature set sequence, an influence coefficient result representing the cross-layer load propagation relationship is output. The spatiotemporal feature set sequence includes multiple spatiotemporal feature sets corresponding to multiple directed acyclic graphs. A reinforcement learning model is used to determine the target scaling and deployment strategy for each microservice based on the enhanced state, and to execute the replica scaling and deployment operation of each microservice according to the target scaling and deployment strategy. The enhanced state includes the influence coefficient result, the performance index data of each microservice and each server running each microservice in the current time period, and the call performance index data between each microservice.

[0008] In some embodiments of the present invention, a spatiotemporal graph neural network model is used to extract spatiotemporal features from the features of each node and each directed edge in the temporal directed acyclic graph to obtain a spatiotemporal feature set sequence, and the influence coefficient result representing cross-layer load propagation is output based on the spatiotemporal feature set sequence.

[0009] In some embodiments of the present invention, the spatiotemporal graph neural network model includes a multi-layer spatiotemporal module stack structure and an output layer connected in sequence. The multi-layer spatiotemporal module stack structure includes a multi-layer spatiotemporal module connected in sequence, used to extract spatiotemporal features from the features of each node and each directed edge in the temporal directed acyclic graph, so as to output a spatiotemporal feature set sequence. The output layer is used to output influence coefficient results based on the spatiotemporal feature set sequence. Each spatiotemporal module includes a temporal convolutional layer for temporal dimension feature extraction, a graph attention layer for spatial dimension feature extraction, and an activation function layer connected in sequence.

[0010] In some embodiments of the present invention, the temporal convolutional layer includes a one-dimensional causal convolution and a gated linear unit as an activation function. Temporal features are extracted from the features of each node and each directed edge in the temporal directed acyclic graph, or from the output of the previous spatiotemporal module, using the following formula, to output the temporal features of each node and each directed edge, or the temporal features of the output of the previous spatiotemporal module: T(v)=P☉σ(Q) Where T(v) represents the time feature, v represents the past ρ time step feature of the input, P represents the basic feature extracted by one-dimensional causal convolution, ⊙ represents the Hada code product, σ represents the sigmoid function, and Q represents the gate signal; For each node in the output of the temporal convolutional layer, the graph attention layer comprises a multi-layer graph attention network, and spatial dimension features are extracted using the following formula: in, Let represent the attention coefficients, i.e., spatial features, of node i and its neighbor node j in the l-th layer graph attention network; ReLU represents the activation function, a represents the attention weight vector, and W represents the learnable linear transformation matrix. and These represent the features of nodes i and j after processing by the (l-1)th layer of the graph attention network. and All are outputs from the temporal convolutional layer; e ij The temporal features of the edge between nodes i and j are represented by the output of the temporal convolutional layer; || represents the feature concatenation operation, and N(i) represents the set of neighboring nodes of node i. For each node in the output of the graph attention layer, the activation function layer performs a weighted sum of the temporal and spatial features of each node's neighboring nodes using the following formula: in, σ' represents the spatiotemporal characteristics of node i; σ′ represents the activation function, used to introduce nonlinearity.

[0011] In some embodiments of the present invention, the reinforcement learning model includes a policy network and a value network; employing the reinforcement learning model to determine the target scaling and deployment strategies for each microservice based on the reinforcement state includes: Actions are generated through a policy network, and these actions include multiple scaling and deployment policies corresponding to each microservice. The reward for the action is determined through a value network, and the reward includes a reward based on the response latency of the microservice cluster, which includes the communication latency between the microservices in the microservice cluster and the processing latency of the individual microservices. Based on the enhanced state, the scaling and deployment strategies that maximize the reward are used as the target scaling and deployment strategies for each microservice.

[0012] In some embodiments of the present invention, the policy network and the value network are iteratively updated by optimizing a loss function based on dynamically weighted time difference error, wherein the dynamically weighted time difference error is represented by the following formula: Where, δ t Indicates time difference error; ω t This represents the dynamic weight, obtained based on the influence coefficient results; r t This represents the reward for the current time slot, where γ represents the reward discount factor. Represents a value function. This indicates the enhancement status of the current time slot.

[0013] In some embodiments of the present invention, the performance metrics data of each microservice includes the number of replicas of each microservice, as well as the resource allocation data and resource utilization of each replica. The call performance metrics data between each microservice includes the call frequency, data transmission volume, and communication latency between each microservice.

[0014] Another aspect of the present invention provides a microservice scaling and deployment apparatus, the apparatus comprising: a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the apparatus implementing the steps of the aforementioned method when the computer instructions are executed by the processor.

[0015] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned method.

[0016] Another aspect of the present invention provides a computer program product including computer instructions that, when executed by a processor, implement the steps of the aforementioned method.

[0017] The microservice scaling and deployment method and apparatus of this invention can accurately capture the cascading impact of upstream services on target services and the complex dependencies between microservices in each service call chain by modeling cross-layer load propagation influence coefficients through STGNN in the joint decision-making process of microservice replica scaling and deployment. This improves the collaborative efficiency of resource allocation and QoS assurance in a microservice environment. Furthermore, reinforcement learning is used to collaboratively optimize the number of microservice replicas and the deployment location of each replica on edge nodes, dynamically realizing the joint decision-making of automatic scaling and deployment of microservices with the optimization goal of reducing service response time and accelerating QoS recovery. This solves the problems of resource waste and response delay conflicts caused by independent decision-making in scaling and deployment in traditional methods.

[0018] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.

[0019] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0020] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.

[0021] Figure 1 This is a schematic diagram of a microservice scaling and deployment architecture in one embodiment of the present invention; Figure 2 This is a flowchart illustrating a microservice scaling and deployment method in one embodiment of the present invention. Figure 3 This is a schematic diagram of the network structure of a spatiotemporal graph neural network model in one embodiment of the present invention; Figure 4 This is a schematic diagram of the network structure and principle of a reinforcement learning model in one embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.

[0023] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.

[0024] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0025] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0026] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0027] This invention proposes a method and apparatus for scaling and deploying microservices, as well as a corresponding architecture. The method is implemented using this architecture and is primarily aimed at server clusters (microservice clusters) running microservice applications deployed in edge environments. However, it can also be applied to microservice clusters in other environments. A microservice is an independently deployed application unit, typically composed of multiple service instances or replicas. Figure 1 This is a schematic diagram of a microservice scaling and deployment architecture in one embodiment of the present invention, as shown below. Figure 1As shown, the architecture includes a monitoring and data collection module, an STGNN (Spatio-Temporal Graph Neural Network) feature extraction module, a joint decision-making module, and an execution module. The monitoring and data collection module, STGNN feature extraction module, and joint decision-making module jointly make scaling and deployment decisions for each microservice by integrating information from all microservices in the microservice cluster. The execution module dynamically adjusts the number and location of replicas of each microservice deployed on edge nodes through the resource management and scheduling module interface of the edge microservice cluster, achieving real-time matching of replica distribution with load changes. In the specific implementation, the monitoring and data collection module and the execution module can be adjusted according to the actual edge microservice platform management interface, as long as they can complete data collection and decision execution.

[0028] Figure 2 This is a flowchart illustrating a microservice scaling and deployment method in one embodiment of the present invention. Figure 2 As shown, the method includes the following steps: Step S210: Construct a time-series directed acyclic graph (DAG) based on time-series microservice data. The time-series microservice data includes microservice data for multiple time periods obtained through periodic collection. The microservice data for each time period includes performance metrics data of each microservice in the microservice cluster, as well as the call relationships and call performance metrics data between the microservices. The time-series DAG includes multiple DAGs corresponding to multiple time periods. Each DAG includes multiple nodes representing their respective microservices and multiple directed edges representing the call relationships between their respective microservices. The performance metrics data of each microservice serve as the features of each node, and the call performance metrics data between the microservices serve as the features of each directed edge.

[0029] Specifically, this method uses a monitoring and collection module to periodically or periodically collect microservice data from a microservice cluster, obtaining microservice data across multiple time periods or time slots, thus forming time-series microservice data. The periodically collected microservice data can be stored in a time-series database for subsequent processing. Then, using an STGNN feature extraction module, directed acyclic graphs (DAGs) are constructed based on the performance metrics of each microservice and the call relationships and performance metrics between them within the time-series microservice data, resulting in a time-series DAG. This allows for the capture and modeling of call relationships or dependencies between microservices using graph neural networks.

[0030] The performance metrics for each microservice include the number of replicas (instances) of each microservice, as well as resource allocation and utilization for each replica (instance). These data serve as characteristics of each node in the directed acyclic graph (DAG). The performance metrics for inter-microservice communication include call frequency, data transmission volume, and communication latency (network latency). These data serve as characteristics of directed edges pointing to nodes in the DAG. Resource allocation data includes CPU size and memory quota, while resource utilization includes CPU utilization and memory usage.

[0031] Step S220: Spatiotemporal feature extraction is performed on the features of each node and each directed edge in the temporal directed acyclic graph to obtain a spatiotemporal feature set sequence, and the influence coefficient result representing the cross-layer load propagation relationship is output based on the spatiotemporal feature set sequence, wherein the spatiotemporal feature set sequence includes multiple spatiotemporal feature sets corresponding to multiple directed acyclic graphs.

[0032] Specifically, this step can be implemented using the spatiotemporal graph neural network model in the STGNN feature extraction module. By learning the impact of edge load on node performance in the temporal directed acyclic graph constructed by the spatiotemporal graph neural network, the influence coefficient matrix of the representation influence coefficient results between microservices is output. This matrix is ​​used to quantify the degree of influence of load propagation (cross-layer load propagation) between microservices and calling edges (directed edges), which is used to guide the generation of subsequent elastic scaling and replica placement or deployment strategies, thereby achieving dynamic resource adjustment.

[0033] In some embodiments, the spatiotemporal graph neural network model includes a stacked structure of multiple spatiotemporal modules (ST-Blocks) and an output layer connected in sequence. The stacked structure of multiple spatiotemporal modules includes multiple spatiotemporal modules connected in sequence for extracting spatiotemporal features from the features of each node and each directed edge in the temporal directed acyclic graph to output a sequence of spatiotemporal feature sets. The output layer is used to output influence coefficient results based on the sequence of spatiotemporal feature sets. Each spatiotemporal module includes a temporal convolutional layer for extracting temporal dimension features, a graph attention layer for extracting spatial dimension features, and an activation function layer connected in sequence. Figure 3 This is a schematic diagram of the network structure of a spatiotemporal graph neural network model in one embodiment of the present invention, as shown below. Figure 3 The multi-layered spatiotemporal module stacking structure consists of two layers of spatiotemporal modules connected in sequence.

[0034] Specifically, a temporal convolutional layer is used to extract the dynamic load features of call edges in the temporal directed acyclic graph (DAG) as they change over time. Graph Attention Layers (GAT) are then used to deeply mine and capture the spatial dependencies between these call edges, enabling the capture of hidden relationships even between distant microservices. In other words, it considers not only the impact of load changes on upstream services but also on the load changes of distant microservices, as well as the impact of the call chain initiating the request and the size of the request data, thus more accurately capturing the dynamic impact of load changes on microservices. Through this multi-layered spatiotemporal module stacking structure, the model can learn more complex temporal-spatial dependencies. Through the synergistic effect of temporal and spatial feature extraction in this multi-layered spatiotemporal module stacking structure, the output layer generates an influence coefficient matrix, characterizing the complex relationship of cross-layer load propagation between microservices. This matrix quantifies the degree of influence of each call edge on the scaling decisions or strategies of each microservice.

[0035] In some embodiments, the temporal convolutional layer includes a one-dimensional causal convolution and a gated linear unit (GLU) as the activation function. Temporal features of each node and each directed edge in the temporal directed acyclic graph, or the output of the previous spatiotemporal module, can be extracted using the following formula, outputting the temporal features of each node and each directed edge, or the temporal features of the output of the previous spatiotemporal module: T(v)=P☉σ(Q) Where T(v) represents the temporal feature, v represents the past ρ time steps of the input, P represents the basic feature extracted by one-dimensional causal convolution, ⊙ represents the Hada code product, σ represents the sigmoid function, and Q represents the gating signal. One-dimensional causal convolution ensures that the features at the current moment depend only on past inputs, and combined with gated linear units, key temporal features can be preserved. The features of all nodes and all edges (node ​​feature set sequences and edge feature set sequences) of each graph in the input temporal directed acyclic graph are processed by temporal convolutional layers, or the output of the previous spatiotemporal module is processed by temporal convolutional layers, thus fully extracting the temporal dimension features.

[0036] For each node in the output of the temporal convolutional layer, the graph attention layer comprises a multi-layer graph attention network, and spatial dimension features can be extracted using the following formula: in, Let represent the attention coefficients, i.e., spatial features, of node i and its neighbor node j in the l-th layer graph attention network; ReLU represents the activation function, a represents the attention weight vector, and W represents the learnable linear transformation matrix. and These represent the features of nodes i and j after processing by the (l-1)th layer of the graph attention network. and All are outputs from the temporal convolutional layer; e ij The temporal features of the edge between nodes i and j are represented by the output of the temporal convolutional layer; || represents the feature concatenation operation, and N(i) represents the set of neighboring nodes of node i. The graph attention network used in this method differs from traditional graph convolutional networks in that it introduces edge temporal features (temporal features such as call frequency, communication latency, and data transmission volume) during node information aggregation, and measures the importance of neighboring nodes to a node by calculating attention coefficients. By introducing edge temporal features, the expressive power of call relationships between remote nodes (microservices) is enhanced, and differentiated weights are assigned to different edges.

[0037] For each node in the output of the graph attention layer, the activation function layer can perform a weighted sum of the temporal and spatial features of each node's neighboring nodes using the following formula: in, The spatiotemporal feature representation of node i is represented by σ′, which is composed of the spatiotemporal feature representations of all nodes in each graph, forming a spatiotemporal feature set sequence. σ′ represents the activation function, used to introduce nonlinearity. Through this mechanism, the graph attention layer can adaptively adjust the information propagation path according to the current context, significantly improving the model's ability to identify key load propagation paths in the microservice call chain.

[0038] Step S230: Using a reinforcement learning model, determine the target scaling and deployment strategy for each microservice based on the enhanced state, and execute the replica scaling and deployment operation for each microservice according to the target scaling and deployment strategy. The enhanced state includes the influence coefficient result, the performance index data of each microservice and each server running each microservice in the current time period, and the call performance index data between each microservice.

[0039] Specifically, this step can be implemented through a joint decision-making module and an execution module. The joint decision-making module determines the optimal target scaling and deployment strategy, while the execution module executes the corresponding scaling and deployment operations based on the determined target scaling and deployment strategy. This dynamically adjusts the number of replicas for each microservice and the deployment location of each replica, ensuring real-time matching of the strategy with changes in cluster load. For the automatic scaling and deployment problem of microservices in a microservice cluster, this problem can be modeled as a Markov Decision Process (MDP). The aim is to minimize response latency and maximize QoS recovery (reducing the time required for QoS recovery) by optimizing the number of replicas of each microservice and the deployment location of each replica (i.e., on which servers they are deployed) in resource-constrained and frequently changing network conditions, while simultaneously improving the overall stability and elastic recovery capability of the cluster, thereby improving the cluster service quality. A reinforcement learning algorithm is used to solve the above problem, obtaining the optimal scaling and deployment strategy for each microservice to achieve the above optimization objectives, thus enabling resource scheduling optimization in edge computing environments. This optimal scaling and deployment strategy includes the optimal number of replicas and the optimal deployment location of each replica.

[0040] Enhanced state can be used Let represent, where s t This indicates the system environment status of the microservice cluster in the current time slot, including performance metrics data for each microservice and each server running each microservice, as well as performance metrics data for calls between microservices. Y t This represents the influence coefficient results, which reflect the cross-layer load propagation or transmission relationship of various service call chains in a microservice cluster. As a core input to the joint decision-making module and reinforcement learning model, the influence coefficient results, by incorporating historical data into reinforcement learning, can model the cross-layer propagation relationship of load in the call chain. This is more conducive to the joint optimization of scaling and deployment strategies, improving the accuracy and adaptability of scaling and deployment decisions.

[0041] In some embodiments, the reinforcement learning model includes a policy network and a value network. Using a reinforcement learning model to determine the target scaling and deployment strategies for each microservice based on the reinforcement state includes the following steps: Actions are generated through a policy network, and these actions include multiple scaling and deployment policies corresponding to each microservice. The reward for the action is determined through a value network, and the reward includes a reward based on the response latency of the microservice cluster, which includes the communication latency between the microservices in the microservice cluster and the processing latency of the individual microservices. Based on the enhanced state, the scaling and deployment strategies that maximize the reward are used as the target scaling and deployment strategies for each microservice.

[0042] Specifically, Figure 4 This is a schematic diagram of the network structure and principle of a reinforcement learning model in one embodiment of the present invention, as shown below. Figure 4 As shown, the Actor-Critic reinforcement learning algorithm is used as the decision framework, which includes a policy (Actor) network and a value (Critic) network. The Actor network generates scaling and deployment policies as actions based on the augmentation state, while the Critic network evaluates the reward that the policy can obtain in the current state to guide the iterative optimization of the policy.

[0043] In some embodiments, the policy network and the value network are iteratively updated by optimizing a loss function based on dynamically weighted time difference error. By introducing dynamically weighted time difference error, the learning of the reinforcement learning model can be further accelerated. The dynamically weighted time difference error (TD error) is expressed by the following formula: Where, δ t Indicates time difference error; ω t The dynamic weight, calculated based on the influence coefficient, reflects the current load impact of each upstream microservice on the target microservice; r t The reward for the current time slot is represented by γ, which can be obtained by weighting and summing the response times of each service call chain in the microservice cluster according to the importance of the call chain; γ represents the reward discount factor, which determines the importance attached to future rewards. The value function represents the expected cumulative reward for a given state under the current policy. This represents the enhanced state of the current time slot. This mechanism allows reinforcement learning models to pay more attention to key state transitions that significantly impact the system performance of the microservice cluster during training.

[0044] Furthermore, the policy network can be updated using the following loss function: Among them, L π This represents the loss of the policy network. Represents the current batch of sample sets The number of samples in The policy function represents the state under enhancement. Take action a t The probability of this action can be determined by p. m,n The resulting matrix representation, p m,n This represents the number of replicas or instances of microservice m deployed on server n, where θ represents the strategy parameter. This represents the derivative of the policy parameters, used to calculate the direction and magnitude of parameter updates.

[0045] The value network can be trained by minimizing the following loss function of the weighted mean squared error: Among them, L V This represents the loss of the value network. Through the above mechanism, this method effectively integrates cross-layer load propagation information and system environment status, improves the model training convergence speed and the performance of the generated final policy, and enables more efficient and timely microservice scaling and deployment decisions in the dynamic environment of edge computing.

[0046] In summary, the microservice scaling and deployment method proposed in this invention, during the joint decision-making process for scaling and deploying microservice replicas, utilizes STGNN to model the cross-layer load propagation impact coefficient. This accurately captures the cascading impact of upstream services on the target service and the complex dependencies between microservices in each service call chain, improving the collaborative efficiency of resource allocation and QoS assurance in edge environments. Furthermore, reinforcement learning is used to collaboratively optimize the number of microservice replicas and the deployment location of each replica on edge nodes, dynamically achieving joint decision-making for automatic scaling and deployment of microservices with the optimization goal of reducing service response time and accelerating QoS recovery. This solves the resource waste and response latency conflicts caused by independent decision-making in scaling and deployment in traditional methods.

[0047] In addition, the various modules in the microservice scaling and deployment architecture can be adjusted and expanded according to different scenarios and optimization goals. The STGNN feature extraction module can adapt to the spatiotemporal feature inputs of different edge scenarios, the joint decision module can be implemented using a variety of algorithms, such as static thresholding, deep learning and optimization algorithms, etc., and the influence coefficient matrix can adjust the weight factors according to business needs, such as call chain priority and data volume weight, so as to meet the service quality assurance needs of diverse microservice applications and dynamically improve the service quality in diverse application scenarios.

[0048] Corresponding to the above method, the present invention also provides a microservice scaling and deployment apparatus, which includes a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the apparatus implements the steps of the aforementioned method.

[0049] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned method. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.

[0050] This invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the steps of the aforementioned method.

[0051] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.

[0052] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0053] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for scaling up and deploying microservices, characterized in that, The method includes: A time-series directed acyclic graph (DAG) is constructed based on time-series microservice data. The time-series microservice data includes microservice data for multiple time periods obtained through periodic collection. The microservice data for each time period includes performance metrics data of each microservice in the microservice cluster, as well as the call relationships and call performance metrics data between the microservices. The time-series DAG includes multiple DAGs corresponding to multiple time periods. Each DAG includes multiple nodes representing their respective microservices and multiple directed edges representing the call relationships between their respective microservices. The performance metrics data of each microservice serve as the features of each node, and the call performance metrics data between the microservices serve as the features of each directed edge. Spatiotemporal features are extracted from the features of each node and each directed edge in the temporal directed acyclic graph to obtain a spatiotemporal feature set sequence. Based on the spatiotemporal feature set sequence, an influence coefficient result representing the cross-layer load propagation relationship is output. The spatiotemporal feature set sequence includes multiple spatiotemporal feature sets corresponding to multiple directed acyclic graphs. A reinforcement learning model is used to determine the target scaling and deployment strategy for each microservice based on the enhanced state, and to execute the replica scaling and deployment operation of each microservice according to the target scaling and deployment strategy. The enhanced state includes the influence coefficient result, the performance index data of each microservice and each server running each microservice in the current time period, and the call performance index data between each microservice.

2. The method according to claim 1, characterized in that, A spatiotemporal graph neural network model is used to extract spatiotemporal features from each node and each directed edge in the temporal directed acyclic graph to obtain a spatiotemporal feature set sequence. Based on the spatiotemporal feature set sequence, the influence coefficient result representing cross-layer load propagation is output.

3. The method according to claim 2, characterized in that, The spatiotemporal graph neural network model includes a multi-layer spatiotemporal module stack structure and an output layer connected in sequence. The multi-layer spatiotemporal module stack structure includes multiple spatiotemporal modules connected in sequence, used to extract spatiotemporal features from the features of each node and each directed edge in the temporal directed acyclic graph, so as to output a spatiotemporal feature set sequence. The output layer is used to output the influence coefficient result based on the spatiotemporal feature set sequence. Each spatiotemporal module includes a temporal convolutional layer for temporal dimension feature extraction, a graph attention layer for spatial dimension feature extraction, and an activation function layer connected in sequence.

4. The method according to claim 3, characterized in that, The temporal convolutional layer includes a one-dimensional causal convolution and a gated linear unit as the activation function. Temporal features are extracted from the features of each node and each directed edge in the temporal directed acyclic graph, or from the output of the previous spatiotemporal module, using the following formula: [Formula omitted for brevity] T(v)=P⊙σ(Q) Where T(v) represents the time feature, v represents the past ρ time step feature of the input, P represents the basic feature extracted by one-dimensional causal convolution, ⊙ represents the Hada code product, σ represents the sigmoid function, and Q represents the gate signal; For each node in the output of the temporal convolutional layer, the graph attention layer comprises a multi-layer graph attention network, and spatial dimension features are extracted using the following formula: in, Let represent the attention coefficients, i.e., spatial features, of node i and its neighbor node j in the l-th layer graph attention network; ReLU represents the activation function, a represents the attention weight vector, and W represents the learnable linear transformation matrix. and These represent the features of nodes i and j after processing by the (l-1)th layer of the graph attention network. and All are outputs from the temporal convolutional layer; e ij The temporal features of the edge between nodes i and j are represented by the output of the temporal convolutional layer; || represents the feature concatenation operation, and N(i) represents the set of neighboring nodes of node i. For each node in the output of the graph attention layer, the activation function layer performs a weighted sum of the temporal and spatial features of each node's neighboring nodes using the following formula: in, σ' represents the spatiotemporal characteristics of node i; σ′ represents the activation function, used to introduce nonlinearity.

5. The method according to claim 1, characterized in that, The reinforcement learning model includes a policy network and a value network; A reinforcement learning model is used to determine the target scaling and deployment strategies for each microservice based on the reinforcement state, including: Actions are generated through a policy network, and these actions include multiple scaling and deployment policies corresponding to each microservice. The reward for the action is determined through a value network, and the reward includes a reward based on the response latency of the microservice cluster, which includes the communication latency between the microservices in the microservice cluster and the processing latency of the individual microservices. Based on the enhanced state, the scaling and deployment strategies that maximize the reward are used as the target scaling and deployment strategies for each microservice.

6. The method according to claim 5, characterized in that, The policy network and the value network are iteratively updated by optimizing a loss function based on dynamically weighted time difference error, wherein the dynamically weighted time difference error is represented by the following formula: Where, δ t Indicates time difference error; ω t This represents the dynamic weight, obtained based on the influence coefficient results; r t This represents the reward for the current time slot, where γ represents the reward discount factor. Represents a value function. This indicates the enhancement status of the current time slot.

7. The method according to any one of claims 1 to 6, characterized in that, The performance metrics data for each microservice include the number of replicas of each microservice, as well as the resource allocation data and resource utilization of each replica. The performance metrics data for calls between microservices include the call frequency, data transmission volume, and communication latency between microservices.

8. A microservice scaling and deployment apparatus, comprising a processor, a memory, and computer instructions stored in the memory, characterized in that, The processor is configured to execute the computer instructions, and when the computer instructions are executed, the device implements the steps of the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1 to 7.