Grayscale changing method and device, storage medium and electronic equipment
By obtaining the status data of the target service and using the grayscale change optimization model to predict the change action, the problem of cumbersome operations under the k8s Pod deployment method is solved, and convenient grayscale changes and response speed are achieved.
Patent Information
- Application Number
- CN202510348580.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
In the deployment mode using k8s Pod, the modification of the service description requires pod reconstruction and manual operation, resulting in cumbersome operation and lagging response speed, and lack of convenient grayscale change solutions.
By obtaining the status data of the target service, calling the grayscale change optimization model for change action prediction, determining the target prediction probability value, and performing grayscale changes based on the probability value, using neural network models such as Double Deep Q-Network for action prediction and optimization.
A convenient grayscale change process is realized, the response speed and operation efficiency are improved, manual intervention is reduced, and the system flexibility and response speed are improved.
Smart Images

Figure CN120301766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a grayscale change method, device, storage medium, and electronic device. Background Art
[0002] Currently, in the deployment mode of using native Pods of k8s (Kubernetes, an open-source container orchestration platform), where a Pod is the smallest scheduling unit in Kubernetes, consisting of one or more containers, and services are usually deployed and managed in the form of Pods in Kubernetes, the service description adopts an end-state configuration method. Modifications to numerous service descriptions all require Pod reconstruction and service restart. In this regard, related technologies usually implement processes such as changes through manual means, resulting in cumbersome operations and lagging response speeds. Based on this, there is currently no good solution for how to conveniently perform grayscale changes to improve the response speed. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a grayscale change method, device, storage medium, and electronic device to solve problems such as cumbersome operations in related technologies; that is, embodiments of the present invention can conveniently determine a target change action through a target grayscale change optimization model, and thus conveniently perform grayscale changes through the target change action, which can effectively improve the response speed.
[0004] According to one aspect of the embodiments of the present invention, there is provided a grayscale change method, the method including:
[0005] Obtain target state data of a target service;
[0006] Call a target grayscale change optimization model to perform change action prediction on the target state data, and obtain target prediction probability values of each change action among multiple change actions;
[0007] Based on the target prediction probability values of each change action, determine a target change action from the multiple change actions, and perform grayscale change on the target service according to the target change action.
[0008] According to another aspect of the embodiments of the present invention, there is provided a grayscale change device, the device including:
[0009] An obtaining unit, configured to obtain target state data of a target service;
[0010] A processing unit, configured to call a target grayscale change optimization model to perform change action prediction on the target state data, and obtain target prediction probability values of each change action among multiple change actions;
[0011] The processing unit is further configured to determine a target change action from the multiple change actions based on the target prediction probability values of the respective change actions, and perform a gray-scale change on the target service according to the target change action.
[0012] According to another aspect of the embodiments of the present invention, there is provided an electronic device, which includes a processor and a memory storing a program. Wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the method mentioned above.
[0013] According to another aspect of the embodiments of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method mentioned above.
[0014] After obtaining the target status data of the target service, the embodiments of the present invention can call the target gray-scale change optimization model to predict change actions for the target status data, and obtain the target prediction probability values of the respective change actions among the multiple change actions; further, based on the target prediction probability values of the respective change actions, a target change action can be determined from the multiple change actions, and a gray-scale change is performed on the target service according to the target change action. It can be seen that the embodiments of the present invention can conveniently determine the target change action through the target gray-scale change optimization model, and thus conveniently perform gray-scale change through the target change action, which can effectively improve the response speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In the following description of exemplary embodiments with reference to the drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:
[0016] Figure 1 A flowchart showing a method for gray-scale change according to an exemplary embodiment of the present invention is shown;
[0017] Figure 2 A schematic diagram showing service description data according to an exemplary embodiment of the present invention is shown;
[0018] Figure 3 A flowchart showing another method for gray-scale change according to an exemplary embodiment of the present invention is shown;
[0019] Figure 4 A flowchart showing yet another method for gray-scale change according to an exemplary embodiment of the present invention is shown;
[0020] Figure 5 A schematic block diagram showing a gray-scale change device according to an exemplary embodiment of the present invention is shown;
[0021] Figure 6A structural block diagram of an exemplary electronic device that can be used to implement an embodiment of the present invention is shown. Detailed implementation manners
[0022] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0023] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0024] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.
[0025] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] It should be noted that the execution subject of the grayscale change method provided in the embodiments of the present invention can be one or more electronic devices, and the embodiments of the present invention do not limit this; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and at least one terminal and at least one server are included in the multiple electronic devices, the grayscale change method provided in the embodiments of the present invention can be jointly executed by the terminal and the server. Optionally, in other embodiments, the execution subject of the grayscale change method can also be an agent in the electronic device, that is, the electronic device can execute the grayscale change method through the agent; or, the electronic device can also be called an agent, etc.; the present invention does not limit this.
[0028] Correspondingly, the terminals mentioned here can include but are not limited to: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, etc. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. Optionally, the electronic device that executes the grayscale change method can be any Node node in the target Kubernetes cluster, or a device located outside the target Kubernetes cluster, and the embodiments of the present invention do not limit this; optionally, the target Kubernetes cluster can be any Kubernetes cluster, that is to say, the embodiments of the present invention can realize the grayscale change of any service in any Kubernetes cluster through the proposed grayscale change method.
[0029] Based on the above description, the embodiments of the present invention propose a grayscale change method, which can be executed by the above-mentioned electronic device (terminal or server); or, this grayscale change method can be jointly executed by the terminal and the server. For the convenience of description, in the following, it is assumed that the electronic device executes this grayscale change method as an example for explanation; as Figure 1 shown, this grayscale change method may include the following steps S101 - S103:
[0030] S101, obtain the target status data of the target service.
[0031] Optionally, the target service can be any service, and the embodiments of the present invention do not limit this. Optionally, the electronic device can comprehensively monitor the service indication data of each service in the target Kubernetes cluster through an efficient real-time observation mechanism (list&watch). When the service indication data of any service is detected to change, any service can be used as the target service; or, the target service can be any arbitrarily specified service, etc.; the embodiments of the present invention do not limit this. Optionally, the number of target services can be one or more, and the embodiments of the present invention do not limit this; for the convenience of description, a single target service will be used as an example for illustration hereinafter. Optionally, the number of target Kubernetes clusters can be one or more, and the embodiments of the present invention do not limit this; that is to say, the target Kubernetes cluster can be any Kubernetes cluster, etc.
[0032] Optionally, a service may have a service description data, and the service description data of a service may include, but is not limited to: the basic information, deployment information, scheduling information, and monitoring information of the corresponding service, etc., such as Figure 2As shown; embodiments of the present invention do not limit this. Optionally, the basic information of a service may include but is not limited to: the service type of the corresponding service (such as java (an object-oriented programming language for writing cross-platform application software) type, Go (an open-source programming language) type, etc.), the technology stack (such as java, go, etc.), the network partition to which it belongs, the main traffic sources (such as search engine traffic, social media traffic, event traffic, etc.), the service importance level (i.e., the service level), and the corresponding RD (Research and Development), OP (Operation and Maintenance), QA (Quality Assurance) responsible persons, etc.; embodiments of the present invention do not limit this. Optionally, the deployment information of a service may include but is not limited to the basic image type of the corresponding service (such as private image, public image, etc.), the application image code library and its version, important configuration files, startup commands, process & port information (such as the number of processes, port numbers, etc.), etc.; embodiments of the present invention do not limit this. Optionally, the scheduling information of a service may include but is not limited to: the numerical sizes of CPU (Central Processing Unit), memory, and disk applied for by the corresponding service, the resource ratio of the sidecar (an architectural design that splits service functions into independent processes) container (i.e., the configuration of each container, such as the CPU ratio, etc.), the network bandwidth used (i.e., the network bandwidth of all pods corresponding to the corresponding service), the configuration of affinity and anti-affinity (a scheduling strategy, such as preventing pods of different services with anti-affinity from being configured on the same Node, and supporting pods of different services with affinity to be configured on the same Node), taint and tolerance configuration (another scheduling strategy), label selector, etc.; embodiments of the present invention do not limit this. Optionally, the monitoring information of a service may include but is not limited to: the log collection path & period of the corresponding service, the process port configuration (such as process monitoring and port monitoring after the service starts), the business metric values of the corresponding service for at least one business core metric, and the specified monitoring & alarm configuration of the corresponding service, etc.; embodiments of the present invention do not limit this.
[0033] Optionally, the service indication data of a service may include, but is not limited to, at least one of the following: the base image type of the corresponding service, the application image code repository, the code version, the target configuration file (such as any important configuration file), the startup command, the port number, the scheduling information, etc.; the embodiments of the present invention do not limit this. In other words, the service indication data of a service may include the metric information of the corresponding service under each service detection change metric in at least one service detection change metric; optionally, at least one service detection change metric may be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this; exemplarily, at least one service detection change metric may include, but is not limited to, at least one of the following: the base image type metric, the application image code repository metric, the code version metric, at least one configuration file metric, at least one scheduling metric (such as the memory metric, the CPU metric, etc.), etc. Based on this, the service indication data of a service may include at least one data in the service description data of the corresponding service.
[0034] In the embodiments of the present invention, the obtaining methods of the target state data of the target service may include, but are not limited to, the following several types:
[0035] The first obtaining method: The electronic device detects changes in the service indication data of each service in at least one service, and when detecting changes in the service indication data of the target service, obtains the current state data of the target service, and uses the current state data of the target service as the target state data, and the target service may be any service in at least one service. Based on this, the electronic device can obtain the target state data of the target service in the actual production operation environment in real time; that is to say, the electronic device can comprehensively monitor the service indication data of the service through an efficient real-time observation mechanism (list&watch). Once any change in the service indication data of the target service is detected, the target state data of the target service can be obtained to trigger an online change process. Based on this, in the deployment environment of Kubernetes native Pod, the design of the service indication data follows the final state configuration principle; this means that once the service indication data is modified or new code is uploaded and passes the Code Review (CR) process, the online change process can be automatically triggered, and this mechanism greatly simplifies the cumbersome steps in the traditional deployment process, enabling the service to respond more quickly and flexibly to changes in business requirements.
[0036] The second obtaining method: The storage space of the electronic device may store the changed service indication data and the target state data of the target service. In this case, the electronic device can obtain the target state data from its own storage space, etc.
[0037] Optionally, the status data of a service may include, but is not limited to, at least one of the following: the monitoring values of the corresponding service under each status monitoring metric in at least one status monitoring metric, the service metric values of each status service metric in at least one status service metric corresponding to the corresponding service, the start status indication values of each process port corresponding to the corresponding service (such as 1 for started and 0 for not started, etc.), the log collection and parsing information of the corresponding service (such as the number of exception logs and / or whether there is exception log indication information, etc.), the importance level (i.e., service level) of the corresponding service, and the online traffic of the corresponding service, etc.; the embodiments of the present invention do not limit this. Optionally, at least one status monitoring metric may include, but is not limited to, at least one of the following: CPU usage rate, memory usage rate, disk usage rate, etc.; the embodiments of the present invention do not limit this. Optionally, at least one status service metric corresponding to a service may be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this; exemplarily, a status service metric may be a business transaction volume metric for indicating the size of the transaction volume of a business, etc. Optionally, a status data may be any data in the state space, and the state space may contain sufficient information to characterize the current state of the system. Optionally, the status data of a service may be obtained in real time through means such as monitoring and log collection.
[0038] Optionally, a status data may be data directly obtained from a monitoring system, etc., or data after preprocessing and / or feature extraction, etc.; the embodiments of the present invention do not limit this. It should be noted that the embodiments of the present invention do not limit the specific implementation manners of preprocessing and feature extraction.
[0039] S102, call the target gray change optimization model to predict change actions for the target status data, and obtain the target prediction probability values of each change action among multiple change actions.
[0040] Optionally, a gray-scale change optimization model may be a neural network model. The embodiments of the present invention do not limit the specific model structure of the gray-scale change optimization model; for example, a gray-scale change optimization model may be a Double Deep Q-Network (DDQN, a neural network model), etc. Optionally, a gray-scale change optimization model may include an input layer, a hidden layer, an output layer, etc., and the embodiments of the present invention do not limit this. Among them, the input layer can receive feature vectors (i.e., state data) in the state space. These feature vectors can be data directly obtained from the monitoring system or data after preprocessing or feature extraction; the hidden layer can include multiple fully connected layers and / or convolutional layers for extracting deep features of the state data. The number of hidden layers and the number of neurons in each layer can be adjusted according to the complexity of the problem and computing resources; the output layer can include the number of neurons corresponding to the action space (i.e., the number of neurons corresponding to multiple change actions). Each neuron outputs the predicted probability value of the corresponding change action. For example, for a discrete action space, the output layer can use activation functions such as softmax to output the probability distribution (i.e., the predicted probability value) of each change action, etc.
[0041] Optionally, multiple change actions can be set according to experience or according to actual needs, and the embodiments of the present invention do not limit this. Optionally, a change action can be to perform gray-scale release according to at least one gray-scale release (i.e., gray-scale change) ratio (at this time, it is a continuous gray-scale change action, such as performing gray-scale release according to a gray-scale release ratio, or a change stage can correspond to a gray-scale release ratio, etc.), or to adjust the gray-scale release ratio for gray-scale release (such as increasing or decreasing the gray-scale release ratio according to a preset ratio, or maintaining the gray-scale release ratio, etc.), or to pause gray-scale release, or to roll back to the old version, etc.; the embodiments of the present invention do not limit this. For example, a change action may further include but is not limited to the gray-scale change duration, the change range (such as a single computer room or multiple computer rooms, etc.), etc.; based on this, detailed configurations such as the gray-scale change duration, the change range, and the volume ratio (i.e., the gray-scale release ratio) can be used as executable actions. Optionally, the preset ratio can be set according to experience or according to actual needs, and the embodiments of the present invention do not limit this. Optionally, the way of continuous gray-scale release can be divided into dimensions and ranges such as single-computer-room volume increase and multi-computer-room volume increase, and the embodiments of the present invention do not limit this. Among them, a change action can be an action in the action space, and the action space can be used to define all possible actions, that is, all allowed actions.
[0042] Optionally, the electronic device may input the target status data into the target gray-scale change optimization model (i.e., input the input layer of the target gray-scale change optimization model) to output the target prediction probability values of each change action through the target gray-scale change optimization model; alternatively, the target status data may also be preprocessed and / or feature-extracted to input the obtained data into the input layer of the target gray-scale change optimization model, so as to realize the prediction of change actions for the target status data, etc.; the embodiments of the present invention do not limit this.
[0043] S103. Based on the target prediction probability values of each change action, determine the target change action from multiple change actions, and perform gray-scale change on the target service according to the target change action.
[0044] Optionally, when determining the target change action from multiple change actions based on the target prediction probability values of each change action, the electronic device may determine the change action with the largest target prediction probability value from multiple change actions, and use the change action with the largest target prediction probability value as the target change action.
[0045] In the embodiments of the present invention, after obtaining the target status data of the target service, the target gray-scale change optimization model may be called to perform change action prediction on the target status data, and the target prediction probability values of each change action among multiple change actions may be obtained; further, based on the target prediction probability values of each change action, the target change action may be determined from multiple change actions, and gray-scale change may be performed on the target service according to the target change action. It can be seen that the embodiments of the present invention can conveniently determine the target change action through the target gray-scale change optimization model, and thus conveniently perform gray-scale change through the target change action, which can effectively improve the response speed.
[0046] Based on the above description, the embodiments of the present invention also propose another gray-scale change method. Correspondingly, this gray-scale change method may be executed by the above-mentioned electronic device (terminal or server); or, this gray-scale change method may be jointly executed by the terminal and the server, etc. For the convenience of description, in the following, it is taken as an example that the electronic device executes this gray-scale change method for illustration; please refer to Figure 3 This gray-scale change method may include the following steps S301-S307:
[0047] S301. Obtain a training data set, and a training data includes a training status data.
[0048] Among them, the training data set may include at least one training data. Optionally, a training data may further include, but is not limited to, at least one of the following: a training action label, a training service indication data (which is a changed service indication data for guiding the gray-scale change of the corresponding service, so that the corresponding service after the gray-scale change conforms to the training service indication data), the next state data of the corresponding training state data (i.e., the state data after the gray-scale change, and a training state data may be a state data before the gray-scale change), and the label reward value corresponding to the corresponding training data, etc.; the embodiments of the present invention do not limit this.
[0049] In the embodiments of the present invention, the acquisition methods of the training data set may include, but are not limited to, the following several:
[0050] The first acquisition method: There are multiple training data stored in the own storage space of the electronic device. In this case, the multiple training data can be added to the training data set to obtain the training data set; alternatively, the experience replay mechanism can be used to store and reuse past experiences, that is, at least one training data can be selected from the multiple training data in the way of experience replay, and the selected at least one training data can be added to the training data set to obtain the training data set. At this time, the time correlation between the data can be broken, and the stability of the training can be improved, etc.
[0051] Optionally, the above multiple training data can be obtained through data collection; for example, in a simulated production environment, the monitoring data, service indication data, etc. can be made consistent with the online data, so as to collect the training state data, training action labels, etc. according to experience or actual needs to collect multiple training data, etc. It should be noted that the embodiments of the present invention do not limit the collection process of the multiple training data.
[0052] The second acquisition method: The electronic device can obtain the training download link of the training data set and use the data set downloaded based on the training download link as the training data set to obtain the training data set, etc.
[0053] S302, call the initial gray-scale change optimization model to predict the change actions for the training state data in each training data included in the training data set, and obtain the predicted probability values of each change action under each training state data.
[0054] Among them, the change action prediction can also be called the change work probability prediction.
[0055] In an embodiment of the present invention, the electronic device may input each training status data (i.e., the training status data in each training data) into the initial grayscale change optimization model to output the predicted probability values of each change action under each training status data (which may also be referred to as the predicted probability values of each change action under each training data).
[0056] S303. Determine the model loss value of the initial grayscale change optimization model based on the predicted probability values of each change action under each training status data.
[0057] In one implementation, a training data may further include a training action label. Optionally, a training action label may be an action identifier (an action identifier can be used to indicate a change action, such as an action name or an action number, etc.), or a training action label probability, etc. The embodiments of the present invention do not limit this; among them, a training action label probability may include the label probability values of each change action, that is, the training action label probability corresponding to a training data may include the label probability values of each change action under the corresponding training data (which may also be referred to as the label probability values of each change action under the corresponding training status data). Based on this, the training action label in a training data can be used to indicate the training label action corresponding to the corresponding training data; when a training action label is an action identifier, the training label action corresponding to a training data can be the action indicated by the action identifier in the corresponding training data; when a training action label is a training action label probability, the training label action corresponding to a training data can be the action corresponding to the maximum label probability value in the training action label probability in the corresponding training data. Correspondingly, the training action label in a training data can also be used to indicate the training action label probability corresponding to the corresponding training data; it should be understood that when the training action label in a training data is an action identifier, the training action label probability corresponding to the corresponding training data can be determined through the training action label in the corresponding training data, so that the training action label in a training data can be used to indicate the training action label probability corresponding to the corresponding training data, such as determining that the label probability value of the change action indicated by the training action label in the corresponding training data is 1, and the label probability value of each change action other than the change action indicated by the training action label in the corresponding training data is 0, and so on.
[0058] Based on this, when determining the model loss value of the initial gray-scale change optimization model based on the prediction probability values of each change action under each training state data, the electronic device can respectively determine the training action labels corresponding to each training data based on the training action labels in each training data, and use the prediction probability values of each change action under each training state data and the training action label probability values corresponding to each training data to calculate the model loss value of the initial gray-scale change optimization model. For example, the cross-entropy loss function is used to calculate the model loss value of the initial gray-scale change optimization model by using the prediction probability values of each change action under each training state data and the training action label probability values corresponding to each training data.
[0059] In another implementation, when determining the model loss value of the initial gray-scale change optimization model based on the prediction probability values of each change action under each training state data, the electronic device can respectively determine the training prediction actions under each training state data based on the prediction probability values of each change action under each training state data. For example, the change action with the largest prediction probability value in any training state data is used as the training prediction action in that training state data. The training prediction action in a training state data can also be referred to as the training prediction action corresponding to the training data where the corresponding training state data is located, that is, the training prediction action corresponding to the corresponding training data. Correspondingly, in the simulated production environment, gray-scale changes can be performed respectively according to the training prediction actions under each training state data to determine the prediction reward values corresponding to each training data; thereby determining the label reward values corresponding to each training data, and determining the model loss value of the initial gray-scale change optimization model based on the prediction reward values and label reward values corresponding to each training data, and so on.
[0060] Optionally, a training data may further include training service indication data. Then, when performing gray-scale change according to the training prediction action in any training status data, the gray-scale change may be performed according to the training service indication data in the training data where any training status data is located and the training prediction action in any training status data. That is to say, the training prediction action in any training status data may be adopted, and the gray-scale change may be performed according to the training service indication data in the training data where any training status data is located. That is, the service corresponding to any training status data (the service corresponding to a training status data may be the service corresponding to the training data where the corresponding training status data is located, and a training data may correspond to a service) is gray-scale changed according to the training service indication data in the training data where any training status data is located. Furthermore, the prediction reward value corresponding to the training data where any training status data is located may be determined to implement the determination of the prediction reward values corresponding to each training data. Based on this, after performing the gray-scale change according to the training prediction action in any training status data, the prediction reward value corresponding to the training data where any training status data is located (hereinafter all referred to as any training data) may be determined.
[0061] Optionally, the prediction reward value corresponding to a training data may be determined by a reward function, that is, a reward value may be determined by a reward function. Optionally, the reward function may be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this. Optionally, the reward function may be defined according to system continuous performance, change timeliness, etc. Exemplarily, the shorter the change time, the higher the reward. The impact range of change failure (such as impact on traffic, impact time, etc.) may be used as a negative value (that is, the larger the impact range of change failure, the higher the negative value). The higher the reduction amount of the index value of each reward monitoring index, the higher the reward (for example, if the CPU usage rate decreases, the reward is a positive value; if the CPU usage rate increases, the reward is a negative value, etc.). The embodiments of the present invention do not limit this. In other words, the reward function may be determined based on reward items such as change time, impact range of change failure, and at least one reward monitoring index. The embodiments of the present invention do not limit the specific reward items in the reward function, that is, the specific design of the reward function may be adjusted and optimized according to actual business requirements. Optionally, at least one reward monitoring index may be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this.
[0062] Optionally, a training data may further include a label reward value. In this case, the label reward value corresponding to any training data can be determined from any training data to achieve the determination of the label reward value corresponding to any training data. Alternatively, a training data may further include a training action label. In this case, in a simulated production environment, gray-scale changes can be made according to the training label actions indicated by the training action labels in any training data to determine the label reward value corresponding to any training data. At this time, the specific implementation of making gray-scale changes according to the training label actions indicated by the training action labels in any training data to determine the label reward value corresponding to any training data may be the same as the specific implementation of making gray-scale changes according to the training prediction actions in any training state data to determine the prediction reward value corresponding to any training data. The embodiments of the present invention will not be described in detail herein, and so on.
[0063] Optionally, when determining the model loss value of the initial gray-scale change optimization model based on the prediction reward values and label reward values corresponding to each training data, the mean square error (MSE) loss function can be used to determine the model loss value of the initial gray-scale change optimization model based on the prediction reward values and label reward values corresponding to each training data. Alternatively, the squared error loss function can also be used to determine the model loss value of the initial gray-scale change optimization model based on the prediction reward values and label reward values corresponding to each training data, and so on. The embodiments of the present invention do not limit this.
[0064] Optionally, in other embodiments, a gray-scale change optimization model may further output a prediction reward value. For example, the initial gray-scale change optimization model may further output the prediction reward value corresponding to any training data, and the label reward value corresponding to any training data can be determined, so as to determine the model loss value of the initial gray-scale change optimization model based on the prediction reward values and label reward values corresponding to each training data. Optionally, when the label reward value corresponding to any training data is determined based on real-time monitoring in the environment, the embodiments of the present invention can determine the model loss value (i.e., LOSS) through the error between the reward feedback by the environment (i.e., the label reward value) and the reward output by the model (i.e., the prediction reward value), so as to optimize the model.
[0065] In summary, the electronic device can call the initial gray-scale change optimization model to predict the change actions for the training status data in each training data included in the training data set respectively, so as to determine the model loss value of the initial gray-scale change optimization model, that is, the predicted probability value of each change action under each training status data and / or the predicted reward value corresponding to each training data can be obtained. Thus, based on the predicted probability value of each change action under each training status data and / or the predicted reward value corresponding to each training data, the model loss value of the initial gray-scale change optimization model is determined. For example, the model loss value of the initial gray-scale change optimization model is determined based on the predicted probability value of each change action under each training status data, or the model loss value of the initial gray-scale change optimization model is determined based on the predicted reward value corresponding to each training data (in this case, the label reward value corresponding to each training data can be further determined to determine the model loss value), and so on. Based on this, the specific determination method of the model loss value in the embodiments of the present invention is not limited.
[0066] S304. Optimize the model parameters in the initial gray-scale change optimization model in the direction of reducing the model loss value to obtain the initial gray-scale change optimization model after model optimization, and determine the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization.
[0067] Optionally, when determining the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization, the training data set can be continuously obtained, such as obtaining the training data set from multiple training data stored in the experience replay buffer, etc., so as to continue training the initial gray-scale change optimization model after model optimization through the training data set until the model convergence condition is reached (such as the model loss value is less than the preset model loss threshold, or the number of training rounds reaches the preset training round threshold, etc.), and the gray-scale change optimization model when the model convergence condition is reached is used as the target gray-scale change optimization model to realize determining the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization. Optionally, both the preset model loss threshold and the preset training round threshold can be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this.
[0068] It should be noted that in the process of optimizing the model parameters, the embodiments of the present invention do not limit the selection of the optimizer and the setting of the learning rate.
[0069] S305. Obtain the target status data of the target service.
[0070] S306. Call the target gray-scale change optimization model to predict the change actions for the target status data to obtain the target predicted probability value of each change action among multiple change actions.
[0071] S307. Determine a target change action from multiple change actions based on the target prediction probability values of the respective change actions, and perform a gray-scale change on the target service according to the target change action.
[0072] Optionally, the electronic device may also determine the initial monitoring values of the respective change preparation monitoring metrics among the multiple change preparation monitoring metrics before the change of the target service; and based on the initial monitoring values of the respective change preparation monitoring metrics (the initial monitoring value of a change preparation monitoring metric is the initial monitoring value of the corresponding change preparation monitoring metric before the change of the target service), determine the disconnection monitoring values of the respective change preparation monitoring metrics when the online traffic is cut off in a single computer room. Further, based on the disconnection monitoring values of the respective change preparation monitoring metrics (the disconnection monitoring value of a change preparation monitoring metric is the disconnection monitoring value of the corresponding change preparation monitoring metric when the online traffic is cut off in a single computer room), it can be determined whether the target service meets the cut-off flow conditions at the computer room level, and when it is determined that the target service meets the cut-off flow conditions at the computer room level, trigger the execution of the above-mentioned gray-scale change of the target service according to the target change action. Optionally, this stage can be referred to as the pre-change preparation stage.
[0073] Optionally, the multiple change preparation monitoring metrics can be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this; exemplarily, the multiple change preparation monitoring metrics may include, but are not limited to: CPU usage rate, memory usage rate, and disk usage rate, etc. Optionally, when determining the disconnection monitoring values of the respective change preparation monitoring metrics when the online traffic is cut off in a single computer room based on the initial monitoring values of the respective change preparation monitoring metrics, for any one of the multiple change preparation monitoring metrics, the electronic device may determine the number of computer rooms corresponding to the target service, and based on the number of computer rooms corresponding to the target service and the initial monitoring value of any one change preparation monitoring metric, determine the disconnection monitoring value of any one change preparation monitoring metric when the online traffic is cut off in a single computer room; exemplarily, assuming that any one change preparation monitoring metric is the CPU usage rate, then the disconnection monitoring value of any one change preparation monitoring metric when the online traffic is cut off in a single computer room may be the initial monitoring value of any one change preparation monitoring metric × the number of computer rooms corresponding to the target service / (the number of computer rooms corresponding to the target service - 1), etc.
[0074] Optionally, when judging whether the target service meets the room-level cutting condition based on the cut-off monitoring value of each change preparation monitoring indicator, if the cut-off monitoring value of each change preparation monitoring indicator is less than the cut-off monitoring threshold of the corresponding change preparation monitoring indicator, it can be determined that the target service meets the room-level cutting condition; if there is a change preparation monitoring indicator among multiple change preparation monitoring indicators whose cut-off monitoring value is greater than or equal to the cut-off monitoring threshold of the corresponding change preparation monitoring indicator, it can be determined that the target service does not meet the room-level cutting condition. Optionally, the cut-off monitoring threshold of a change preparation monitoring indicator can be set according to experience or according to actual needs, and the embodiments of the present invention are not limited to this.
[0075] Based on this, in the pre-change preparation stage, the embodiment of the present invention can determine whether the time is ripe to promote the change or initiate loss-stopping measures; at the same time, the service traffic water level data must meet the computer room-level flow cutting conditions to ensure that when the change operation is performed in a single computer room, the online traffic can be completely cut off, while other computer rooms can smoothly take over and maintain normal online services.
[0076] Optionally, when performing grayscale changes to the target service according to the target change action, when the target change action is to suspend grayscale release, the grayscale change can be suspended; when the target change action is to roll back to the old version, the target service can be rolled back. Correspondingly, when the target change action is a continuing grayscale change action, the electronic device can determine multiple change stages, and make step-by-step changes according to the multiple change stages and the target change action, so as to achieve grayscale changes to the target service according to the target change action. For example, the target Kubernetes cluster can be controlled according to the target change action to perform grayscale changes to the target service, and one change stage corresponds to one change level. Optionally, multiple change stages include a pre-release instance stage (one instance can be a pod), a low-traffic stage, a single computer room mass release stage, and a multi-computer room mass release stage. Figure 4 As shown; wherein, the pre-launch instance stage refers to the change process of the pre-launch instance, the small traffic stage refers to the change and volume expansion process of the online instances with a preset number, the single computer room volume expansion stage refers to the change and volume expansion process of the online instances of a single computer room, and the multi-computer room volume expansion stage refers to the change and volume expansion process of the online instances of multiple computer rooms. Optionally, the pre-launch instances and the preset number can be set according to experience or according to actual needs, and the embodiments of the present invention are not limited to this. Among them, the pre-launch instance refers to an instance that is not connected to online traffic, that is, a non-online instance; the online instance can refer to an instance that supports online traffic. Optionally, the change levels of the pre-launch instance stage, the small traffic stage, the single computer room volume expansion stage, and the multi-computer room volume expansion stage can be increased successively, that is, the change scope of each change stage can be expanded successively.
[0077] Optionally, during the change process, the electronic device can continuously determine the current monitoring data, and based on the change monitoring strategy and the current monitoring data in the current change stage, determine whether the current change process is abnormal, so as to perform a stop-loss operation according to the change stop-loss strategy in the current change stage when detecting that the change process is abnormal. Optionally, the change monitoring strategy and the change stop-loss strategy in a change stage can both be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this; for example, the change monitoring strategy in a change stage can be used to indicate the monitoring value thresholds of each change monitoring index in at least one change monitoring index (such as determining that the change process is abnormal if the monitoring value exceeds the corresponding monitoring value threshold, and at this time, it can also be determined whether the current change process is abnormal based on historical monitoring data), and / or used to indicate the monitoring value change trend thresholds of each change monitoring index within a period of time (such as determining that the change process is abnormal if the monitoring value change trend exceeds the corresponding monitoring value change trend threshold), and so on; for another example, the change stop-loss strategy in a change stage can be used to indicate stopping the change process, or can also be used to indicate continuing to observe the monitoring data within a specified duration, or can also be used to indicate rolling back to the old version, and so on. Optionally, at least one change monitoring index, the monitoring value threshold of a change monitoring index, and the monitoring value change trend threshold of a change monitoring index within a period of time can all be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this. Optionally, the change monitoring strategies in different change stages can be the same or different, and the embodiments of the present invention do not limit this; correspondingly, the change stop-loss strategies in different change stages can be the same or different, and the embodiments of the present invention do not limit this; exemplarily, the change stop-loss strategy in the change stage with a higher change level can be more stringent to avoid causing greater impacts and thus improve system stability. Optionally, the current monitoring data can refer to the current monitoring data in the current change stage; optionally, the current monitoring data in a change stage can refer to the current monitoring data of the instances within the change scope in the corresponding change stage (such as the pre-release instances in the pre-release instance stage, the preset number of instances in the small traffic stage, all the instances of the target service in the current changed computer room in the single computer room ramp-up stage or the multi-computer room ramp-up stage, etc.), or can also refer to the current monitoring data of the instances within the specified change scope, and so on, and the embodiments of the present invention do not limit this; exemplarily, the pre-release instance stage can refer to the current monitoring data of the pre-release instances, and the online instance stage (i.e., the subsequent several stages) other than the pre-release instance stage can refer to the current monitoring data of the online instances (i.e., all the instances providing online traffic), or can also refer to the current monitoring data of the currently changed online instances, and so on.
[0078] Based on this, the electronic device can first promote the deployment of the pre-release instance, that is, first execute the pre-release instance phase to implement changes to the pre-release instance and closely monitor its monitoring data. Only when the monitoring data of the pre-release instance shows normal (that is, it is determined that the current change process is not abnormal according to the change monitoring policy under the preset instance phase), or there is no significant difference from the monitoring data of the online instance (such as the difference between the monitoring values under each change monitoring index is less than the corresponding monitoring difference threshold, etc.), will it progress to the small-traffic phase. Correspondingly, the change scope will be gradually expanded, that is, gradually progress to the next change phase. Correspondingly, as the online service is gradually changed (that is, the gradual change of the small-traffic phase, the single-data-center traffic increase phase, and the multi-data-center traffic increase phase), environmental data such as process log monitoring and business monitoring will continuously be fed back to the electronic device to continuously determine the current monitoring data. These monitoring data provide valuable real-time information for the electronic device to help it evaluate the effect and impact of the change, and based on the change trend of these monitoring data, it can again carefully decide whether to continue triggering traffic increase or need to pause the change, observe for a period of time and then make a decision. In other words, during the change process, the electronic device always closely monitors the monitoring data fed back by the environment. Once any abnormality and / or alarm event is detected, it can immediately take actions, such as stopping the current change process, quickly shielding the traffic of the problem instance to prevent the problem from spreading, and setting the concurrency according to the urgency and impact range of the alarm event, and triggering the rollback process to restore the service to the state before the change, which can effectively improve the emergency response speed, that is, avoid the lag of the emergency response speed, and so on.
[0079] In the embodiment of the present invention, the single-data-center traffic increase phase (which can also be called the single-data-center implementation phase) is the last line of defense for realizing traffic switching and loss prevention operations. At this time, the changed single data center and other data centers (that is, other unchanged data centers) will together bear the traffic pressure of the entire system, which may have a certain impact on upstream and downstream services. Therefore, once an abnormality occurs in the changed single data center, it is necessary to immediately execute the traffic switching plan, quickly transfer the traffic to other data centers, and immediately start the single-data-center rollback operation to minimize the loss to the greatest extent. Correspondingly, the multi-data-center traffic increase phase (which can also be called the multi-data-center full traffic increase phase) is the final phase of the change implementation. After the single data center (that is, the first changed single data center) has been stably running for more than a specified duration (which can be any duration, such as 40 minutes, etc.) and there are no obvious abnormalities in each monitoring index and business performance (that is, no abnormality is detected under the change monitoring policy in the current change phase), the change operation of this phase can be executed. In this phase, single-data-center changes can be carried out one by one for the remaining unchanged data centers. If a rollback is required in this phase, it will inevitably cause a short interruption or loss of the online service. Therefore, it is necessary to carefully formulate and execute the rollback plan according to the real-time traffic water level data to ensure the continuity and stability of the service.
[0080] It should be understood that in the actual production operation environment, more than 80% of system failures can be traced back to change operations. In view of this, whenever the service description (i.e., service indication data) is adjusted, a series of subsequent change operations must have the ability of instant failure alarm and intelligent loss stop. When formulating change decisions and loss stop measures, it is necessary to comprehensively consider whether the service process meets the conditions for multi-data center traffic switching and whether a detailed loss stop response plan has been pre-developed. These pieces of information are not only crucial but also serve as the core components of the service description, and are observed and analyzed in real time by electronic devices. In addition, all change scenarios triggered by electronic devices (also known as intelligent systems) must be equipped with a complete and verified rollback mechanism to ensure that in case of any abnormal situation, it can quickly and accurately return to the stable state before the change. According to the importance and urgency of the service, the observation time required at each stage will be adjusted accordingly.
[0081] In the embodiment of the present invention, after obtaining the training data set, the initial gray-scale change optimization model can be called to predict the change action of the training status data in each training data included in the training data set respectively, and obtain the predicted probability value of each change action under each training status data. And based on the predicted probability value of each change action under each training status data, the model loss value of the initial gray-scale change optimization model is determined. Further, the model parameters in the initial gray-scale change optimization model can be optimized in the direction of reducing the model loss value to obtain the initial gray-scale change optimization model after model optimization, and based on the initial gray-scale change optimization model after model optimization, the target gray-scale change optimization model is determined. Based on this, the target status data of the target service can be obtained, and the target gray-scale change optimization model can be called to predict the change action of the target status data to obtain the target predicted probability value of each change action among multiple change actions. Then correspondingly, based on the target predicted probability value of each change action, the target change action can be determined from multiple change actions, and the target service can be gray-scaled according to the target change action. It can be seen that the embodiment of the present invention can further improve the model performance of the gray-scale change optimization model through the training data set to implement the automated intelligent hierarchical gray-scale change method, that is, through continuous learning and optimization, it can automatically select the optimal change strategy (i.e., the target change action) according to historical data and the current environmental state. And during the change process, in order to cope with possible abnormal states, the embodiment of the present invention designs an automatic change loss stop strategy, so that when an abnormal state (such as an extended service response time, an increased error rate, etc.) is detected, the change loss stop strategy can be automatically triggered, thereby minimizing the online loss to ensure that high attention is always paid to the business stability during the change process and the smooth progress of the change is determined.
[0082] Based on the description of the related embodiments of the above gray-scale change method, an embodiment of the present invention also proposes a gray-scale change device, which can be a computer program (including program code) running in an electronic device; as Figure 5 shown, the gray-scale change device may include an acquisition unit 501 and a processing unit 502. The gray-scale change device can execute Figure 1 or Figure 3 the gray-scale change method shown, that is, the gray-scale change device can run the above units:
[0083] The acquisition unit 501 is used to acquire the target status data of the target service;
[0084] The processing unit 502 is used to call the target gray-scale change optimization model to predict change actions for the target status data, and obtain the target prediction probability values of each change action among multiple change actions;
[0085] The processing unit 502 is further used to determine a target change action from the multiple change actions based on the target prediction probability values of the respective change actions, and perform gray-scale change on the target service according to the target change action.
[0086] In one implementation, the acquisition unit 501 can also be used to: acquire a training data set, and one training data includes one training status data;
[0087] The processing unit 502 can also be used to:
[0088] Call the initial gray-scale change optimization model to predict change actions for the training status data in each training data included in the training data set, and obtain the prediction probability values of each change action under each training status data;
[0089] Based on the prediction probability values of each change action under each training status data, determine the model loss value of the initial gray-scale change optimization model;
[0090] Optimize the model parameters in the initial gray-scale change optimization model in the direction of reducing the model loss value to obtain the initial gray-scale change optimization model after model optimization, and determine the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization.
[0091] In another implementation, when the processing unit 502 determines the model loss value of the initial gray-scale change optimization model based on the prediction probability values of each change action under each training status data, it can specifically be used to:
[0092] Based on the prediction probability values of each change action under each training status data, determine the training prediction actions under each training status data;
[0093] In a simulated production environment, perform gray-scale changes according to the training prediction actions under each training status data respectively, to determine the predicted reward values corresponding to each training data;
[0094] Determine the label reward values corresponding to each training data, and based on the predicted reward values and label reward values corresponding to each training data, determine the model loss value of the initial gray-scale change optimization model.
[0095] In another implementation manner, when the processing unit 502 performs gray-scale change on the target service according to the target change action, it may specifically be used for:
[0096] When the target change action is a continuous gray-scale change action, determine multiple change stages, and perform step-by-step changes according to the multiple change stages and the target change action, so as to implement the gray-scale change of the target service according to the target change action, and one change stage corresponds to one change level;
[0097] During the change process, continuously determine the current monitoring data, and based on the change monitoring policy in the current change stage and the current monitoring data, determine whether the current change process is abnormal, so as to perform a stop-loss operation according to the change stop-loss policy in the current change stage when detecting that the change process is abnormal.
[0098] In another implementation manner, the multiple change stages include a pre-release instance stage, a small-traffic stage, a single-data-center ramp-up stage, and a multi-data-center ramp-up stage; wherein, the pre-release instance stage refers to the change process of pre-release instances, the small-traffic stage refers to the change and ramp-up process of a preset number of online instances, the single-data-center ramp-up stage refers to the change and ramp-up process of online instances in a single data center, and the multi-data-center ramp-up stage refers to the change and ramp-up process of online instances in multiple data centers.
[0099] In another implementation manner, the processing unit 502 may also be used for:
[0100] Determine the initial monitoring values of each change preparation monitoring index among multiple change preparation monitoring indexes before the change of the target service;
[0101] Based on the initial monitoring values of each change preparation monitoring index, determine the disconnection monitoring values of each change preparation monitoring index when disconnecting the online traffic in a single data center;
[0102] Based on the disconnection monitoring values of the monitoring metrics prepared for each change, determine whether the target service meets the cut-over condition at the computer room level, and when it is determined that the target service meets the cut-over condition at the computer room level, trigger the execution of the gray-scale change of the target service according to the target change action.
[0103] In an embodiment of the present invention, after obtaining the target status data of the target service, the target gray-scale change optimization model can be called to predict the change actions for the target status data, and the target prediction probability values of each change action among multiple change actions can be obtained; further, based on the target prediction probability values of each change action, the target change action can be determined from multiple change actions, and the target service can be gray-scale changed according to the target change action. It can be seen that in the embodiment of the present invention, the target change action can be conveniently determined through the target gray-scale change optimization model, and thus the gray-scale change can be conveniently performed through the target change action, which can effectively improve the response speed.
[0104] Based on the descriptions of the above method embodiments and apparatus embodiments, an exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.
[0105] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0106] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0107] Reference Figure 6 , the structural block diagram of the electronic device 600 that can be used as the server or client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0108] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0109] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information into the electronic device 600. The input unit 606 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 607 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include but is not limited to a magnetic disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0110] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above. For example, in some embodiments, the grayscale change method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 can be configured to execute the grayscale change method in any other appropriate manner (e.g., by means of firmware).
[0111] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0112] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0114] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).
[0115] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0116] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs that run on the respective computers and have a client - server relationship with each other.
[0117] Moreover, it should be understood that the above - disclosed are only the preferred embodiments of the present invention, and of course cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made in accordance with the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A gray-scale change method, characterized in that, Including: Obtain the target status data of the target service; Invoke the target gray-scale change optimization model to predict change actions for the target status data, and obtain the target prediction probability values of each change action among multiple change actions; Based on the target prediction probability values of each change action, determine the target change action from the multiple change actions, and perform gray-scale change on the target service according to the target change action.
2. The method according to claim 1, wherein The method further includes: Obtain a training data set, where one training data includes one training status data; Invoke the initial gray-scale change optimization model to predict change actions for the training status data in each training data included in the training data set, and obtain the prediction probability values of each change action under each training status data; Based on the prediction probability values of each change action under each training status data, determine the model loss value of the initial gray-scale change optimization model; Optimize the model parameters in the initial gray-scale change optimization model in the direction of reducing the model loss value to obtain the initial gray-scale change optimization model after model optimization, and determine the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization.
3. The method according to claim 2, wherein The determining the model loss value of the initial gray-scale change optimization model based on the prediction probability values of each change action under each training status data includes: Based on the prediction probability values of each change action under each training status data respectively, determine the training prediction actions under each training status data; In a simulated production environment, perform gray-scale change according to the training prediction actions under each training status data respectively to determine the prediction reward values corresponding to each training data; Determine the label reward values corresponding to each training data, and based on the prediction reward values and label reward values corresponding to each training data, determine the model loss value of the initial gray-scale change optimization model.
4. The method according to any one of claims 1 to 3, characterized in that, The performing gray-scale change on the target service according to the target change action includes: When the target change action is a continuous gray-scale change action, determine multiple change stages, and perform step-by-step change according to the multiple change stages and the target change action to implement the gray-scale change on the target service according to the target change action, where one change stage corresponds to one change level; During the change process, continuously determine the current monitoring data, and based on the change monitoring strategy in the current change stage and the current monitoring data, determine whether the current change process is abnormal, so as to perform a stop-loss operation according to the change stop-loss strategy in the current change stage when it is detected that the change process is abnormal.
5. The method according to claim 4, characterized in that, The multiple change stages include a pre-release instance stage, a small-traffic stage, a single-data center scale-up stage, and a multi-data center scale-up stage; among them, the pre-release instance stage refers to the change process of pre-release instances, the small-traffic stage refers to the change and scale-up process of a preset number of online instances, the single-data center scale-up stage refers to the change and scale-up process of online instances in a single data center, and the multi-data center scale-up stage refers to the change and scale-up process of online instances in multiple data centers.
6. The method according to any one of claims 1 to 3, characterized in that The method further includes: Determining an initial monitoring value of each change preparation monitoring metric among a plurality of change preparation monitoring metrics before the target service change; Based on the initial monitoring values of the respective change preparation monitoring metrics, determining a disconnection monitoring value of each change preparation monitoring metric when the traffic on the single computer room disconnection line is cut off; Based on the disconnection monitoring values of the respective change preparation monitoring metrics, determining whether the target service meets the cut-flow condition at the computer room level, and when it is determined that the target service meets the cut-flow condition at the computer room level, triggering the execution of the gray-scale change of the target service according to the target change action.
7. A gray-scale change device, characterized in that, The apparatus includes: An acquisition unit, configured to acquire target state data of a target service; A processing unit, configured to call a target gray-scale change optimization model to perform change action prediction on the target state data, and obtain a target prediction probability value of each change action among a plurality of change actions; The processing unit is further configured to determine a target change action from the plurality of change actions based on the target prediction probability values of the respective change actions, and perform gray-scale change on the target service according to the target change action.
8. The device according to claim 7, characterized in that, The acquisition unit is further configured to: acquire a training data set, and one training data includes one training state data; The processing unit is further configured to: call an initial gray-scale change optimization model, and perform change action prediction on the training state data in each training data included in the training data set respectively, to obtain prediction probability values of the respective change actions under each training state data; Based on the prediction probability values of the respective change actions under each training state data, determining a model loss value of the initial gray-scale change optimization model; Optimizing the model parameters in the initial gray-scale change optimization model in a direction of reducing the model loss value, obtaining an initial gray-scale change optimization model after model optimization, and determining the target gray-scale change optimization model based on the initial gray-scale change optimization model after model optimization.
9. An electronic device, characterized in that, Including: A processor; And A memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-6.