Application Management Method and Device, Electronic Device, and Computer-Readable Storage Medium

By integrating the application operation status detection results and container status, the problems of low single-machine detection efficiency and relying on manual operations in the existing technology are solved, and more efficient and accurate application management is achieved, ensuring the stability and high availability of services.

CN114625478BActive Publication Date: 2025-06-24ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210126119.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-10
Publication Date
2025-06-24
Estimated Expiration
2042-02-10

AI Technical Summary

Technical Problem

In the prior art, application management based on stand-alone detection is inefficient and relies on manual operations, resulting in the failure of containers to restart in time to affect user use.

Method used

By obtaining the application's running status detection results and container status, comprehensively determine the application with a predetermined running status, and perform corresponding operations to avoid error restarts and service stability problems caused by stand-alone detection.

Benefits of technology

Improves the efficiency and accuracy of application management, reduces the dependence of manual operations, and ensures the stability and high availability of services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625478B_ABST
    Figure CN114625478B_ABST
Patent Text Reader

Abstract

The present application discloses an application management method, an apparatus, an electronic device, and a computer-readable storage medium. The method includes: obtaining a running state detection result for each application in at least one group of applications; obtaining the container states of a group of containers that respectively run the group of applications; determining, according to the running state detection results of the respective applications and the container states of the containers that run the applications, the applications having a predetermined running state in each group of applications; and performing a predetermined operation on the applications having the predetermined running state. The embodiments of the present application can make a comprehensive decision for management by comprehensively considering the entirety of applications with the same application identifier, avoiding the problem that the service provided by the group of applications has poor stability due to lack of overall control when performing application management only based on the running state detection results of individual applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud computing technology, and in particular to an application management method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of cloud computing technology, more and more applications can provide services to users with the help of a wide range of cloud computing resources. In recent years, cloud-native technologies based on cloud computing have emerged. With the help of cloud computing environment, native technologies are designed based on cloud computing systems, so that they can run in a better state on cloud computing resources and fully utilize the advantages of cloud platform distribution and elasticity. In the cloud-native system, containers are one of the basic elements of the cloud-native system. Containerization can provide basic guarantees for cloud-native microservices, and containers can be managed through container orchestration systems such as K8S.

[0003] In a cloud-native system, applications or services can be generated as microservices provided to users through containerization. Therefore, an application can improve its operating efficiency, reduce the load of a single container, and even provide disaster recovery performance in the event of a failure by making multiple copies and running them in multiple containers. For example, when an application running in a container is in an unhealthy state, such as when a process is suspended or a service is abnormal, the container is actually unable to provide services to the outside as an independent microservice unit. Therefore, it is usually possible to issue an alarm for such an unhealthy container through pre-settings, and the maintenance personnel manually restart the container to restore the normal operation of the application. However, such manual maintenance repetitive operations rely on manual inspection and timely operations, which is not only very inefficient, but also may cause one or some containers to fail to restart in time due to insufficient manpower, affecting the normal use of users. Summary of the invention

[0004] The embodiments of the present application provide an application management method and device, an electronic device, and a computer-readable storage medium to solve the defects of the prior art based on single-machine detection, which is complex in configuration and low in efficiency.

[0005] To achieve the above object, an embodiment of the present application provides an application management method, the method comprising:

[0006] Obtaining a running status detection result for each application in at least one group of applications;

[0007] Get the container status of a group of containers in which the group of applications are respectively run;

[0008] Determine the applications with a predetermined running state in each group of applications according to the running state detection results of each application and the container state of the container in which the application runs;

[0009] Perform a predetermined operation on the applications with a predetermined running state.

[0010] An embodiment of the present application further provides an application management device, including:

[0011] A first status acquisition module, configured to acquire the running state detection results for each application in at least one group of applications;

[0012] A second status acquisition module, configured to acquire the container status of a group of containers in which the group of applications runs respectively;

[0013] A determination module, configured to determine the applications with a predetermined running state in each group of applications according to the running state detection results of each application and the container state of the container in which the application runs;

[0014] An execution module, configured to perform a predetermined operation on the applications with a predetermined running state.

[0015] An embodiment of the present application further provides an electronic device, including:

[0016] A memory, configured to store programs;

[0017] A processor, configured to run the programs stored in the memory, and when the programs run, execute the application management method provided by the embodiment of the present application.

[0018] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program executable by a processor is stored, wherein when the program is executed by the processor, it implements the application management method provided by the embodiment of the present application.

[0019] The application management method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present application obtain the running state detection results of a group of applications with the same application identifier, and obtain the container status of the containers corresponding to the group of applications, so as to make a comprehensive judgment based on the running state detection results of the applications and the container status of the containers in which the applications run, to determine the applications with a predetermined running state in the group of applications, and perform operations on the applications with a predetermined running state thus determined. Therefore, it is possible to make a comprehensive management decision by considering the whole of the applications with the same application identifier, avoiding the problem that the management of the applications is performed only based on the running state detection results of a single application, resulting in poor service stability of the group of applications due to lack of overall control.

[0020] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specific embodiments of this application are specifically given. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0022] Figure 1 is a schematic diagram of the application scenario of the application program management solution provided by the embodiment of this application;

[0023] Figure 2 is a flowchart of an embodiment of the application program management method provided by this application;

[0024] Figure 3 is a flowchart of another embodiment of the application program management method provided by this application;

[0025] Figure 4 is a schematic structural diagram of an embodiment of the application program management device provided by this application;

[0026] Figure 5 is a schematic structural diagram of an embodiment of the electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0028] Embodiment 1

[0029] The solution provided by the embodiment of this application can be applied to any system with cloud application management capabilities, such as a cloud service system including a cloud application management module, and so on. Figure 1 is a schematic diagram of the application scenario of the application program management solution provided by the embodiment of this application, Figure 1 The shown scenario is only one of the examples of the principle of the technical solution of this application.

[0030] With the development of cloud computing technology, more and more applications can utilize extensive cloud computing resources to provide services to users. In recent years, cloud-native technologies based on cloud computing have emerged. They are designed natively based on the cloud computing system, enabling them to operate in a better state on cloud computing resources and fully leveraging the distributed and elastic advantages of the cloud platform. In the cloud-native system, containers are one of the basic elements. Containerization can provide basic guarantees for cloud-native microservices, and through container orchestration systems such as K8S, containers can be managed.

[0031] In the cloud-native system, an application or service can be containerized to generate microservices provided to users. Therefore, an application can improve its running efficiency, reduce the load on a single container, and even enhance its disaster tolerance performance in case of failures by creating multiple replicas and running them in multiple containers. For example, when an application running in a container is in an unhealthy state, such as a process being suspended or a service being abnormal, the container can no longer provide services as an independent microservice unit. Usually, an alarm can be set in advance for such an unhealthy container, and maintenance personnel can manually restart the container to restore the normal running state of the application. However, such manual maintenance and repetitive operations rely on manual inspection and timely operation. The efficiency is not only very low but also may affect the normal use of users due to insufficient manpower, resulting in some containers not being restarted in time.

[0032] In the prior art, container probes have been proposed to be set in the container management system, enabling the probes to periodically enter specified containers according to user configurations to detect whether the container services are normal. For example, the probe can be a command script program. By executing the command script program in the specified container and judging whether the application in the container is in a normal running state based on the execution result of the command script program. For example, users can configure the command script program according to a set of containerized applications so that the execution of the command script program can reflect the running state of the applications in the container. After configuration, the configured detection command script can be implanted into the containers of all application replicas that provide services externally, and for example, when a predetermined exit code is successfully generated during execution, it can be determined that the detection is successful, that is, the running state of the application in the container is normal.

[0033] Conversely, if the exit code generated by the execution of the command script program is not zero, it can indicate that the script execution fails. Therefore, in the prior art, based on this result, it can be judged that the application in the container is running abnormally, and thus the container can be directly restarted to restore the normal running of the application in the container.

[0034] However, in this existing technical solution, the running state of the application in the container is actually indirectly reflected by the execution result of the implanted detection command script. Therefore, if the configuration of the command script is unreasonable, it will affect the accuracy of detecting the running state of the application, and may even lead to incorrect judgments. Therefore, in the prior art, it is usually required that the maintainer of the application has rich experience and professional knowledge to perform such configuration.

[0035] In addition, especially in the cloud-native system, multiple application replicas are formed based on an application and run in individual containers respectively to jointly provide services externally. That is, in fact, this group of application replicas constitutes a whole to provide each part of the services externally respectively. Therefore, in such a case, if a single container is restarted unreasonably or even incorrectly due to the unreasonable configuration of the detection command script, it will directly affect a part of the service. Moreover, if the application running in a single container is relatively important in the whole, the direct restart of such a container may have a greater impact on the normal operation of the entire service, and even cause losses to users.

[0036] In the solution of the prior art, the solution of managing a group of applications based on the detection results of a single container has potential safety hazards in practical applications. Since in the cloud-native system, it is necessary to measure the impact of the running states of the application replicas in multiple containers providing microservices on the overall service from an overall perspective. For this reason, maintaining a single-machine probe detection for each container requires the maintainer to have a deep understanding of the application and also requires relatively professional knowledge. In addition, since the relationships among the individual containers providing microservices in the cloud-native system are complex, such a one-by-one single-machine configuration is also prone to incorrect configuration by the maintainer due to cumbersome configuration or complex content. Once such an incorrect configuration takes effect as determined by the maintainer, it is likely to cause the containers providing microservices to be restarted incorrectly or at the wrong time, resulting in the externally provided service being affected or even unable to provide services externally.

[0037] In particular, the probe detection for a single container currently adopted in the prior art is uniformly configured without discrimination for each application in a group of applications during configuration. Therefore, once the configuration takes effect, it will take effect in all containers of this group of applications. In other words, if the probe detection method is incorrectly configured due to the mistake of the maintainer, it may cause all the applications running in all containers to be restarted, resulting in all applications being unable to provide services externally, seriously affecting the stability and high availability of the service.

[0038] For example, as Figure 1 shown in Figure 1shows an application scenario to which the application program management method of the present application can be applied. In Figure 1 In the scenario shown, on the cloud server, three containers 1-3 can be generated for three replicas of an application program to run the three replica programs respectively, so that the three application program replicas 1-3 running in the containers 1-3 can be provided as a whole to the outside in the form of microservices. Therefore, in order to ensure the stability of the service, a detection command script program can be configured for the application programs 1-3. The script program can be executed in the containers 1-3 respectively. Therefore, in the prior art, the running status of the application programs 1-3 can be judged according to the execution results of the above command script program in the containers 1-3. In particular, in the embodiments of the present application, the maintenance personnel usually configure the script program only according to experience, and after the script program is configured, it can take effect for all three application programs 1-3. Therefore, when, for example, the execution result of the script program in container 1 generates an exit code of zero, or generates a successful execution result identifier, it can be considered that the running status of the application program in this container is normal. On the contrary, when the execution result of the script program in a certain container is non-zero, or returns a failed execution result, then in the prior art, it will be judged that the running status of the application program in this container is abnormal, and thus a restart instruction can be directly issued to this container to restore the running of the application program. However, as described above, in the prior art, if the script program fails to run in container 1 due to a configuration error of the script program, and in fact the application program 1 is always running normally in this container 1, but the script application program will return a non-zero exit code due to the running failure, then this will cause this container to be restarted. In particular, after the restart, the script program will continue to be used to detect the running status of the application program. Therefore, inevitably, these script programs will return failed execution results every time they run in the container, so that the containers 1-3 of the application programs 1-3 are restarted repeatedly, and finally the services provided by the application programs 1-3 to the outside are abnormal.

[0039] In response to this, in the embodiments of the present application, after obtaining the execution result of the script program, the application program is not directly managed based on this execution result, but the current container status of the container is further obtained, and the running status of the application program in the container is comprehensively judged based on this container status and the execution result of the script program. In particular, in the embodiments of the present application, when judging the running status of the application program, the execution results of the script programs of all the application programs in a group of this application program are obtained, so that the application program with a predetermined status can be managed while trying to ensure the normal service based on the predetermined strategy of this group of application programs.

[0040] Specifically, the technical solution of the embodiment of the present application can use the Kubernetes framework system as the container management solution. In this solution, the embodiment of the present application can achieve non-invasive configuration, draw on the livenessProbe probe capability in the art for in-depth functional expansion and technological evolution, and thus propose an enhanced livnessProbe probe controller. In this solution, the controller can adopt a standard Kubernetes detection module and can be based on the ControllerRuntime architecture. Therefore, in the embodiment of the present application, the overall implementation of the cloud service system is based on centralized deployment, and the influence scope of the entire deployment domain is a Kubernetes cluster. As described above, for each single container on each server in the cluster, the container status is reported periodically, and the container status can be further comprehensively judged based on the enhanced livenessProbe probe control module.

[0041] Under the microservices architecture, applications are deployed in a containerized form, and it is generally considered that applications run inside containers. Therefore, the detection script component can be arranged to perform detection on a single machine and report the status to the container level, and then the management module of the application program management method of the embodiment of the present application, such as a cloud service system, can make a comprehensive decision. After the enhanced livenessProbe controller according to the embodiment of the present application senses the change in the container status, it executes internal logical calculations and decides whether to restart this container. Finally, when it is determined that the application program is running abnormally, an instruction can be issued through the enhanced livenessProbe controller, and the corresponding container can be triggered to restart. In other words, in the embodiment of the present application, after the user edits and configures the detection script, only the detection is performed on the single-machine container side, and the detection results are reported by the container, for example, and the management module, such as the detection controller, comprehensively judges the execution results of the script and the container status, especially can further comprehensively determine the containers that need to be restarted based on the judgment results of this group of application program replicas, so as to be able to manage and control the application program from the perspective of ensuring service stability and reliability.

[0042] For example, in the embodiments of the present application, the following three methods, namely ExecAction, HTTPGetAction, and TCPSocketAction, can be used to obtain the detection results of the running status of the application. For example, the detection result Success can indicate passing the detection, Failure indicates failing the detection, and Unknown indicates that the detection has not been carried out normally. Specifically, for the ExecAction detection method, a specified script command can be executed in the container. If the execution is successful and the exit code is 0, the detection is successful; for the HTTPGetAction method, the HTTP Get method can be called through the IP address, port number, and path of the container. If the response status code is greater than or equal to 200 and less than 400, the container is considered healthy; for the TCPSocketAction, a TCP check can be performed through the IP address and port number of the container. If a TCP connection can be established, it indicates that the container is healthy.

[0043] Therefore, in the solution of the embodiments of the present application, after the user completes the configuration of the probe, it can take effect immediately at the level of all containers. This is a potential risk to the reliability of the application in the prior art solutions. For example, as described above, if the probe is configured incorrectly, adopting the community logic will cause the application to be restarted in full, which is a fatal point in terms of stability and high availability of the service. However, in the solution of the embodiments of the present application, the detection execution result is reported to the container control module through the container side, so that the container control module can consider the high reliability of the service provided by the overall application from the centralized concept and the global perspective of all application programs, especially from the overall perspective of a group of application programs.

[0044] In addition, in the embodiments of the present application, since the overall reliability of all application program replicas providing services can be considered, a threshold for the number of containers restarted each time can be further introduced to prevent the restart of all containers caused by configuration errors. For example, in the embodiments of the present application, the maximum unavailable ratio maxUnAvailable value can be set, so that it can provide a fallback protection in combination with the actual number of containers of the current application.

[0045] In addition, in the embodiments of the present application, since the number of application program copies providing different services is different, when setting the maximum unavailable ratio, the current number of application copies can also be considered to dynamically adjust the ratio status. For example, if the current number of container copies is 10 and maxUnAvailable = 20% is set, only 2 containers are allowed to be triggered to restart at the same time. In actual applications, this threshold can be flexibly set by maintenance personnel according to the actual application requirements. For example, if the number of copies < 10, then maxUnAvailable = 1; if the number of copies >= 10, then maxUnAvailable = 20%. Of course, in the embodiments of the present application, the pdbproducer controller in the open source OpenKruise can also perform real-time dynamic maintenance on the maxUnAvailable value according to the capacity and rule policies of the service.

[0046] In addition, as described above, if the probe scheme configuration fails, for example, the configuration of the command script is incorrect, the script will return a failed result every time it is executed in the container, but in fact the container and the application program are both running normally. Therefore, in the embodiments of the present application, the number of times the application program detection result is failed within a predetermined time period can be further counted, and if the number of failures within a certain period of time reaches a predetermined threshold, the above-mentioned configuration error may occur. Therefore, an alarm can be issued to the user in this case.

[0047] In addition, in the embodiments of the present application, when judging the actual running state of the application program, since the detection result and the container state are comprehensively considered, different priorities can be further assigned to the combination of the detection result and the container state to achieve a more accurate judgment. For example, the detection result can include: execution success and execution failure, and the container state can include: container available and container unavailable. Therefore, in the embodiments of the present application, the priority of the first combination of execution failure and container unavailable can be set to be greater than the second combination of execution failure and container available.

[0048] The application program management solution provided by the embodiments of the present application obtains the running state detection results of a group of application programs with the same application program identifier, and obtains the container state of the containers corresponding to the group of application programs, so as to make a comprehensive judgment based on the running state detection results of the application programs and the container state of the containers in which the application programs run, to determine the application programs with a predetermined running state in the group of application programs, and perform operations on the application programs with the predetermined running state thus determined. Therefore, it is possible to make a comprehensive decision on management by considering the whole of the application programs with the same application program identifier, avoiding the problem that the management of the application programs is performed only based on the running state detection results of a single application program, resulting in poor service stability of the group of application programs due to lack of overall control.

[0049] The above embodiments illustrate the technical principles and exemplary application frameworks of the embodiments of the present application. The following further describes the specific technical solutions of the embodiments of the present application through multiple embodiments.

[0050] Embodiment 2

[0051] Figure 2 As shown in the flowchart of an embodiment of the application management method provided by the present application, the execution subject of this method can be various terminal or server devices with cloud application management capabilities, or can also be a device or chip integrated on these devices. As Figure 2 shown, the application management method includes the following steps:

[0052] S201, obtain the running state detection results for each application in at least one group of applications.

[0053] In step S201, the running state detection results of each application in at least one group of applications can be obtained. In the embodiments of the present application, the obtained detection results can be used as the running state detection results in step S201 by executing a pre-configured detection scheme in the container running each application. Specifically, in the embodiments of the present application, the detection results can be obtained for a group of applications with the same application identifier. Therefore, in a microservices architecture, the detection status of a group of microservices application replicas that provide services as a whole to the outside can be obtained.

[0054] S202, obtain the container status of a group of containers in which the group of applications are respectively running.

[0055] In step S202, the container status of the containers in which the applications corresponding to the detection results obtained in step S201 are running can be obtained. For example, in a cloud-native architecture, each application can run in a container to provide different service contents respectively. Therefore, in the embodiments of the present application, in addition to obtaining the execution results of probes, such as, as detection results, the status of the container in which the application is running is further obtained in step S202 for comprehensive judgment.

[0056] S203, determine the applications with a predetermined running state in each group of applications according to the running state detection results of each application and the container status of the containers in which the applications are running.

[0057] In step S203, the actual running state of the application can be comprehensively determined based on the detection result obtained in step S201 and the container state obtained in step S202. For example, in practical applications, the user configures the probe to run in the container of the application to detect the running state of the application. Therefore, the detection result of the probe can only indirectly reflect the running state of the application. Moreover, as described above, the accuracy of the detection result of the probe depends on the specific configuration of the probe by the maintenance personnel. Therefore, if the configuration is unreasonable or even incorrect, it will cause the probe to fail to execute in the container. In the prior art, if only such an execution result is used as the sole basis for judging the running state of the application, it will lead to incorrect handling of the application, such as restarting the application when it is actually running normally. In step S203 of the present application, a comprehensive judgment can be made based on the detection result and further considering the actual state of the container to eliminate the problem of incorrect detection results caused by the detection scheme configuration problem on the single-machine side.

[0058] S204, perform a predetermined operation on the application with a predetermined running state.

[0059] Therefore, in step S204, a predetermined operation can be performed on the application determined to have a predetermined running state based on the result determined in step S203. For example, when it is determined in step S203 that the running state of the application is abnormal or fails, the application can be restarted in step S204 to restore the normal running of the application.

[0060] The application management method provided by the embodiments of the present application obtains the running state detection results of a group of applications with the same application identifier, and obtains the container state of the containers corresponding to the group of applications, so as to make a comprehensive judgment based on the running state detection results of the applications and the container state of the containers in which the applications run, to determine the applications with a predetermined running state in the group of applications, and perform operations on the applications with a predetermined running state thus determined. Therefore, it is possible to make a comprehensive decision for management by comprehensively considering the entirety of the applications with the same application identifier, avoiding the problem of poor service stability provided by the group of applications due to lack of overall control when performing application management only based on the running state detection results of a single application.

[0061] Embodiment III

[0062] Figure 3The flowchart of another embodiment of the application management method provided by this application. The execution subject of this method can be various terminal or server devices with cloud native application management capabilities, or can also be a device or chip integrated on these devices. As Figure 3 shown, the application management method includes the following steps:

[0063] S301, execute a predetermined command script in the containers of each application.

[0064] In step S301, a predetermined command script can be executed in the containers running each application to detect the running status of the application. For example, in the prior art, maintenance personnel can configure a command script program and execute the configured script program in the container as a means to detect the application running in the container.

[0065] S302, determine the running status detection result of each application according to the execution result of the command script.

[0066] In step S302, the running status detection result of the application in the container can be determined according to the execution result of the command script executed in step S301. For example, as Figure 1 shown, the user can pre-configure a probe script program for the applications running in containers 1-3, and execute the script program in containers 1-3 respectively in step S301. Thus, in step S302, it can be obtained that the execution result in container 1 is that a zero exit code is generated, or a successful execution result identifier is generated. Therefore, in step S302, it can be considered that the running status of the application in this container is normal. When the execution result of the script program in a certain container is non-zero, or a failed execution result is returned in step S302, then in step S302, the failed execution result can be used as the running status detection result of the application.

[0067] S303, obtain the container status of a group of containers in which this group of applications are respectively running.

[0068] In step S303, the container status of the container running the application targeted by the detection result obtained in step S302 can be obtained. For example, in the cloud native system, each application can run in a container to provide different service contents respectively.

[0069] For example, when the execution result of the command script executed in step S301 obtained in step S302 is non-zero, or a failed execution result is returned, then in the prior art, it will be determined that the running state of the application program in the container is abnormal, and thus a restart instruction can be directly issued to the container to resume the running of the application program. However, as described above, in the prior art, if the script program fails to run in container 1 due to incorrect configuration of the script program, and in fact, application program 1 has been running normally in this container 1, but the script application program will return a non-zero exit code due to the running failure, then this will cause the container to be restarted. In particular, after the restart, the script program will continue to be used to detect the running state of the application program. Therefore, inevitably, these script programs will return failed execution results every time they run in the container, causing containers 1-3 of application programs 1-3 to be restarted repeatedly, and ultimately resulting in abnormal services provided by application programs 1-3 to the outside world.

[0070] Therefore, in the embodiments of the present application, in addition to obtaining the execution result of, for example, a probe as the detection result, the state of the container in which the application program runs is further obtained in step S303 for comprehensive judgment.

[0071] S304. Determine the application program with a predetermined running state according to the priority of the combination of the running state detection result of the application program and the container state.

[0072] In step S304, the running state of the application program can be judged based on the combination of the detection result determined in step S302 and the container state obtained in step S303. In particular, in the embodiments of the present application, in order to comprehensively consider the detection result and the container state when judging the actual running state of the application program, different priorities can be further assigned to the combination of the detection result and the container state to achieve more accurate judgment. For example, the detection result can include: execution success and execution failure, and the container state can include: container available and container unavailable. Therefore, in the embodiments of the present application, the priority of the first combination of execution failure and container unavailable can be set to be greater than the second combination of execution failure and container available. Therefore, in the scenario shown in, for example Figure 1 when it is determined in step S302 that the execution result of the command script in container 1 is failed and the execution result of the command script in container 2 is failed, and the state of container 1 is obtained as available and the state of container 2 is unavailable in step S303, then in step S304, it can be determined to preferentially restart container 2 based on the fact that the combination priority of the script execution result and the container state in container 2 is higher than that in container 1, so as to resume the running of application program 2 in container 2.

[0073] S305. Determine an operation threshold according to the total number of applications in a set of applications.

[0074] S306. Manage the set of applications according to the proportion of applications with a predetermined running state in each set of applications.

[0075] In step S305, the operation threshold can be determined according to the total number of applications with the same identifier detected in step S301. In particular, in the solution of the embodiment of the present application, after the user completes the configuration of the probe, it can take effect immediately at all container levels, which is a potential risk to the reliability of the application in the solution of the prior art. For example, as described above, if the probe configuration is incorrect, adopting the community logic will cause all applications to be restarted in full, which is a fatal point in terms of stability and high availability of the service. Therefore, in the embodiment of the present application, in step S302, the container side can report the detection execution result to the container control module, and further obtain the container status in step S303, so that the container control module can, based on the centralized idea, consider the high reliability of the service provided by the overall application from the global perspective of a set of application detection results and container status, especially from the overall perspective of a set of applications.

[0076] Therefore, in step S305, the overall reliability of all application replicas providing services can be considered to determine the number threshold for each operation in a set of applications, so as to prevent the restart of all containers caused by configuration errors.

[0077] For example, in step S305, the maximum unavailable ratio value can be set in combination with the actual number of containers of the current application for overall fallback protection of the application. Specifically, since the number of application programs providing different services varies, when setting the maximum unavailable ratio in step S305, the total number of current application programs can be considered to dynamically adjust the ratio status. For example, when the current number of container replicas is 10, the maximum unavailable ratio value can be set to 20% in step S305, so that only 2 containers are allowed to be triggered for restart at the same time in step S306. In actual applications, this threshold can be flexibly set by maintenance personnel according to the actual application requirements. For example, when the total number of a group of application programs is less than 10, the maximum unavailable quantity value can be fixed at 1, that is, only 1 container is allowed to be triggered for restart at the same time in step S306. When the total number of a group of application programs is greater than or equal to 10, the maximum unavailable ratio value can be set to 20%, that is, only 20% of all application programs are allowed to be restarted at the same time in step S306. Of course, in the embodiments of the present application, the open source OpenKruise's pdbproducer controller can also perform real-time dynamic maintenance on this threshold according to the capacity and rule policies of the service.

[0078] In addition, if the user makes an error in configuring the detection scheme, when executing the script in the container in step S301, a failed result will surely be returned in step S302, but in fact both the container and the application program are running normally. Therefore, in step S306, the number of times the application program detection result is failed within a predetermined time period can be further counted, and if the number of failures within a period of time reaches a predetermined threshold, the above-mentioned configuration error may occur. Therefore, in this case, an alarm can be issued to the user without restarting the container.

[0079] The application program management method provided by the embodiments of the present application obtains the running status detection results of a group of application programs with the same application program identifier, and obtains the container status of the containers corresponding to the group of application programs, so as to make a comprehensive judgment based on the running status detection results of the application programs and the container status of the containers in which the application programs run, to determine the application programs with a predetermined running status in the group of application programs, and perform operations on the determined application programs with the predetermined running status. Therefore, it is possible to make a comprehensive decision on management by considering the whole of the application programs with the same application program identifier, avoiding the problem that the management of application programs is performed only based on the running status detection results of a single application program, resulting in poor service stability provided by the group of application programs due to lack of overall control.

[0080] Embodiment 4

[0081] Figure 4The structural schematic diagram of the application management device embodiment provided by this application can be used to execute the method steps as shown in Figure 2 and Figure 3 As shown in Figure 4 The application management device may include: a first status acquisition module 41, a second status acquisition module 42, a determination module 43, and an execution module 44.

[0082] The first status acquisition module 41 may be used to obtain the operation status detection results for each application in at least one group of applications.

[0083] The first status acquisition module 41 may obtain the operation status detection results of the applications in at least one group of applications. In the embodiments of this application, the first status acquisition module 41 may use a pre-configured detection scheme executed in the container running each application as the operation status detection result. Specifically, in the embodiments of this application, the detection results may be obtained for a group of applications with the same application identifier. Therefore, in a microservices system, the detection status of a group of microservices application copies that provide services outward as a whole may be obtained.

[0084] Specifically, the first status acquisition module 41 may execute a predetermined command script in the container running each application to detect the operation status of the application and determine the operation status detection result of each application according to the execution result of the command script. For example, in the prior art, maintenance personnel may configure a command script program and execute the configured script program in the container as a means for detecting the application running in the container. Therefore, the first status acquisition module 41 may determine the operation status detection result of the application in the container according to the execution result of the executed command script. For example, as shown in Figure 1 A user may pre-configure a detection script program for the applications running in containers 1-3 and execute the script program in containers 1-3 respectively. As a result, it can be obtained that the execution result in container 1 is that a zero exit code is generated, or a successful execution result identifier is generated. Therefore, the first status acquisition module 41 may consider that the operation status of the application in this container is normal. When the first status acquisition module 41 obtains that the execution result of the script program in a certain container is non-zero, or returns a failed execution result, then the first status acquisition module 41 may use the failed execution result as the operation status detection result of the application.

[0085] The second status acquisition module 42 may be used to obtain the container status of a group of containers in which the group of applications are respectively running.

[0086] The second status acquisition module 42 can acquire the container status of the container of the application program for which the detection result acquired by the first status acquisition module 41 runs. For example, in the cloud native system, each application program can run in a container to provide different service contents respectively. Therefore, in the embodiments of the present application, in addition to acquiring the execution result of a probe, for example, as the detection result, the second status acquisition module 42 further acquires the status of the container in which the application program runs, so as to facilitate comprehensive judgment.

[0087] The determination module 43 can be used to determine the application programs with a predetermined running status in each group of application programs according to the running status detection results of each application program and the container status of the container in which the application program runs.

[0088] The determination module 43 can comprehensively judge the actual running status of the application program according to the detection result acquired by the first status acquisition module 41 and the container status acquired by the second status acquisition module 42. For example, in actual applications, the user configures the probe so that it can run in the container of the application program to detect the running status of the application program. Therefore, the detection result of the probe can only indirectly reflect the running status of the application program. Moreover, as described above, the accuracy of the detection result of the probe depends on the specific configuration of the maintenance personnel for the probe. Therefore, if the configuration is unreasonable or even incorrect, it will cause the probe to fail to execute in the container. In the prior art, if only such an execution result is used as the sole judgment basis for the running status of the application program, then the processing of the application program error will be caused, for example, restarting the application program when it is actually running normally. The determination module 43 can make a comprehensive judgment based on the detection result and further consider the actual status of the container to eliminate the problem of incorrect detection results caused by the detection scheme configuration problem on the single-machine side.

[0089] Specifically, the determination module 43 can determine the operation threshold according to the total number of application programs with the same identifier that are the acquisition objects of the first status acquisition module 41. In particular, in the solution of the embodiment of the present application, after the user completes the configuration of the probe, it can take effect immediately at all container levels. In the solution of the prior art, this is a potential risk to the reliability of the application. For example, as described above, if the probe configuration is incorrect, adopting the community logic will cause the application to be restarted in full volume, which is a fatal point in terms of stability and high availability of the service. Therefore, in the embodiment of the present application, the first status acquisition module 41 can report the detection execution result on the container side to the determination module 43, and the second status acquisition module 42 can further acquire the container status, so that the determination module 43 can, with a centralized idea, based on the detection results of a group of application programs and the container status, consider the high reliability of the service provided by the overall application from the global perspective of all application programs, especially from the overall perspective of a group of application programs.

[0090] Therefore, the determination module 43 can consider the overall reliability of all application program replicas providing services to determine the quantity threshold for each operation in a group of application programs to prevent the restart of all containers caused by configuration errors.

[0091] For example, the determination module 43 can set the maximum unavailable ratio value in combination with the actual number of containers of the current application for the fallback protection of the overall application program. Specifically, since the number of application programs providing different services is different, when the determination module 43 sets the maximum unavailable ratio, it can consider the total number of current application programs and dynamically adjust the ratio status. For example, when the current number of container replicas is 10, the maximum unavailable ratio value can be set to 20%, then the execution module 44 only allows 2 containers to be triggered for restart at the same time. In actual applications, this threshold can be flexibly set by the maintenance personnel according to the actual application requirements. For example, when the total number of a group of application programs is less than 10, the maximum unavailable quantity value can be fixed at 1, that is, only 1 container is allowed to be triggered for restart at the same time. When the total number of a group of application programs is greater than or equal to 10, the maximum unavailable ratio value can be set to 20%, that is, only 20% of all application programs are allowed to be restarted at the same time. Of course, in the embodiment of the present application, the pdbproducer controller in the open source OpenKruise can also perform real-time dynamic maintenance on this threshold according to the capacity and rule strategy of the service.

[0092] The execution module 44 can be used to perform a predetermined operation on an application program with a predetermined running state.

[0093] Therefore, the execution module 44 may perform a predetermined operation on an application determined to have a predetermined running state based on the result determined by the determination module 43. For example, when the determination module 43 determines that the running state of the application is abnormal or fails, the execution module 44 may restart the application to restore its normal operation.

[0094] If the user makes an error in configuring the detection scheme, when executing the script in the container, the first status acquisition module 41 will inevitably obtain a failed result, but in fact, both the container and the application are running normally. Therefore, the execution module 44 can further count the number of times the application detection result is failed within a predetermined time period, and if the number of failures within a period of time reaches a predetermined threshold, the above-mentioned configuration error may occur. Therefore, in this case, an alarm can be issued to the user without restarting the container.

[0095] The application management device provided by the embodiments of the present application obtains the running state detection results of a group of applications with the same application identifier, and obtains the container status of the container corresponding to the group of applications, so as to make a comprehensive judgment based on the running state detection results of the applications and the container status of the container in which the applications run, to determine the applications with a predetermined running state in the group of applications, and perform operations on the applications determined to have a predetermined running state in this way. Therefore, it is possible to make a comprehensive decision for management by considering the whole of the applications with the same application identifier, avoiding the problem that the management of the applications is performed only based on the running state detection results of a single application, resulting in poor service stability provided by the group of applications due to lack of overall control.

[0096] Embodiment Five

[0097] The internal functions and structures of the application management device are described above, and the device can be implemented as an electronic device. Figure 5 It is a schematic structural diagram of an embodiment of the electronic device provided by the present application. As Figure 5 shown, the electronic device includes a memory 51 and a processor 52.

[0098] The memory 51 is used to store programs. In addition to the above programs, the memory 51 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0099] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0100] The processor 52, not limited to a central processing unit (CPU), may also be a processing chip such as a graphics processing unit (GPU), a field programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. The processor 52 is coupled to the memory 51 and executes the program stored in the memory 51. When the program runs, it executes the application program management method of the second or third embodiment above.

[0101] Furthermore, as Figure 5 shown, the electronic device may further include other components such as a communication component 53, a power supply component 54, an audio component 55, and a display 56. Figure 5 Only some components are schematically shown in Figure 5 and it does not mean that the electronic device only includes

[0102] The components shown.

[0103] The power supply component 54 provides power for various components of the electronic device. The power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.

[0104] The audio component 55 is configured to output and / or input audio signals. For example, the audio component 55 includes a microphone (MIC). When the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals may be further stored in the memory 51 or transmitted via the communication component 53. In some embodiments, the audio component 55 further includes a speaker for outputting audio signals.

[0105] The display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.

[0106] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0107] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An application management method, wherein, Each application runs in a container, and each group of the applications has the same application identifier. The method includes: Obtaining a running state detection result for each application in at least one group of applications; Obtaining the container states of a group of containers in which the applications in the group are respectively running; Determining, based on the running state detection results of the applications and the container states of the containers in which the applications are running, the applications with a predetermined running state in each group of applications; Performing a predetermined operation on the applications with a predetermined running state, including determining an operation threshold according to the total number of applications in a group of applications, where the operation threshold is the maximum number of containers corresponding to restarting the applications in the group allowed at the same time, and managing the group of applications according to the proportion of the applications with a predetermined running state in each group of applications; wherein, the managing of the group of applications further includes: dynamically adjusting and setting the maximum unavailable proportion state based on the current number of application replicas, and performing fallback protection management on the applications by combining the current number of containers and the maximum unavailable proportion.

2. The application management method according to claim 1, wherein, The obtaining a running state detection result for each application in at least one group of applications includes: Executing a predetermined command script in the containers of the applications; Determining the running state detection results of the applications according to the execution results of the command script.

3. The application management method according to claim 1, wherein, The managing the group of applications according to the proportion of the applications with a predetermined running state in each group of applications includes: When the proportion is less than the operation threshold, performing a restart operation on the applications with a predetermined running state.

4. The application management method according to claim 1, wherein, The determining the operation threshold according to the total number of applications in the group of applications includes: When the total number of applications in the group of applications is less than a preset threshold, determining a first preset value as the operation threshold; When the total number of applications in the group of applications is greater than or equal to the preset threshold, determining the product of a second preset value and the total number of applications as the operation threshold.

5. The application management method according to claim 2, wherein, The determining, based on the running state detection results of the applications and the container states of the containers in which the applications are running, the applications with a predetermined running state in each group of applications includes: Determining the applications with a predetermined running state according to the priority of the combination of the running state detection results of the applications and the container states.

6. The application management method according to claim 5, wherein, The running state detection results of the applications include: the command script is executed successfully and the command script is executed failed, and the container states include: the container is ready and the container is unavailable, and The combination includes: a first combination of the command script is executed failed and the container is unavailable; and a second combination of the command script is executed failed and the container is available, and The priority of the first combination is higher than that of the second combination.

7. The application management method according to claim 1, wherein, The performing a predetermined operation on the applications with a predetermined running state includes: Calculating the number of times of the applications determined to have a predetermined running state within a predetermined time period; Performing a predetermined operation on the applications with a predetermined running state according to the number of times.

8. An application management device, wherein, Each application runs in a container, each group of the applications has the same application identifier, and the device includes: A first status acquisition module, configured to acquire a running status detection result for each application in at least one group of applications; A second status acquisition module, configured to acquire the container status of a group of containers in which the group of applications are respectively running; A determination module, configured to determine, according to the running status detection results of the applications and the container status of the containers in which the applications are running, the applications with a predetermined running status in each group of applications; An execution module, configured to perform a predetermined operation on the applications with a predetermined running status, including determining an operation threshold according to the total number of applications in a group of the applications, where the operation threshold is the maximum number of containers corresponding to restarting the group of applications allowed at the same time, and managing the group of applications according to the proportion of the applications with a predetermined running status in each group of applications; wherein, the managing of the group of applications further includes: dynamically adjusting and setting the maximum unavailable proportion status based on the current number of application replicas, and performing fallback protection management on the applications by combining the current number of containers and the maximum unavailable proportion.

9. An electronic device, including: A memory, configured to store a program; A processor, configured to run the program stored in the memory to execute the application management method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon that is executable by a processor, wherein, When the program is executed by the processor, it implements the application management method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Application program operation monitoring method, medium and electronic equipment

    CN109710492A

  • Testing method, manufacturing method and device, medium and electronic equipment

    CN110647470A