Process state recovery method and device, storage medium and electronic equipment
By setting up management processes in the service node and using checkpoint files to restore process status, the problem of insufficient operation stability of visual service applications is solved, and rapid recovery and resource optimization are achieved.
Patent Information
- Application Number
- CN202510506098.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing visual service application operation mode is insufficient instability, resulting in unexpected interruption of service processes, resulting in data loss and resource waste.
By setting up management processes in the service node, listening to node communication ports, receiving visual service requests, and using checkpoint files to restore the status of visual service processes and target task processes when an interrupt is detected.
It realizes accurate recovery of different processes during the service node failure recovery process, reduces recovery time, improves system availability and user experience, and avoids data loss and resource waste.
Smart Images

Figure CN120029828A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the computer field, and in particular, to a method and device for recovering a process state, a storage medium, and an electronic device. Background Art
[0002] A high-performance computing cluster can usually provide powerful high-performance computing services to one or more clients. When the cluster receives a visualization service request, it usually needs to run visualization applications and service applications in the service node at the same time, and needs to operate based on the visualization service protocol.
[0003] However, some visualization applications may take a long time to run, during which the service process may be unexpectedly interrupted due to various hardware failures (such as CPU failure, power failure, etc.) or active requests from the client. If the service process needs to be re-executed from the beginning after the interruption, it will not only waste a lot of time and computing resources, but also may cause data loss during the re-execution process, which is unacceptable. It can be seen that the existing operation mode of visualization service applications has the technical problem of insufficient stability. There is currently no effective solution to this problem.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present application provide a method and device for recovering a process state, a storage medium, and an electronic device, so as to at least solve the technical problem that the existing operation mode of visualization service application has insufficient stability.
[0006] According to one aspect of an embodiment of the present invention, a method for recovering a process state is provided, comprising: receiving a visualization service request sent by a target client to the node communication port during monitoring by a management process on a current service node; in a case where a checkpoint file matching the visualization service request is found, recovering the visualization service process according to first recovery information recorded in the checkpoint file, wherein the visualization service process is used to provide a visualization service interface to the target client; in a case where the target task process is recovered according to second recovery information recorded in the checkpoint file, wherein the target task process is a task process used to execute a target service task in the visualization service interface; in a case where process communication between the visualization service process and the target task process is recovered, providing visualization service to the target client through the visualization service process and the target task process.
[0007] According to another aspect of an embodiment of the present invention, a process status recovery device is also provided, including: a receiving unit, used to receive a visualization service request sent by a target client to the node communication port of the current service node during the process in which the management process monitors the node communication port of the above-mentioned service node; a first recovery unit, used to restore the visualization service process according to the first recovery information recorded in the above-mentioned checkpoint file when a checkpoint file matching the above-mentioned visualization service request is found, wherein the above-mentioned visualization service process is used to provide a visualization service interface to the above-mentioned target client; a second recovery unit, used to restore the target task process according to the second recovery information recorded in the above-mentioned checkpoint file, wherein the above-mentioned target task process is a task process for executing the target service task in the above-mentioned visualization service interface; a service unit, used to provide visualization service to the above-mentioned target client through the above-mentioned visualization service process and the above-mentioned target task process when the process communication between the above-mentioned visualization service process and the above-mentioned target task process is restored.
[0008] Optionally, the above-mentioned first recovery unit is used to: create a first sub-process through the above-mentioned management process; obtain the communication description information included in the above-mentioned first recovery information according to the above-mentioned first sub-process, and perform a communication recovery operation according to the above-mentioned communication description information; obtain the process state description information included in the above-mentioned first recovery information according to the above-mentioned first sub-process, and perform a process state recovery operation on the above-mentioned visualization service process according to the above-mentioned communication description information.
[0009] Optionally, the above-mentioned first recovery unit is used to: determine a first communication description identifier that matches the above-mentioned first sub-process; obtain a first historical communication description identifier that matches the above-mentioned visualization service process from the above-mentioned first recovery information through the above-mentioned first sub-process; and update the above-mentioned first communication description identifier according to the above-mentioned first historical communication description identifier.
[0010] Optionally, the above-mentioned second recovery unit is used to: create a second sub-process through the above-mentioned management process; obtain the communication description information included in the above-mentioned second recovery information according to the above-mentioned second sub-process, and perform communication recovery operations according to the above-mentioned communication description information; obtain the process status description information included in the above-mentioned second recovery information according to the above-mentioned second sub-process, and perform process status recovery operations on the above-mentioned target task process according to the above-mentioned communication description information.
[0011] Optionally, the second recovery unit is used to: determine a second communication description identifier that matches the second sub-process, and establish a process communication connection between the first sub-process and the second sub-process based on the first communication description identifier and the second communication description identifier; obtain a second historical communication description identifier that matches the visualization service process from the second recovery information through the second sub-process; and update the second communication description identifier based on the second historical communication description identifier.
[0012] Optionally, the above-mentioned process state recovery device also includes: a third recovery unit, used for at least one of the following: restoring the memory state corresponding to the above-mentioned current process according to the historical memory state parameters in the above-mentioned process state description information that matches the current process; restoring the CPU state corresponding to the above-mentioned current process according to the historical CPU state parameters in the above-mentioned process state description information that matches the current process; restoring the reference process state corresponding to the above-mentioned current process according to the historical reference state parameters in the above-mentioned process state description information that matches the current process, wherein the above-mentioned reference process state includes at least one of the following: process priority state, device resource state, file handle state; wherein the above-mentioned current process includes the above-mentioned visualization service process and the above-mentioned target task process.
[0013] Optionally, the above-mentioned process status recovery device also includes: an interruption unit, which is used to suspend the above-mentioned target task process when the above-mentioned management process detects a target event, and save the communication description information used to indicate the communication status of the above-mentioned target task process and the process description information used to indicate the process status of the above-mentioned target task process to the above-mentioned second recovery information in the above-mentioned checkpoint file; suspend the above-mentioned visualization service process, and save the communication description information used to indicate the communication status of the above-mentioned visualization service process and the process description information used to indicate the process status of the above-mentioned visualization service process to the above-mentioned first recovery information in the above-mentioned checkpoint file; end the above-mentioned target task process and the above-mentioned visualization service process.
[0014] Optionally, the above-mentioned interruption unit is also used for at least one of the following: when the above-mentioned management process receives a service suspension instruction according to the above-mentioned node communication port, it is determined that the above-mentioned target event is detected; when the above-mentioned management process detects communication interruption indication information according to the above-mentioned node communication port, it is determined that the above-mentioned target event is detected; when the above-mentioned management process detects that the current timestamp is a target timestamp that matches the target inspection period, it is determined that the above-mentioned target event is detected.
[0015] Optionally, the above-mentioned interrupt unit is also used for, when the above-mentioned current service node is the target service node determined by the service cluster scheduling system to respond to the above-mentioned target client, the above-mentioned management process obtains the first visualization service request of the above-mentioned target client according to the allocated above-mentioned node communication port, wherein the above-mentioned service cluster scheduling system includes multiple service nodes, and when the above-mentioned service cluster scheduling system receives the visualization service request of at least one client, the above-mentioned service cluster scheduling system allocates the above-mentioned service nodes corresponding to each of the above-mentioned clients respectively; starts the visualization service process through the above-mentioned management process, and starts the above-mentioned target task process through the above-mentioned management process; when the process communication connection between the above-mentioned visualization service process and the above-mentioned target task process is established, and the service communication connection between the above-mentioned visualization service process and the above-mentioned target client is established, the above-mentioned visualization service process and the above-mentioned target task process are provided to the above-mentioned target client through the above-mentioned visualization service process and the above-mentioned target task process.
[0016] Optionally, the device for recovering the process state further includes: a starting unit, used to start the management process when the current service node is the target node determined by the service cluster scheduling system to respond to the target client; monitor the node communication port allocated to the target client through the management process; and maintain the monitoring state of the management process when the management process does not detect the first visual service request of the target client according to the node communication port.
[0017] Optionally, the startup unit is used to: create a third sub-process through the management process, and obtain at least one visualization request parameter carried in the first visualization service request according to the third sub-process; start the visualization service process according to at least one visualization request parameter; create a fourth sub-process through the management process, and configure the visualization service environment according to the fourth sub-process; and when the visualization service environment is configured, start the target task process according to the fourth sub-process.
[0018] Optionally, the startup unit is further used to: obtain a third communication description identifier matching the node communication port according to the third sub-process, and update the fourth communication description identifier of the third sub-process according to the third communication description identifier; when the update of the fourth communication description identifier of the third sub-process is completed, determine that the service communication connection between the visualization service process and the target client is established; according to the fourth sub-process, obtain the fourth communication description identifier from the configuration information of the visualization service environment; the visualization service process and the target task process establish the process communication connection according to the fourth communication description identifier.
[0019] Optionally, the device for recovering the process state further includes: a reference service unit, used to determine the reference node communication ports respectively allocated to the reference clients when the current service node is the reference service node determined by the service cluster scheduling system to respond to the reference client; start a reference management process for monitoring the reference node communication port; and when the reference management process receives a second visualization service request from the reference client according to the reference node communication port, start a reference visualization service process and a reference task process for responding to the second visualization service request according to the reference management process.
[0020] Optionally, the above-mentioned reference service unit is also used to: create a reference service container that matches the above-mentioned reference node communication port; run a reference management process for listening to the above-mentioned reference node communication port in the above-mentioned reference service container; in the above-mentioned reference service container, start the above-mentioned reference visualization service process and the above-mentioned reference task process for responding to the above-mentioned second visualization service request according to the above-mentioned reference management process.
[0021] Optionally, the reference service unit is also used for at least one of the following: sending a service transfer request to the service cluster scheduling system when the number of the management processes running in the current service node is greater than or equal to a quantity threshold; sending a service transfer request to the service cluster scheduling system when at least one node status parameter corresponding to the current service node meets a parameter condition; wherein the service transfer request is used to request the service cluster scheduling system to redetermine the reference service node from the multiple service nodes.
[0022] Optionally, the device for recovering the process state further includes: a statistical unit, which is used to obtain at least one start timestamp matching the above-mentioned visualization service and a termination timestamp corresponding to each of the at least one start timestamp when the management process receives an end service instruction according to the node communication port, wherein the start timestamp is used to indicate the start running time node of the above-mentioned visualization service process, and the termination timestamp is used to indicate the termination running time node of the above-mentioned visualization service process; determine the target service duration according to the at least one start timestamp and the corresponding termination timestamps; and send service duration prompt information to the target client according to the target service duration.
[0023] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned process state recovery method when running.
[0024] According to another aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above process state recovery method.
[0025] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the process state recovery method through the computer program.
[0026] In an embodiment of the present invention, when the visualization service is interrupted, only the management process is kept in the running state, and the management process occupies very few resources. This design significantly reduces the consumption of system resources, thereby achieving the effect of saving CPU and memory resources and energy saving. In addition, after the client interrupts the visualization service request, if the visualization request is initiated again, the user does not need to perform complex manual recovery operations, but can directly connect to the original node communication port through the client to trigger the visualization application to recover from the checkpoint. This process not only simplifies the user's operating steps, but also ensures that the user can seamlessly access the terminated visualization application. Therefore, the present invention achieves significant improvements in resource optimization and user experience, and solves the problems of complex recovery and high resource occupancy after the visualization service is interrupted in the prior art.
[0027] Furthermore, by recording the first recovery information and the second recovery information in the checkpoint file, which are used to recover the visualization service process and the target task process respectively, accurate recovery of different processes is achieved during the service node fault recovery process. On the one hand, by using the recovery information recorded in the checkpoint file, the visualization service process and the target task process can be quickly located and recovered, which reduces the recovery time and improves the availability of the system; on the other hand, by recovering the process communication between the visualization service process and the target task process, the integrity and consistency of the service are ensured. Furthermore, through the accurate recovery mechanism of the checkpoint file and the rapid recovery of the process communication, the efficiency of the service node in the fault recovery process is significantly improved, and the technical problem of insufficient stability of the operation mode of the visualization service application in the related technology is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0029] Figure 1It is a hardware structure block diagram of a server device of a process state recovery method according to an embodiment of the present application;
[0030] Figure 2 is a flowchart of a method for restoring a process state according to an embodiment of the present application;
[0031] Figure 3 is a schematic diagram of a method for restoring a process state according to an embodiment of the present application;
[0032] Figure 4 is a schematic diagram of another method for restoring a process state according to an embodiment of the present application;
[0033] Figure 5 is a flowchart of another method for restoring a process state according to an embodiment of the present application;
[0034] Figure 6 is a flowchart of another method for restoring a process state according to an embodiment of the present application;
[0035] Figure 7 is a structural schematic diagram of a process state recovery device according to an embodiment of the present application;
[0036] Figure 8 It is a structural diagram of an electronic device for recovering a process state according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0039] The following is an explanation of the technical terms involved in this case:
[0040] Checkpoint-restart: is a process-level fault recovery technology that allows the system to recover quickly after a failure. This technology is particularly suitable for application scenarios that require high availability or fault tolerance.
[0041] X server: It is the display server in the X Window System, which is used to monitor the graphical interface display requests sent by Xclient, and draw and display the graphical interface on the screen.
[0042] X client: X client. Usually various GUI applications, such as Firefox browser, x term, x clock, etc.
[0043] File Descriptor: A file descriptor in the Linux system is an abstract indicator used to represent an open file or other input / output resource. Each open file or resource is assigned a unique non-negative integer, which is used to identify a specific file or resource in a system call.
[0044] Local socket: It can realize communication between different processes on the same host. The link is the socket file. The server creates the socket file and listens to it. The client establishes a socket connection with the socket file path as a parameter to communicate with the server.
[0045] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 1 is a hardware structure block diagram of a server device of a process state recovery method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the figure) a first processor 102 (the first processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a first memory 104 for storing data, wherein the above-mentioned server device may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0046] The first memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as computer programs corresponding to the data processing method of the memory in the embodiment of the present application. The first processor 102 executes various functional applications and data processing by running the computer program stored in the first memory 104, that is, implementing the above method. The first memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the first memory 104 may further include a memory remotely arranged relative to the first processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0047] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0048] As an optional implementation, Figure 2 As shown, the above process status recovery method can be applied to a service node in a service cluster, including:
[0049] S202, receiving a visualization service request sent by a target client to the node communication port of the current service node during the process of the management process monitoring the node communication port of the current service node;
[0050] S204, when a checkpoint file matching the visualization service request is found, restore the visualization service process according to the first restoration information recorded in the checkpoint file, wherein the visualization service process is used to provide a visualization service interface to the target client;
[0051] S206, restoring the target task process according to the second restoration information recorded in the checkpoint file, wherein the target task process is a task process used to execute the target service task in the visual service interface;
[0052] S208 , when the process communication between the visualization service process and the target task process is restored, providing visualization service to the target client through the visualization service process and the target task process.
[0053] It should be noted that the above implementation of the present application can be applied to the current service node in the service cluster. The above current service node may refer to a service node assigned to process a specific client request in a distributed system or cluster environment. Optionally, the above current service node may be a physical server or a virtual machine with independent computing resources and communication capabilities.
[0054] Furthermore, in the above step S202, the above management process may be a process in the current service node responsible for monitoring and managing the running status of the service node, and may be used to monitor the node communication port, receive client requests, and suspend or resume the process when necessary.
[0055] It can be understood that the above visualization service request can be a request sent by the target client to request the current service node to provide a visualization service interface. This request may include user operation instructions, data query requests or other instructions related to visualization services. In an optional manner, the above visualization service request can be specifically a service request based on transmission using the X protocol or the vnc protocol, which is used to request a computing node (i.e., a specific service node) in a high-performance computing cluster to perform a visualization job.
[0056] The implementation method of the present application can be applied to any service node in a high-performance computing cluster. The client sends a visualization service request to request the computing node (service node) in the cluster to perform a visualization job. When the cluster receives the visualization service request, it is usually necessary to run the visualization application and the service application in the service node at the same time and operate based on the visualization service protocol.
[0057] However, some visualization applications may take a long time to run, during which the service process may be unexpectedly interrupted due to various hardware failures (such as CPU failure, power failure, etc.) or active requests from the client. If the service process needs to be re-executed from the beginning after the interruption, it will not only waste a lot of time and computing resources, but also may cause data loss during the re-execution process, which is unacceptable. It can be seen that the existing operation mode of visualization service applications has the technical problem of insufficient stability.
[0058] In response to the above technical problems, in the above implementation of the present application, a management process responsible for monitoring and managing the running status of the service node can be kept running in the current service node, and then when the application processes of the above-mentioned visualization application and service application are closed, the process status of the application processes of the visualization application and service application in the current service node can be restored through the above-mentioned management process.
[0059] During the recovery process, the above steps S204 to S208 may be further included. The following is an explanation of the recovery method of the above process status.
[0060] In the above step S204, the checkpoint file may be a file storing process status information, which is used to save key data when the process is suspended for subsequent recovery. The above checkpoint file may include recovery information, such as communication description information, process status description information, etc. In this implementation, the specific information type included in the above checkpoint file is not limited.
[0061] In an embodiment of the present application, the above-mentioned checkpoint file may include first recovery information and second recovery information. Among them, the first recovery information can be used to indicate the recovery data related to the visualization service process in the checkpoint file. It includes communication description information and process status description information, which are used to restore the state of the visualization service process. The second recovery information can be used to refer to the recovery data related to the target task process in the checkpoint file. It also includes communication description information and process status description information, which are used to restore the state of the target task process.
[0062] It should be noted that the visualization service process can be a process specifically used to provide a visualization service interface to the target client. It is responsible for generating and managing the visualization interface and interacting with the target client. In a specific manner, the above-mentioned visualization service process can be an Xvnc process. It should be noted that in the implementation mode of the present application, Xvnc is a VNC (Virtual Network Computing) server based on the X Window System, which allows users to remotely access and operate a graphical interface running on another computer through a network. Xvnc combines the functions of X Server and VNC Server, and acts as an X Server to provide display services for applications, and as a VNC Server to provide desktop access for remote clients.
[0063] Further, the target task process can be a process in the visual service interface for executing a specific service task. It is responsible for processing the request of the target client and feeding back the result to the visual service process. In the implementation mode of the present application, the above target task process can match the above visual service request.
[0064] For example, in the case where the visualization service request is specifically a text editing request, the target task process may be a Gedit process. Gedit is a lightweight text editor suitable for processing text file editing tasks. When the target client sends a text editing request, the current service node may start the Gedit process as the target task process. The Gedit process receives the text content input by the user sent by the target client, performs text editing operations (such as inserting, deleting, searching and replacing, etc.), and saves the edited results to a local file. Subsequently, the Gedit process feeds back the edited text content to the visualization service process, and the visualization service process returns the edited text content to the target client through network communication for display.
[0065] For another example, when the visualization service request is specifically an image processing request, the target task process may be a GIMP (GNU Image Manipulation Program) process. When the current service node receives an image processing request sent by the target client, the current service node may start the GIMP process as the target task process. The GIMP process receives the user-uploaded image file sent by the target client, performs image processing operations (such as cropping, adjusting colors, applying filters, etc.), and saves the processed image as a new file. Subsequently, the GIMP process feeds back the processed image to the visualization service process, and the visualization service process then returns the processed image to the target client through network communication for display.
[0066] For another example, when the visualization service request is specifically a data analysis request, the target task process can be an R language process. The R language process can be used to handle complex data analysis tasks. When the target client sends a data analysis request, the system starts the R language process as the target task process. The R language process receives the data files uploaded by the user (such as CSV files), performs operations such as data cleaning, statistical analysis or visualization, and saves the analysis results (such as statistical charts or analysis reports) as output files. Subsequently, the R language process feeds back the analysis results to the visualization service process, and the visualization service process displays the results to the user through a visualization interface.
[0067] The correspondence between the target task process and the visualization service request is merely an exemplary description, and the specific types and correspondences between the target task process and the visualization service request are not limited in the implementation of the present application.
[0068] Furthermore, in the above step S208, the management process can establish a communication connection between the visualization service process and the target task process, and when the visualization service process establishes a connection with the target client through the restored communication state, it starts to provide a visualization service interface to the target client.
[0069] The following is a description of the further implementation of the above process status recovery method. Figure 3 and Figure 4 , describing the process running status in various service nodes before and after the process status is restored.
[0070] like Figure 4 As shown in the figure, before the visualization service is interrupted, there may be multiple processes running in the system, including management process, Xvnc process and target task process. These processes communicate through port 5901 and jointly provide visualization services to the target client (vcn client).
[0071] Then as Figure 3 As shown in the figure, when the visualization service is interrupted (for example, the vcn client actively exits), other processes except the management process are terminated. Before the service is interrupted, the system will record the first recovery information and the second recovery information in the checkpoint file. The first recovery information is used to restore the visualization service process (such as the Xvnc process), and the second recovery information is used to restore the target task process. Due to its lightweight design, the management process occupies very few system resources, such as CPU and memory, thereby achieving resource conservation and energy saving effects.
[0072] Furthermore, after the client interrupts the visualization service request, if a visualization request is initiated again, the client can directly connect to the original node communication port (port 5901) to trigger the visualization application to recover from the checkpoint. Figure 4 The figure shows the system structure after recovery. The management process continues to listen to port 5901, and the VNC client connects through port 5901, triggering the above recovery process. When the communication between the management process, Xvnc process and the target task process is restored, the process status of the Xvnc process and the target task process can be restored to the state before the interruption according to the checkpoint file.
[0073] In the above-mentioned embodiment of the present application, when the visualization service is interrupted, only the management process is kept in the running state, and the management process occupies very few resources. This design significantly reduces the consumption of system resources, thereby achieving the effect of saving CPU and memory resources and energy saving. In addition, after the client interrupts the visualization service request, if the visualization request is initiated again, the user does not need to perform complex manual recovery operations, but can directly connect to the original node communication port through the client to trigger the visualization application to recover from the checkpoint. This process not only simplifies the user's operating steps, but also ensures that the user can seamlessly access the terminated visualization application. Therefore, the present invention achieves significant improvements in both resource optimization and user experience, and solves the problem of complex recovery and high resource occupancy after the visualization service is interrupted in the prior art.
[0074] Furthermore, by recording the first recovery information and the second recovery information in the checkpoint file, which are used to recover the visualization service process and the target task process respectively, accurate recovery of different processes is achieved during the service node fault recovery process. On the one hand, by using the recovery information recorded in the checkpoint file, the visualization service process and the target task process can be quickly located and recovered, which reduces the recovery time and improves the availability of the system; on the other hand, by recovering the process communication between the visualization service process and the target task process, the integrity and consistency of the service are ensured. Furthermore, through the accurate recovery mechanism of the checkpoint file and the rapid recovery of the process communication, the efficiency of the service node in the fault recovery process is significantly improved, and the technical problem of insufficient stability of the operation mode of the visualization service application in the related technology is solved.
[0075] The following describes the specific method for restoring the process status in the above steps S204 to S208. First, the following describes the process of restoring the visualization service process based on the first restoration information.
[0076] In an optional implementation manner, the above-mentioned restoring the visualization service process according to the first restoration information recorded in the checkpoint file includes:
[0077] S1, create the first child process through the management process;
[0078] S2, acquiring communication description information included in the first recovery information according to the first sub-process, and performing a communication recovery operation according to the communication description information;
[0079] S3: acquiring process state description information included in the first recovery information according to the first sub-process, and performing a process state recovery operation on the visualization service process according to the communication description information.
[0080] In the above implementation of the present application, a first sub-process may be created through a management process, and the visualization service process may be restored through the first sub-process. The restoration process may be isolated using the sub-process to avoid affecting the management process itself.
[0081] In an optional implementation manner, the acquiring, according to the first sub-process, the communication description information included in the first recovery information, and performing the communication recovery operation according to the communication description information, includes:
[0082] S1, determining a first communication description identifier matching a first subprocess;
[0083] S2, obtaining, through the first subprocess, from the first recovery information, a first historical communication description identifier that matches the visualization service process;
[0084] S3: Update the first communication description identifier according to the first historical communication description identifier.
[0085] In the above implementation, in the first sub-process, a communication description identifier matching the sub-process may be determined. The above communication description identifier may be used to identify and distinguish communications between different processes.
[0086] Then, through the first sub-process, a historical communication description identifier matching the visualization service process is obtained from the first recovery information. The first recovery information is a previously saved checkpoint file, which contains communication status information when the process executes the checkpoint. Further, according to the obtained first historical communication description identifier, the first communication description identifier in the first sub-process is updated. The above-mentioned updating method can be to copy or map the information in the historical communication description identifier to the new communication description identifier.
[0087] In the above implementation of the present application, the communication between processes is restored by obtaining and using the communication description information through the first subprocess. By determining, obtaining and updating the communication description identifier, it can be ensured that the communication connection can be correctly rebuilt after the process is restored, thereby restoring the integrity and consistency of the service. This method not only improves the availability and stability of the system, but also simplifies the complexity of process recovery, making the process of process recovery more efficient and reliable.
[0088] In an optional implementation, when the current process is a visualization service process, the above-mentioned process state recovery operation includes at least one of the following:
[0089] Method 1: restore the memory state corresponding to the current process according to the historical memory state parameters in the process state description information matching the current process;
[0090] Optionally, the above-mentioned memory state may include one or more of a program counter, register value, stack content, and allocated memory area corresponding to the visualization service process. Restoring the memory state by the above-mentioned method 1 can ensure that the visualization service process can return to the working state before the interruption, including all data and variables in the memory.
[0091] Method 2: restore the CPU state corresponding to the current process according to the historical CPU state parameters in the process state description information matching the current process;
[0092] Optionally, the CPU state may include but is not limited to general registers, floating point registers, program counters, flag registers, etc. By performing CPU state recovery in the second method, it can be ensured that parameters reflecting the execution state of the visualization service process, such as CPU registers and program counters, are restored to the state before the interruption.
[0093] Method three: restore the reference process state corresponding to the current process according to the historical reference state parameters in the process state description information matching the current process, wherein the reference process state includes at least one of the following: process priority state, device resource state, and file handle state.
[0094] Optionally, the process priority status can be used to indicate the importance level of the process in the operating system scheduling. The device resource status can be used to indicate the device resources held by the process, such as the status of file descriptors, device locks, etc. The file handle status can be used to indicate the files and network connections opened by the process, so as to ensure that these resources are still valid and accessible when restored.
[0095] In the above-mentioned implementation manner of the present application, the purpose of the above-mentioned at least one recovery operation is to restore the visualization service process to its previous state as quickly as possible after the interruption of the visualization service process, thereby reducing the impact of the interruption on the user and maintaining the continuity of the service.
[0096] In an optional implementation, the management process is specifically a pre-processed BLCR (Berkeley Lab Checkpoint / Restart) process, wherein the pre-processing method may be to modify the BLCR with reference to the implementation method of the super daemon process xinetd so that the BLCR also has the function of a super daemon process.
[0097] The following describes an implementation method of restoring the process state of the Xvnc process based on BLCR as the management process, the visualization service process as the Xvnc process, and the target task process as the gedit process.
[0098] S1, the user visits the application again;
[0099] Specifically, after saving the application data and ending the application, the user accesses the host's port 5901 through the vnc client again. This is the trigger point for the user to restart the session.
[0100] S2, BLCR process handles the client connection request;
[0101] When the vnc client connects to port 5901, the BLCR process receives the connection request. Since the Xvnc and gedit processes have ended, the BLCR process needs to restart these processes to restore the user's session.
[0102] S3, check the Checkpoint file;
[0103] The BLCR process checks whether a checkpoint file has been created for Xvnc and gedit. If a checkpoint file has been created, the checkpoint file is used to restore the process status; if a checkpoint file has not been created, the Xvnc and gedit processes can be started from scratch based on the BLCR process.
[0104] S4, create the first child process;
[0105] The BLCR process creates the first child process through fork, which is used to restore the Xvnc and gedit processes. In this step, using the child process can isolate the recovery process to avoid affecting the BLCR process itself.
[0106] S5, create a local socket
[0107] A local socket for communication between Xvnc and gedit is created in the first child process. For example, a local socket (such as / tmp / .X11-unix / X0) for communication between Xvnc and gedit can be created in the child process. Assume that the file descriptor of the created local socket is fda (for example, it can be a specific first communication description identifier).
[0108] S6, copy file descriptor;
[0109] In the above step S6, the file descriptor can be a specific communication description information. Specifically, the dup2() function can be used to copy the file descriptor of the socket bound to port 5901 by the BLCR process to the standard input, output and error file descriptors of the first child process, thereby passing the client's connection request to the first child process.
[0110] Next, the file descriptor of the local socket for X protocol communication opened by the Xvnc process when the checkpoint is executed is obtained from the Xvnc checkpoint file, and copied to the first child process, thereby restoring the communication between Xvnc and gedit.
[0111] Specifically, the dup2() function is called three times in the first child process to copy the file descriptor of the socket currently bound to port 5901 of the BLCR process to the file descriptors 0 (standard input), 1 (standard output), and 2 (standard error) of the first child process respectively; the checkpoint file will record the information of the file descriptors that the process has opened when the checkpoint is executed, and the file descriptor of the local socket opened by the Xvnc process for X protocol communication when the checkpoint is executed is obtained from the Xvnc checkpoint file, assuming it is fdb (that is, a specific first historical communication description identifier), and the dup2(fda, fdb) function is called in the first child process to copy the created local socket file descriptor fda to the file descriptor fdb of the child process, so that the local socket file descriptor in the first child process is restored. This socket is used for communication between Xvnc and gedit, and the first child process can subsequently communicate with the gedit process through this file descriptor.
[0112] S7, restore the Xvnc process status; restore the CPU status and memory data of the Xvnc process from the Xvnc checkpoint file. This is to restore the Xvnc process to the state when the checkpoint was executed.
[0113] Through the above implementation of the present application, the checkpoint file can be managed and the process status can be restored through the BLCR process. This solution can effectively manage and restore the running status of the visualization service application without sacrificing the user experience.
[0114] The following describes a process of restoring the target task process based on the second restoration information.
[0115] In an optional implementation manner, the above-mentioned recovering the target task process according to the second recovery information recorded in the checkpoint file includes:
[0116] S1, create the second child process through the management process;
[0117] S2, acquiring communication description information included in the second recovery information according to the second sub-process, and performing a communication recovery operation according to the communication description information;
[0118] S3, obtaining the process state description information included in the second recovery information according to the second sub-process, and performing a process state recovery operation on the target task process according to the communication description information.
[0119] In the above implementation of the present application, a second sub-process can be further created through the management process, and the target task process can be restored through the second sub-process. The sub-process can then be used to isolate the restoration process to avoid affecting the management process itself.
[0120] In an optional implementation manner, the acquiring, according to the second sub-process, the communication description information included in the second recovery information, and performing the communication recovery operation according to the communication description information, includes:
[0121] S1, determining a second communication description identifier that matches the second sub-process, and establishing a process communication connection between the first sub-process and the second sub-process according to the first communication description identifier and the second communication description identifier;
[0122] S2, obtaining, through the second subprocess, from the second recovery information, a second historical communication description identifier that matches the target task process;
[0123] S3: Update the second communication description identifier according to the second historical communication description identifier.
[0124] In the above implementation, in the second sub-process, a communication description identifier matching the sub-process may be determined. The above communication description identifier may be used to identify and distinguish communications between different processes.
[0125] Then, through the second sub-process, the historical communication description identifier that matches the target task process is obtained from the second recovery information. The second recovery information is a previously saved checkpoint file, which contains the communication status information when the process executes the checkpoint. Further, according to the obtained second historical communication description identifier, the second communication description identifier in the second sub-process is updated. The above-mentioned updating method can be to copy or map the information in the historical communication description identifier to the new communication description identifier.
[0126] In an optional implementation, when the current process is the target task process, the above-mentioned process state recovery operation includes at least one of the following:
[0127] Method 1: restore the memory state corresponding to the current process according to the historical memory state parameters in the process state description information matching the current process;
[0128] Optionally, the above-mentioned memory state may include one or more of the program counter, register value, stack content, and allocated memory area corresponding to the visualization service process. Restoring the memory state by the above-mentioned method 1 can ensure that the target task process can return to the working state before the interruption, including all data and variables in the memory.
[0129] Method 2: restore the CPU state corresponding to the current process according to the historical CPU state parameters in the process state description information matching the current process;
[0130] Optionally, the CPU state may include but is not limited to general registers, floating point registers, program counters, flag registers, etc. By performing CPU state recovery in the second method, it can be ensured that parameters such as CPU registers and program counters reflecting the execution state of the target task process are restored to the state before the interruption.
[0131] Method three: restore the reference process state corresponding to the current process according to the historical reference state parameters in the process state description information matching the current process, wherein the reference process state includes at least one of the following: process priority state, device resource state, and file handle state.
[0132] Optionally, the process priority status can be used to indicate the importance level of the process in the operating system scheduling. The device resource status can be used to indicate the device resources held by the process, such as the status of file descriptors, device locks, etc. The file handle status can be used to indicate the files and network connections opened by the target task process, so as to ensure that these resources are still valid and accessible when restored.
[0133] The following continues to use BLCR as a management process, the visualization service process is specifically an Xvnc process, and the target task process is specifically a gedit process to explain the implementation method of restoring the process state of the Xvnc process.
[0134] S1, create a second child process; the BLCR process creates a new child process again through fork, which is called the second child process. In this step, the second child process will be used exclusively to restore the state of the gedit process to isolate the recovery process and avoid affecting other processes.
[0135] S2, create a socket connection; create a new socket connection in the second child process and connect it to the local socket created in step a. Through this step, a communication channel is established between the second child process and the visualization service process (Xvnc) to restore the interaction between the processes. Assume that the file descriptor of the newly created socket is fdc (i.e., a specific second communication description identifier).
[0136] S3, obtain the file descriptor of gedit; obtain the file descriptor of the local socket for X protocol communication opened by the gedit process when executing the checkpoint from the gedit checkpoint file, assuming it is fdd (that is, a specific second history communication description identifier).
[0137] S4, copy the file descriptor; call the dup2(fdc, fdd) function in the second child process to copy the file descriptor fdc to the file descriptor fdd. Through this step, the new socket file descriptor fdc in the second child process can be mapped to the file descriptor fdd of the gedit process at the time of checkpoint, thereby restoring the communication state of the gedit process.
[0138] It should be noted that the dup2 function is used to copy the file descriptor and replace the old file descriptor with the new one.
[0139] S5, restore the gedit process status; specifically, the CPU status and memory data of the gedit process may be restored from the gedit checkpoint file.
[0140] Through the above steps, the BLCR process can restore the state of the gedit process by creating a child process and using the checkpoint file. This process not only restores the communication state of the process, but also restores the computing state and memory data of the process, thereby achieving complete recovery of the process.
[0141] The following combination Figure 5 A complete process status recovery procedure is described.
[0142] S502, creating a first child process;
[0143] S504, creating a local socket in the first child process;
[0144] S506, calling the dup2 function in the first child process to copy the socket of port 5901;
[0145] S508, calling the dup2 function in the first child process to copy the local socket;
[0146] S510, restoring the CPU memory data of the first child process;
[0147] S512, creating a second child process;
[0148] S514, creating a local socket connection in the second child process;
[0149] S516, calling the dup2 function in the second child process to copy the local socket;
[0150] S518, restore the CPU memory data of the second child process.
[0151] It should be noted that the specific implementation of the above steps can be carried out in the manner described in the above embodiment. Through the above embodiment, the recovery of the Xvnc and gedit processes is completed, and the connection between the vnc client and Xvnc and the connection between Xvnc and gedit are restored.
[0152] On the one hand, after the user suspends the use of the application, he can directly use the vnc client to connect to the original port to access the application again. In other words, for users, when accessing a suspended application, there is no need to manually resume the application. Using the vnc client to directly connect to the vnc port can trigger the application to resume from the checkpoint, which is very convenient to use.
[0153] On the other hand, from the perspective of fault recovery, the job process may fail to exit due to various software and hardware reasons during operation, which may lead to loss of user data and cause serious consequences. By adopting the above implementation method, checkpoints can be created for the user's job process at regular intervals. When the job process exits abnormally, when the user connects to vnc, the job application can be automatically restored to the state of the most recent checkpoint, so as to minimize the user's loss.
[0154] The following describes the pre-steps for triggering the above-mentioned process status recovery method.
[0155] In an optional implementation manner, before receiving the visualization service request sent by the target client to the node communication port, the method further includes:
[0156] S1, when the management process detects a target event, suspending the target task process, and saving communication description information for indicating the communication state of the target task process and process description information for indicating the process state of the target task process to second recovery information in the checkpoint file;
[0157] S2, pausing the visualization service process, and saving the communication description information used to indicate the communication state of the visualization service process and the process description information used to indicate the process state of the visualization service process to the first recovery information in the checkpoint file;
[0158] S3, end the target task process and visualization service process.
[0159] It is understandable that in the above implementation of the present application, when a target event is detected, the target task process can be firstly suspended, and the state record of the target task process can be completed; then the visualization service process can be suspended, and the state record of the visualization service process can be completed. Finally, the target task process and the visualization service process can be terminated.
[0160] It should be noted that the above target event may include at least one of the following:
[0161] Method 1: When the management process receives a service suspension instruction according to the node communication port, it is determined that the target event is detected;
[0162] In the above-mentioned method 1, when the management process receives a clear service suspension instruction through the node communication port, it is determined that the target event has been detected. For example, the user can send a suspension instruction to the communication port of the service node to the current service node through the management interface. The above suspension instruction may include but is not limited to the exit service operation, the close service operation, the suspension service operation, etc. received by the client. After receiving the suspension instruction, the management process recognizes it as a target event and triggers the saving of the process state and the suspension of the process.
[0163] Method 2: when the management process detects communication interruption indication information according to the node communication port, it is determined that the target event is detected;
[0164] In the above-mentioned method 2, when the management process detects the indication information of communication interruption through the node communication port, it is determined that the target event has been detected. Specifically, in a distributed system, if a service node loses connection with other nodes due to network problems or node failures, the management process on the node may detect the communication interruption. At this time, the management process will believe that the target event has occurred, and then suspend the relevant processes, save their states, and prepare to restore or shut down.
[0165] Mode three: when the management process detects that the current timestamp is a target timestamp that matches the target inspection period, it is determined that the target event is detected.
[0166] In the third method above, when the management process detects that the current timestamp matches the preset target check period, it is determined that the target event has been detected. Specifically, a regular checkpoint operation can be set to save the process status. For example, the status of all key processes can be automatically saved and a checkpoint can be created every 10 minutes.
[0167] It should be noted that, when the target event corresponding to the third method is detected, only the operation of creating a checkpoint may be performed without executing the operations of temporarily executing the process and terminating the process.
[0168] The following further describes how to detect a target event and create a checkpoint.
[0169] In this implementation, the BLCR (Berkeley Lab Checkpoint / Restart) tool can be used to perform checkpoint (save checkpoints) and restore (recovery) operations on the user's application (such as the gedit text editor) to ensure that the user can save the application state when needed and restore the application at a subsequent time point for continued use.
[0170] S1, user requests to save application data; the user sends a request to the high-performance computing (HPC) cluster scheduling system through the client, requesting to save the data of the gedit application currently in use, and can choose to automatically end the application after the saving is completed.
[0171] S2, the HPC cluster scheduling system executes the BLCR command; after receiving the user's request, the HPC cluster scheduling system uses the BLCR tool to perform checkpoint operations on the gedit and Xvnc processes. BLCR first suspends the gedit process and saves gedit's CPU status, memory data, and file descriptor information to a local file (called a gedit checkpoint file). Then, BLCR suspends the Xvnc process and saves Xvnc's CPU status, memory data, and file descriptor information to a local file (called an Xvnc checkpoint file).
[0172] S3, decides whether to terminate the process based on the user's choice; if the user chooses to exit automatically after saving, the Xvnc and gedit processes will be terminated after saving to the checkpoint file. If the user does not choose to exit automatically, the Xvnc and gedit processes will resume after saving to the checkpoint file.
[0173] S4, the BLCR process continues to listen to the port.
[0174] It is understandable that no matter whether the Xvnc and gedit processes are terminated or not, the BLCR process continues to listen to port 5901 so as to be able to respond quickly when the user requests to access the application again.
[0175] In the above implementation of the present application, using BLCR for checkpoint and restore operations provides users with a reliable way to save and restore application states, especially in HPC environments, where this capability is particularly important for preventing data loss and improving resource utilization. Through this process, users can pause applications when needed, save states, and resume applications at the appropriate time to continue previous work without worrying about data loss or resource waste.
[0176] The following further describes the process of establishing the visualization service before the above process recovery operation.
[0177] In an optional implementation manner, before suspending the target task process when the management process detects the target event, the method further includes:
[0178] S1, when the current service node is a target service node determined by the service cluster scheduling system to respond to the target client, the management process obtains a first visualization service request of the target client according to the assigned node communication port, wherein the service cluster scheduling system includes multiple service nodes, and when the service cluster scheduling system receives the visualization service request of at least one client, it allocates the corresponding service nodes to the at least one client respectively;
[0179] S2, start the visualization service process through the management process, and start the target task process through the management process;
[0180] S3, when the process communication connection between the visualization service process and the target task process is established, and the service communication connection between the visualization service process and the target client is established, a visualization service matching the first visualization service request is provided to the target client through the visualization service process and the target task process.
[0181] In the above implementation of the present application, the service cluster scheduling system can be used to indicate a system responsible for managing and scheduling multiple service nodes in a service cluster. The main function of the system is to dynamically allocate service nodes to process these requests based on client requests, thereby achieving load balancing and improving the availability and responsiveness of the system. The service cluster scheduling system is a key component in a distributed system, which ensures high availability and scalability of services.
[0182] In an optional implementation, the service cluster scheduling system may be configured to include the following functions:
[0183] Service node management function: Track the status of all service nodes in the cluster, including the health status and load status of the nodes.
[0184] Request scheduling function: Receive client requests and assign them to appropriate service nodes based on certain strategies (such as polling, minimum number of connections, load balancing, etc.).
[0185] Load balancing function: According to the current load of the service node, the request is reasonably distributed to avoid some nodes being overloaded while other nodes are idle.
[0186] Failover function: When a service node fails, it can automatically transfer requests to other healthy nodes to ensure service continuity.
[0187] Scaling management function: Dynamically increase or decrease service nodes according to the system load to adapt to load changes.
[0188] In the above implementation of the present application, when the service cluster scheduling system receives a visualization service request from a target client, it will allocate a suitable service node to the client according to the status and load of the service nodes in the current cluster.
[0189] The management process runs on the assigned service node and listens to the communication port assigned to the node to receive client requests. Once the client request is received, the management process will start the visualization service process and the target task process to handle the request. The visualization service process is responsible for establishing a communication connection with the client and providing a user interface; the target task process is responsible for executing specific service tasks.
[0190] When the communication connection between the visualization service process and the target task process is established, the two processes will work together to provide the required visualization services to the target client. The visualization service process forwards the client's interaction request to the target task process and displays the processing results of the target task process to the client.
[0191] Through the above implementation of the present application, the role of the service cluster scheduling system is to ensure that the client's request can be processed efficiently and reliably. By dynamically allocating service nodes, the system can achieve load balancing, improve resource utilization, and ensure high availability and scalability of services. In addition, the service cluster scheduling system can also achieve failover when a service node fails to ensure service continuity.
[0192] The following further describes in detail the process of establishing the visualization service before the above process recovery operation.
[0193] In an optional implementation manner, before the management process obtains the first visualization service request of the target client according to the allocated node communication port, the method further includes:
[0194] S1, when the current service node is the target node determined by the service cluster scheduling system to respond to the target client, start the management process;
[0195] S2, monitors the node communication port allocated to the target client through the management process;
[0196] S3: When the management process does not detect the first visualization service request of the target client according to the node communication port, the management process is kept in a listening state.
[0197] In the above implementation of the present application, the management process can be started first, but the subprocess is not created immediately. Instead, the subprocess is created after a vnc client is connected, and then the Xvnc and visualization application subprocesses are started.
[0198] In an optional implementation, the above-mentioned starting the visualization service process through the management process and starting the target task process through the management process includes:
[0199] S1, creating a third sub-process through a management process, and obtaining at least one visualization request parameter carried in a first visualization service request according to the third sub-process;
[0200] S2, starting a visualization service process according to at least one visualization request parameter;
[0201] S3, creating a fourth sub-process through the management process, and configuring a visualization service environment according to the fourth sub-process;
[0202] S4, when the visualization service environment configuration is completed, start the target task process according to the fourth sub-process.
[0203] In the above implementation of the present application, the third sub-process and the fourth sub-process may be created respectively through the management process, so as to further create and start the visualization service process and the target service process respectively.
[0204] In an optional implementation, before providing the visualization service matching the first visualization service request to the target client through the visualization service process and the target task process, the method further includes:
[0205] S1, obtaining a third communication description identifier matching the node communication port according to the third sub-process, and updating a fourth communication description identifier of the third sub-process according to the third communication description identifier;
[0206] S2, when the fourth communication description identifier of the third sub-process is updated, determining that the service communication connection between the visualization service process and the target client is established;
[0207] S3, according to the fourth sub-process, obtaining a fourth communication description identifier from the configuration information of the visualization service environment;
[0208] S4: The visualization service process and the target task process establish a process communication connection according to the fourth communication description identifier.
[0209] In the above implementation mode of the present application, the above communication description identifier may specifically be a file descriptor of a socket.
[0210] The following combination Figure 6 The specific implementation method of creating a visualization service at the current service node is described.
[0211] S602, configure the super daemon function of BLCR;
[0212] In the process of configuring BLCR, you can refer to the implementation of the super daemon xinetd and modify BLCR so that BLCR also has the function of a super daemon. Use BLCR to start Xvnc and the application that the user requests to access. On the one hand, in the process of running other processes through BLCR, you can provide the command cr_run, and the parameter of the cr_run command is the application to be started. Assuming that the application that the user requests to access is gedit, that is, you can execute cr_run gedit to run gedit through BLCR.
[0213] On the other hand, since the requested target application needs to be accessed through vnc during the visualization service process, in addition to starting the requested application, Xvnc also needs to be started (both are started through BLCR), and when starting Xvnc, some vnc-related parameters are required, such as vnc port number, vnc password, etc.
[0214] Therefore, it is necessary to further extend two parameters for the cr_run command, the vnc port number (that is, the port number assigned by the system) and the vnc password. When BLCR starts Xvnc later, these two parameters will be used to start the vnc process.
[0215] S604, processing the client request;
[0216] In the application implementation method, the user submits a request to the HPC scheduling system (i.e., a specific service cluster scheduling system) to use a visualization application. The following takes gedit as an example, assuming that the visualization application requested by the user is gedit, which is a graphical text editor program in the Linux desktop environment.
[0217] The HPC cluster scheduling system selects a suitable computing node to run the graphical application, and randomly allocates an idle port on the computing node host as the port for providing external VNC services. Assume that the port allocated is 5901 on the host.
[0218] Later, the user who submitted the request will use the vnc client to connect to this port of the host to access the running gedit program. For different requests submitted by different users, different hosts or different ports on the same host will be used, so that the vnc ports used by visualization applications requested by different users will not conflict.
[0219] S606, the BLCR process monitors the port and creates a child process according to the monitoring situation;
[0220] After executing the cr_run command, the BLCR process will be started. The process is bound to the vnc port passed in S604 and listens to the port, but Xvnc and gedit are not started immediately until a user uses a vnc client to connect to the port. The role of BLCR here is equivalent to a super daemon process such as xinetd. After it is started, it keeps listening to the port and enables the corresponding application when a client accesses the port. It can be understood that the port listening function in this step can be an application function configured for the BLCR extension. The HPC cluster management system returns the assigned computing node host IP, vnc port number (i.e., the previously assigned port 5901) and vnc password (if a vnc password is used) to the user who submitted the request.
[0221] S608, setting environment variables and starting the process; S610, configuring inter-process communication and client network access communication.
[0222] In the above implementation, the user can use a vnc client to connect to port 5901 of the host;
[0223] Due to the above step S606, the BLCR process has been listening on port 5901. When the vnc client connects to the port, the BLCR process will receive the connection. BLCR generates a child process (the third child process) through fork, and calls the dup2() function three times in the child process to copy the file descriptor of the socket currently bound to port 5901 of the BLCR process to the file descriptors 0 (standard input), 1 (standard output), and 2 (standard error) of the child process, that is, execute dup2(socket0,0), dup2(socket0,1), and dup2(socket0,2). Here, socket0 represents the file descriptor of the socket listening on port 5901. In this way, the file descriptors 0, 1, and 2 of the child process are associated with the file descriptors of the socket of port 5901, and the reading and writing of file descriptors 0, 1, and 2 by the child process are the reading and writing of the socket.
[0224] Furthermore, the parameters received in the above steps are used in the child process to start the Xvnc program through the execve function. In this way, Xvnc can communicate with the vnc client through file descriptors 0, 1, and 2.
[0225] Next, after BLCR starts the child process and starts Xvnc in the child process, the BLCR process executes fork again, starts the second child process (the fourth child process), and executes the export command in the fourth child process to set the Linux operating system environment variables, such as export DISPLAY=:1. The DISPLAY environment variable format is as follows host:NumA.NumB
[0226] The host parameter is used to indicate which host to project to. If it is the local host, host is empty. For other hosts, fill in the corresponding IP address. When projecting to the local host, the NumA parameter is used to indicate the path of the unix socket. If it is 0, it means connecting to the local socket file / tmp / .X11-unix / X0. When projecting to the remote end, it means the value of the port minus 5900. If NumA is 1, it means connecting to port 5901. The parameter NumB is usually 0.
[0227] Since in this implementation, Xvnc and gedit are both running on the same computing node host, that is, the local machine mentioned above, the host parameter should be empty.
[0228] Then, in the fourth child process, start gedit through the execve function. After gedit is started, the gedit process automatically obtains the local socket file used by Xvnc through the environment variable DISPLAY set in the above steps, and gedit connects to the socket file. Then, the content to be displayed can be passed to Xvnc through the connection, and Xvnc then passes the displayed content to the vnc client according to the vnc protocol, so that the user can see the content displayed in the gedit application window on the vnc client.
[0229] The above is the process of BLCR starting Xvnc and gedit, and the user accessing the application through the vnc client.
[0230] In an optional implementation, the process in which the management process monitors the node communication port of the current node further includes:
[0231] S1, when the current service node is a reference service node determined by the service cluster scheduling system to respond to the reference client, determining the reference node communication ports respectively allocated to the reference client;
[0232] S2, start the reference management process for monitoring the reference node communication port;
[0233] S3: when the reference management process receives a second visualization service request from the reference client according to the reference node communication port, the reference management process starts a reference visualization service process and a reference task process for responding to the second visualization service request.
[0234] When the service cluster scheduling system determines that the current service node is the reference service node for responding to the reference client, the system will allocate an independent reference node communication port to each reference client. This step ensures that each client's request can be processed through an independent communication port, thereby achieving multi-user concurrent access.
[0235] Next, the service cluster scheduling system can start a dedicated reference management process for each assigned reference node communication port. The management process is responsible for listening to the communication port assigned to it and preparing to receive and process requests from reference clients.
[0236] Furthermore, when the reference management process receives a second visualization service request from a reference client through the reference node communication port it monitors, it will start the corresponding reference visualization service process and reference task process according to the request. These processes will be specifically used to process the client's request to ensure the personalization and isolation of the service.
[0237] It can be understood that in the above implementation of the present application, the solution is particularly suitable for scenarios where multiple users access visualization applications. When different users request the same or different applications, the system can dynamically start a BLCR super daemon process for each request, and each process listens on a different VNC port. This design not only improves the flexibility and scalability of the system, but also ensures the independence and security of user requests.
[0238] In the above implementation of the present application, a dynamic port allocation strategy may be further adopted. For example, the system may dynamically allocate ports according to the current load conditions and available resources to optimize resource usage and avoid port conflicts.
[0239] Through the above implementation of the present application, the security of remote visualization services is enhanced. By allocating independent communication ports and management processes to each client, the system can better isolate user requests and prevent interference between different users and potential security risks. In addition, the user experience is further optimized. Users can connect to the specific port assigned to them through the VNC client to access the visualization application they requested and enjoy a personalized service experience.
[0240] In an optional implementation, when the current service node is a reference service node determined by the service cluster scheduling system to respond to the reference client, after determining the reference node communication ports respectively allocated to the reference client, the method further includes:
[0241] S1, create a reference service container that matches the reference node communication port;
[0242] In the above steps, containerization technology may be involved. Container technology can be used to create an independent operating environment for each client request, thereby enhancing isolation and avoiding conflicts between different applications. Container technology implements resource isolation and restriction through Linux namespaces and control groups (cgroups).
[0243] S2, running a reference management process in the reference service container for listening to the reference node communication port;
[0244] S3: In the reference service container, starting a reference visualization service process and a reference task process for responding to the second visualization service request according to the reference management process.
[0245] In the specific implementation process, Docker container technology, such as Chroot, Cgroups and Namespace in Linux system, can be used to isolate applications from the system environment and limit resources. Chroot changes the root directory of the process to isolate the file system. Cgroups control group technology limits, records and isolates the physical resources (such as CPU, memory, etc.) used by the process group. Namespace namespace technology provides isolation of resources such as processes, networks, and mount points.
[0246] In an optional implementation manner, when the current service node is a reference service node determined by the service cluster scheduling system to respond to the reference client, the method further includes at least one of the following:
[0247] Method 1: When the number of management processes running in the current service node is greater than or equal to the number threshold, a service transfer request is sent to the service cluster scheduling system;
[0248] In the first approach, when the number of management processes running in the current service node reaches or exceeds a preset number threshold, a service transfer request is sent to the service cluster scheduling system.
[0249] Method 2: When at least one node status parameter corresponding to the current service node meets the parameter condition, a service transfer request is sent to the service cluster scheduling system;
[0250] In the second method, when at least one node status parameter of the current service node meets certain conditions, such as high CPU usage or insufficient memory, a service transfer request is sent to the service cluster scheduling system. This helps to transfer services in time when resources are tight, avoiding service interruption or performance degradation.
[0251] The service transfer request is used to request the service cluster scheduling system to re-determine a reference service node from multiple service nodes.
[0252] In the above implementation of the present application, the current service node may refuse the system's allocation when local service resources are insufficient, and request the system to allocate the client service request to other service nodes according to the load balancing principle.
[0253] In a specific implementation, assume that a service cluster scheduling system uses Docker containerization technology to deploy its services. Each client requests an independent service environment, and the service cluster scheduling system can orchestrate and manage containers through Kubernetes. When the number of containers deployed in a node reaches a threshold, Kubernetes automatically deploys new containers to other nodes to balance the load and ensure high availability of the service.
[0254] In another specific implementation, the service cluster scheduling system can use containerization technology to deploy applications of microservice architecture. The platform monitors the resource usage of each service node, such as CPU and memory usage. When it is detected that the resource usage of a node is close to the upper limit, the platform will automatically transfer some services to other nodes based on the feedback from the node to avoid overload and ensure the stability of the service. It is understandable that during the service transfer process, the checkpoint creation method in the above implementation can be further adopted to determine the checkpoint file of the service to be transferred. And forward the checkpoint file to the reallocated service node, thereby realizing the dynamic transfer of the service.
[0255] In an optional implementation, after providing visualization service to the target client through the visualization service process and the target task process, the method further includes:
[0256] S1, when the management process receives an end service instruction according to the node communication port, obtains at least one start timestamp matching the visualization service and a termination timestamp corresponding to each of the at least one start timestamp, wherein the start timestamp is used to indicate a start running time node of the visualization service process, and the termination timestamp is used to indicate a termination running time node of the visualization service process;
[0257] S2, determining a target service duration according to at least one start timestamp and respective corresponding stop timestamps;
[0258] S3, sending service duration reminder information to the target client according to the target service duration.
[0259] In the above-mentioned implementation mode of the present application, in a high-performance cluster scheduling system, it is usually necessary to meter and bill the resources used by users. In our solution, because the visualization application used by the user is started through BLCR, and the started visualization application is a child process of the BLCR process, when the visualization application process exits, it can also be perceived by its parent process, namely the BLCR process. In this way, BLCR can accurately record the start and end time of the user's use of the visualization application, thereby performing accurate metering and billing.
[0260] S1: When the management process receives an end service instruction according to the node communication port, at least one start timestamp matching the visualization service and a termination timestamp corresponding to each of the at least one start timestamp are obtained. The start timestamp is used to indicate the start running time node of the visualization service process, and the termination timestamp is used to indicate the termination running time node of the visualization service process.
[0261] S2: Determine a target service duration according to at least one start timestamp and the corresponding stop timestamps.
[0262] S3: Sending service duration reminder information to the target client according to the target service duration.
[0263] In the above implementation of the present application, the management process in the current service node can accurately record the start and end time of the visualization service process. This can be achieved by having the BLCR process record the current timestamp when the process starts and exits. The timestamp can be stored in a log file or database for subsequent billing and analysis.
[0264] Furthermore, in the service duration calculation process, the service duration can be determined by calculating the difference between the end timestamp and the start timestamp. This calculation can be performed immediately when the process exits, or when needed.
[0265] In an optional manner, the service duration reminder information may be displayed to the user through a user interface, or may be sent to the user through an email, text message, or the like.
[0266] In the above implementation of the present application, in a high-performance cluster scheduling system, it is usually necessary to meter and bill the resources used by the user. In our solution, because the visualization application used by the user is started through BLCR, and the started visualization application is a child process of the BLCR process, when the visualization application process exits, it can also be sensed by its parent process, i.e., the BLCR process, so that BLCR can accurately record the start and end time of the user's use of the visualization application, thereby performing accurate metering and billing.
[0267] Through the above implementation of the present application, the super daemon process and BLCR can be combined and applied to process management in high-performance computing scenarios to achieve efficient, transparent and automatic checkpoint recovery of the job process.
[0268] Specifically, in the process of comprehensive capture and recovery of process status, combined with the BLCR tool, comprehensive status capture of the process is achieved, including CPU status, memory status and socket status, ensuring that it can be quickly restored from the last saved status when a failure occurs.
[0269] In addition, the implementation of this application also supports manual / automatic checkpoint mechanisms. It supports users to manually create checkpoints and exit applications; it also supports periodic saving of job status through automated timed checkpoint strategies, which reduces computing interruptions and resource waste caused by failures, and improves system stability and reliability.
[0270] In addition, the implementation method of the present application further provides fast and flexible recovery capabilities. When a process ends abnormally, the checkpoint file can be used to quickly restore the process state. The effect of automatic recovery can be achieved, that is, when a user accesses an application through a vnc client, the BLCR process automatically checks whether there is a previously created checkpoint file. If so, Xvnc and the visualization application are restored from the checkpoint; if not, Xvnc and the visualization application are restarted.
[0271] The present invention integrates these innovations into an integrated solution to meet the needs of creating checkpoints and quickly recovering visual job processes in high-performance computing environments. Through this innovative integrated approach, the present invention provides an efficient and reliable process management strategy, which significantly improves the job execution efficiency and stability of high-performance computing clusters.
[0272] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0273] According to another aspect of the embodiment of the present application, a data processing device in a memory for implementing the above-mentioned process state recovery method is also provided. Figure 7 As shown, the device comprises:
[0274] The receiving unit 702 is used to receive a visualization service request sent by a target client to the node communication port of the current service node during the process of the management process monitoring the node communication port of the current service node;
[0275] A first recovery unit 704 is configured to recover the visualization service process according to the first recovery information recorded in the checkpoint file when a checkpoint file matching the visualization service request is found, wherein the visualization service process is used to provide a visualization service interface to the target client;
[0276] A second recovery unit 706, configured to recover the target task process according to the second recovery information recorded in the checkpoint file, wherein the target task process is the task process for executing the target service task in the visual service interface;
[0277] The service unit 708 is used to provide visualization service to the target client through the visualization service process and the target task process when the process communication between the visualization service process and the target task process is restored.
[0278] Optionally, the above-mentioned first recovery unit 704: creates a first sub-process through the above-mentioned management process; obtains the communication description information included in the above-mentioned first recovery information according to the above-mentioned first sub-process, and performs a communication recovery operation according to the above-mentioned communication description information; obtains the process state description information included in the above-mentioned first recovery information according to the above-mentioned first sub-process, and performs a process state recovery operation on the above-mentioned visualization service process according to the above-mentioned communication description information.
[0279] Optionally, the first recovery unit 704: determines a first communication description identifier that matches the first sub-process; obtains a first historical communication description identifier that matches the visualization service process from the first recovery information through the first sub-process; and updates the first communication description identifier according to the first historical communication description identifier.
[0280] Optionally, the second recovery unit 706: creates a second sub-process through the management process; obtains the communication description information included in the second recovery information according to the second sub-process, and performs a communication recovery operation on the target task process according to the communication description information; obtains the process status description information included in the second recovery information according to the second sub-process, and performs a process status recovery operation on the target task process according to the communication description information.
[0281] Optionally, the second recovery unit 706: determines a second communication description identifier that matches the second sub-process, and establishes a process communication connection between the first sub-process and the second sub-process based on the first communication description identifier and the second communication description identifier; obtains a second historical communication description identifier that matches the visualization service process from the second recovery information through the second sub-process; and updates the second communication description identifier based on the second historical communication description identifier.
[0282] Optionally, the above-mentioned process state recovery device also includes: a third recovery unit, used for at least one of the following: restoring the memory state corresponding to the above-mentioned current process according to the historical memory state parameters in the above-mentioned process state description information that matches the current process; restoring the CPU state corresponding to the above-mentioned current process according to the historical CPU state parameters in the above-mentioned process state description information that matches the current process; restoring the reference process state corresponding to the above-mentioned current process according to the historical reference state parameters in the above-mentioned process state description information that matches the current process, wherein the above-mentioned reference process state includes at least one of the following: process priority state, device resource state, file handle state; wherein the above-mentioned current process includes the above-mentioned visualization service process and the above-mentioned target task process.
[0283] Optionally, the above-mentioned process status recovery device also includes: an interruption unit, which is used to suspend the above-mentioned target task process when the above-mentioned management process detects a target event, and save the communication description information used to indicate the communication status of the above-mentioned target task process and the process description information used to indicate the process status of the above-mentioned target task process to the above-mentioned second recovery information in the above-mentioned checkpoint file; suspend the above-mentioned visualization service process, and save the communication description information used to indicate the communication status of the above-mentioned visualization service process and the process description information used to indicate the process status of the above-mentioned visualization service process to the above-mentioned first recovery information in the above-mentioned checkpoint file; end the above-mentioned target task process and the above-mentioned visualization service process.
[0284] Optionally, the above-mentioned interruption unit is also used for at least one of the following: when the above-mentioned management process receives a service suspension instruction according to the above-mentioned node communication port, it is determined that the above-mentioned target event is detected; when the above-mentioned management process detects communication interruption indication information according to the above-mentioned node communication port, it is determined that the above-mentioned target event is detected; when the above-mentioned management process detects that the current timestamp is a target timestamp that matches the target inspection period, it is determined that the above-mentioned target event is detected.
[0285] Optionally, the above-mentioned interrupt unit is also used for, when the above-mentioned current service node is the target service node determined by the service cluster scheduling system to respond to the above-mentioned target client, the above-mentioned management process obtains the first visualization service request of the above-mentioned target client according to the allocated above-mentioned node communication port, wherein the above-mentioned service cluster scheduling system includes multiple service nodes, and when the above-mentioned service cluster scheduling system receives the visualization service request of at least one client, the above-mentioned service cluster scheduling system allocates the above-mentioned service nodes corresponding to each of the above-mentioned clients respectively; starts the visualization service process through the above-mentioned management process, and starts the above-mentioned target task process through the above-mentioned management process; when the process communication connection between the above-mentioned visualization service process and the above-mentioned target task process is established, and the service communication connection between the above-mentioned visualization service process and the above-mentioned target client is established, the above-mentioned visualization service process and the above-mentioned target task process are provided to the above-mentioned target client through the above-mentioned visualization service process and the above-mentioned target task process.
[0286] Optionally, the device for recovering the process state further includes: a starting unit, used to start the management process when the current service node is the target node determined by the service cluster scheduling system to respond to the target client; monitor the node communication port allocated to the target client through the management process; and maintain the monitoring state of the management process when the management process does not detect the first visual service request of the target client according to the node communication port.
[0287] Optionally, the startup unit is used to: create a third sub-process through the management process, and obtain at least one visualization request parameter carried in the first visualization service request according to the third sub-process; start the visualization service process according to at least one visualization request parameter; create a fourth sub-process through the management process, and configure the visualization service environment according to the fourth sub-process; and when the visualization service environment is configured, start the target task process according to the fourth sub-process.
[0288] Optionally, the startup unit is further used to: obtain a third communication description identifier matching the node communication port according to the third sub-process, and update the fourth communication description identifier of the third sub-process according to the third communication description identifier; when the update of the fourth communication description identifier of the third sub-process is completed, determine that the service communication connection between the visualization service process and the target client is established; according to the fourth sub-process, obtain the fourth communication description identifier from the configuration information of the visualization service environment; the visualization service process and the target task process establish the process communication connection according to the fourth communication description identifier.
[0289] Optionally, the device for recovering the process state further includes: a reference service unit, used to determine the reference node communication ports respectively allocated to the reference clients when the current service node is the reference service node determined by the service cluster scheduling system to respond to the reference client; start a reference management process for monitoring the reference node communication port; and when the reference management process receives a second visualization service request from the reference client according to the reference node communication port, start a reference visualization service process and a reference task process for responding to the second visualization service request according to the reference management process.
[0290] Optionally, the above-mentioned reference service unit is also used to: create a reference service container that matches the above-mentioned reference node communication port; run a reference management process for listening to the above-mentioned reference node communication port in the above-mentioned reference service container; in the above-mentioned reference service container, start the above-mentioned reference visualization service process and the above-mentioned reference task process for responding to the above-mentioned second visualization service request according to the above-mentioned reference management process.
[0291] Optionally, the reference service unit is also used for at least one of the following: sending a service transfer request to the service cluster scheduling system when the number of the management processes running in the current service node is greater than or equal to a quantity threshold; sending a service transfer request to the service cluster scheduling system when at least one node status parameter corresponding to the current service node meets a parameter condition; wherein the service transfer request is used to request the service cluster scheduling system to redetermine the reference service node from the multiple service nodes.
[0292] Optionally, the device for recovering the process state further includes: a statistical unit, which is used to obtain at least one start timestamp matching the above-mentioned visualization service and a termination timestamp corresponding to each of the at least one start timestamp when the management process receives an end service instruction according to the node communication port, wherein the start timestamp is used to indicate the start running time node of the above-mentioned visualization service process, and the termination timestamp is used to indicate the termination running time node of the above-mentioned visualization service process; determine the target service duration according to the at least one start timestamp and the corresponding termination timestamps; and send service duration prompt information to the target client according to the target service duration.
[0293] For a specific embodiment, reference may be made to the example shown in the above process status recovery method, which will not be described in detail in this example.
[0294] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned method for processing data in a memory is also provided. The electronic device may be Figure 1The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a mobile phone or a computer as an example. Figure 8 As shown, the electronic device includes a second memory 802 and a second processor 804. The second memory 802 stores a computer program. The second processor 804 is configured to execute the steps of any of the above method embodiments through the computer program.
[0295] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0296] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0297] S1, receiving a visualization service request sent by a target client to the node communication port of the current service node during the process of the management process monitoring the node communication port of the current service node;
[0298] S2, when a checkpoint file matching the visualization service request is found, resuming the visualization service process according to the first recovery information recorded in the checkpoint file, wherein the visualization service process is used to provide a visualization service interface to the target client;
[0299] S3, restoring the target task process according to the second restoration information recorded in the checkpoint file, wherein the target task process is the task process used to execute the target service task in the visual service interface;
[0300] S4, when the process communication between the visualization service process and the target task process is restored, providing visualization service to the target client through the visualization service process and the target task process.
[0301] Alternatively, a person skilled in the art may understand that: Figure 8 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 8 The structure of the electronic device is not limited. Figure 8 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 8 Different configurations are shown.
[0302] Among them, the second memory 802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the process state recovery method and device in the embodiment of the present application. The second processor 804 executes various functional applications and data processing by running the software programs and modules stored in the second memory 802, that is, realizing the above-mentioned process state recovery method. The second memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the second memory 802 may further include a memory remotely arranged relative to the second processor 804, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. As an example, Figure 8 As shown, the second memory 802 may include, but is not limited to, the receiving unit 702, the first recovery unit 704, the second recovery unit 706, and the service unit 708 in the process state recovery device. In addition, other module units in the process state recovery device may also be included but are not limited to, which will not be repeated in this example.
[0303] Optionally, the transmission device 806 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 806 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0304] In addition, the electronic device further includes: a display 808 for displaying the target page; and a connection bus 810 for connecting various module components in the electronic device.
[0305] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. The nodes may form a point-to-point network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the point-to-point network.
[0306] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method provided in the above various optional implementations;
[0307] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0308] S1, receiving a visualization service request sent by a target client to the node communication port of the current service node during the process of the management process monitoring the node communication port of the current service node;
[0309] S2, when a checkpoint file matching the visualization service request is found, resuming the visualization service process according to the first recovery information recorded in the checkpoint file, wherein the visualization service process is used to provide a visualization service interface to the target client;
[0310] S3, restoring the target task process according to the second restoration information recorded in the checkpoint file, wherein the target task process is the task process used to execute the target service task in the visual service interface;
[0311] S4, when the process communication between the visualization service process and the target task process is restored, providing visualization service to the target client through the visualization service process and the target task process.
[0312] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0313] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0314] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.
[0315] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0316] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0317] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0318] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0319] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
[0320] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0321] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0322] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0323] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0324] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0325] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0326] The embodiments of the present application also provide a computer program, which includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in any one of the above method embodiments.
[0327] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.
[0328] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0329] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for restoring a process state, characterized in that: include: In the process of the management process monitoring the node communication port of the current service node, receiving a visualization service request sent by a target client to the node communication port; In the case where a checkpoint file matching the visualization service request is found, resuming the visualization service process according to the first resumption information recorded in the checkpoint file, wherein the visualization service process is used to provide a visualization service interface to the target client; Recovering the target task process according to the second recovery information recorded in the checkpoint file, wherein the target task process is a task process used to execute the target service task in the visual service interface; When the process communication between the visualization service process and the target task process is restored, the visualization service is provided to the target client through the visualization service process and the target task process.
2. The method according to claim 1, characterized in that The step of restoring the visualization service process according to the first restoration information recorded in the checkpoint file includes: Creating a first child process through the management process; Acquire, according to the first sub-process, the communication description information included in the first recovery information, and perform a communication recovery operation according to the communication description information; The process state description information included in the first recovery information is acquired according to the first sub-process, and a process state recovery operation is performed on the visualization service process according to the communication description information.
3. The method according to claim 2, characterized in that The acquiring, according to the first sub-process, the communication description information included in the first recovery information, and performing a communication recovery operation according to the communication description information, includes: Determining a first communication description identifier that matches the first subprocess; Acquiring, by the first subprocess, from the first recovery information, a first historical communication description identifier that matches the visualization service process; The first communication description identifier is updated according to the first historical communication description identifier.
4. The method according to claim 2, characterized in that: The step of restoring the target task process according to the second restoration information recorded in the checkpoint file includes: Creating a second child process through the management process; Acquire, according to the second sub-process, the communication description information included in the second recovery information, and perform a communication recovery operation according to the communication description information; The process state description information included in the second recovery information is obtained according to the second sub-process, and a process state recovery operation is performed on the target task process according to the communication description information.
5. The method according to claim 4, characterized in that The acquiring, according to the second sub-process, the communication description information included in the second recovery information, and performing a communication recovery operation according to the communication description information, comprises: Determine a second communication description identifier that matches the second sub-process, and establish a process communication connection between the first sub-process and the second sub-process according to the first communication description identifier and the second communication description identifier; Acquire, by the second subprocess, from the second recovery information, a second historical communication description identifier that matches the target task process; The second communication description identifier is updated according to the second historical communication description identifier.
6. The method according to any one of claims 2 or 4, characterized in that The execution process state recovery operation includes at least one of the following: Restoring the memory state corresponding to the current process according to the historical memory state parameters in the process state description information matching the current process; Restoring the CPU state corresponding to the current process according to the historical CPU state parameters in the process state description information matching the current process; Restore the reference process state corresponding to the current process according to the historical reference state parameters in the process state description information matching the current process, wherein the reference process state includes at least one of the following: process priority state, device resource state, and file handle state; The current process includes the visualization service process and the target task process.
7. The method according to claim 1, characterized in that Before receiving the visualization service request sent by the target client to the node communication port, the method further includes: In the case where the management process detects a target event, pausing the target task process, and saving communication description information indicating the communication state of the target task process and process description information indicating the process state of the target task process to the second recovery information in the checkpoint file; Pausing the visualization service process, and saving communication description information used to indicate the communication state of the visualization service process and process description information used to indicate the process state of the visualization service process to the first recovery information in the checkpoint file; End the target task process and the visualization service process.
8. The method according to claim 7, characterized in that Before the management process detects the target event, the method further includes at least one of the following: In a case where the management process receives a service suspension instruction according to the node communication port, determining that the target event is detected; In a case where the management process detects communication interruption indication information according to the node communication port, determining that the target event is detected; In a case where the management process detects that the current timestamp is a target timestamp that matches a target inspection period, it is determined that the target event is detected.
9. The method according to claim 7, characterized in that: In the case where the management process detects a target event, before suspending the target task process, the process further includes: In a case where the current service node is a target service node determined by the service cluster scheduling system to respond to the target client, the management process obtains a first visualization service request of the target client according to the assigned node communication port, wherein the service cluster scheduling system includes a plurality of service nodes, and the service cluster scheduling system allocates the corresponding service nodes to at least one of the clients respectively when receiving a visualization service request of at least one client; Start the visualization service process through the management process, and start the target task process through the management process; When the process communication connection between the visualization service process and the target task process is established, and the service communication connection between the visualization service process and the target client is established, the visualization service matching the first visualization service request is provided to the target client through the visualization service process and the target task process.
10. The method according to claim 9, characterized in that Before the management process obtains the first visualization service request of the target client according to the allocated node communication port, the process further includes: In a case where the current service node is a target node determined by the service cluster scheduling system to respond to the target client, starting the management process; Monitoring the node communication port allocated to the target client by the management process; When the management process fails to detect the first visualization service request of the target client according to the node communication port, the monitoring state of the management process is maintained.
11. The method according to claim 9, characterized in that The step of starting the visualization service process through the management process and starting the target task process through the management process includes: Creating a third sub-process through the management process, and acquiring at least one visualization request parameter carried in the first visualization service request according to the third sub-process; Starting the visualization service process according to at least one visualization request parameter; Creating a fourth sub-process through the management process, and configuring a visualization service environment according to the fourth sub-process; When the visualization service environment is configured, the target task process is started according to the fourth sub-process.
12. The method according to claim 11, characterized in that Before providing the target client with a visualization service matching the first visualization service request through the visualization service process and the target task process, the method further includes: Acquire a third communication description identifier matching the node communication port according to the third sub-process, and update a fourth communication description identifier of the third sub-process according to the third communication description identifier; When the fourth communication description identifier of the third subprocess is updated, determining that the service communication connection between the visualization service process and the target client is established; According to the fourth sub-process, acquiring the fourth communication description identifier from the configuration information of the visualization service environment; The visualization service process and the target task process establish the process communication connection according to the fourth communication description identifier.
13. The method according to claim 9, characterized in that In the process of the management process monitoring the node communication port of the current node, the method further includes: In a case where the current service node is a reference service node determined by the service cluster scheduling system for responding to a reference client, determining reference node communication ports respectively allocated to the reference clients; Starting a reference management process for monitoring the communication port of the reference node; When the reference management process receives the second visualization service request from the reference client according to the reference node communication port, a reference visualization service process and a reference task process for responding to the second visualization service request are started according to the reference management process.
14. The method according to claim 13, characterized in that In the case where the current service node is a reference service node determined by the service cluster scheduling system to respond to the reference client, after determining the reference node communication ports respectively allocated to the reference client, the method further includes: Creating a reference service container matching the reference node communication port; Running a reference management process in the reference service container for monitoring the reference node communication port; In the reference service container, the reference visualization service process and the reference task process for responding to the second visualization service request are started according to the reference management process.
15. The method according to claim 13, characterized in that In the case where the current service node is a reference service node determined by the service cluster scheduling system for responding to the reference client, at least one of the following is further included: When the number of the management processes running in the current service node is greater than or equal to a number threshold, sending a service transfer request to the service cluster scheduling system; When at least one node status parameter corresponding to the current service node satisfies a parameter condition, sending a service transfer request to the service cluster scheduling system; The service transfer request is used to request the service cluster scheduling system to re-determine the reference service node from the multiple service nodes.
16. The method according to claim 9, characterized in that After providing the visualization service to the target client through the visualization service process and the target task process, the method further includes: In the case where the management process receives an end service instruction according to the node communication port, obtaining at least one start timestamp matching the visualization service and a termination timestamp corresponding to each of the at least one start timestamp, wherein the start timestamp is used to indicate a start running time node of the visualization service process, and the termination timestamp is used to indicate a termination running time node of the visualization service process; Determine a target service duration according to at least one start timestamp and respective corresponding stop timestamps; Send service duration prompt information to the target client according to the target service duration.
17. A process status recovery device, characterized in that: include: A receiving unit, configured to receive a visualization service request sent by a target client to the node communication port of the current service node during a process in which the management process monitors the node communication port of the current service node; A first recovery unit, configured to recover a visualization service process according to first recovery information recorded in the checkpoint file when a checkpoint file matching the visualization service request is found, wherein the visualization service process is used to provide a visualization service interface to the target client; A second recovery unit, configured to recover a target task process according to second recovery information recorded in the checkpoint file, wherein the target task process is a task process in the visual service interface for executing the target service task; The service unit is used to provide visualization service to the target client through the visualization service process and the target task process when the process communication between the visualization service process and the target task process is restored.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 16 when executed by a processor.
19. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 16 are implemented.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 16 are implemented.
Citation Information
Patent Citations
Method for implementing checkpoint of Linux program at user level based on virtual kernel object
CN101093453A
Cluster fault-tolerance system, apparatus and method
CN101369241A
Dual-process redundancy transient fault tolerating method
CN103064770A
Recommended data display method and system
CN105893558A
Fault recovery method and apparatus
CN106528324A