Function monitoring method, device, equipment, medium and product

By synchronizing heartbeat messages between nodes in a distributed system, the stability problem of function monitoring when local devices malfunction is solved, and accurate monitoring of function status and anomaly location are achieved.

CN121333992APending Publication Date: 2026-01-13INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511351425.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In the event of local equipment malfunction or downtime, existing technologies struggle to effectively monitor the operational status of service functions, resulting in poor service stability.

Method used

A distributed system is used for function monitoring. Function information is synchronized in the distributed node system through heartbeat messages. Other nodes monitor the running status of the function to ensure that the accuracy and consistency of function information can be maintained even if a node fails.

Benefits of technology

It improves the stability and accuracy of function monitoring, enabling continued monitoring of function status even when nodes are abnormal, thus enhancing the accuracy and efficiency of anomaly localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121333992A_ABST
    Figure CN121333992A_ABST
Patent Text Reader

Abstract

The invention provides a function monitoring method, device and equipment, a medium and a product, which can be applied to the field of distributed technologies. The method is applied to any node in the distributed node system, and is used for monitoring the running state of a target function in the distributed node system. The method comprises the steps of determining that a first target function is in a running state under the condition that a heartbeat message for the first target function in other nodes is received; and for a second target function in other nodes, under the condition that the duration between the preset moment of a heartbeat message which is received last time and aims at the second target function and the current moment is greater than a preset duration threshold value, sending a function detection request aiming at the second target function to other nodes to which the second target function belongs, and determining whether the second target function is in a running state or not based on a feedback result of the function detection request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed technology, specifically to a function monitoring method, apparatus, device, medium, and product. Background Technology

[0002] With the increasing prevalence of digitalization, the stability of business functions is receiving more and more attention. For example, for services provided locally by devices, it is often necessary to monitor the operational status of these services in order to perform appropriate processing operations when services malfunction, thereby improving service stability.

[0003] However, the self-monitoring method for local service functions often fails to function properly and has poor stability when local devices malfunction or crash. Summary of the Invention

[0004] In view of the above problems, this application provides a function monitoring method, apparatus, equipment, medium and product for improving the stability of function monitoring.

[0005] According to a first aspect of this application, a function monitoring method is provided, applied to any node in a distributed node system. The method is used to monitor the running status of a target function in the distributed node system. The method includes: upon receiving a heartbeat message for a first target function in another node, determining that the first target function is in a running state; and for a second target function in another node, if the duration between a preset time of a previously received heartbeat message for the second target function and the current time is greater than a preset duration threshold, sending a function detection request for the second target function to the other nodes to which the second target function belongs, and determining whether the second target function is in a running state based on the feedback result of the function detection request.

[0006] Optionally, determining whether the second target function is in a running state based on the feedback result of the function detection request includes: determining that the second target function is in an abnormal state if no heartbeat message for the second target function is received within a preset waiting time; and determining that the second target function is in a running state if a heartbeat message for the second target function is received within a preset waiting time.

[0007] Optionally, the method further includes: synchronizing heartbeat messages for the third objective function in the distributed node system when it is determined that the local third objective function is in a running state.

[0008] Optionally, the method further includes: if it is determined that any local function needs to monitor its running status, the determined function is identified as the fourth objective function, and the information of the fourth objective function is synchronized in the distributed node system.

[0009] Optionally, any node in the distributed node system stores a set of objective function information; the method further includes: if it is determined that the locally stored set of objective function information is inconsistent with the set of objective function information stored by other nodes, performing a synchronization operation on the set of objective function information stored between different nodes in the distributed node system.

[0010] Optionally, the target function information set includes the state information of the target function and the corresponding state determination time information; the step of performing a synchronization operation on the target function information set stored between different nodes in the distributed node system includes: for any target function in the target function information set stored between different nodes in the distributed node system, determining the current state information of the target function based on the corresponding state determination time information, and synchronizing the determined current state information in the distributed node system.

[0011] Optionally, any node in the distributed node system stores a set of objective function information; the method further includes: determining, based on the hash value of the objective function information set, whether the locally stored objective function information set is consistent with the objective function information sets stored by other nodes.

[0012] Optionally, determining whether the locally stored set of target function information is consistent with the set of target function information stored on other nodes based on the hash value of the target function information set includes: determining whether the locally stored set of target function information is consistent with the set of target function information stored on other nodes based on the hash value of a preset group in the target function information set; wherein the preset group contains information about target functions belonging to the same node; the method further includes: performing a synchronization operation on the preset group that is inconsistent between different nodes in the distributed node system when it is determined that the locally stored set of target function information is inconsistent with the set of target function information stored on other nodes.

[0013] A second aspect of this application provides a function monitoring device applied to any node in a distributed node system. The device is used to monitor the running status of a target function in the distributed node system. The device includes: a first module, used to determine that the first target function is running when a heartbeat message for a first target function in another node is received; and a second module, used to send a function detection request for a second target function to other nodes to which the second target function belongs, if the duration between a preset time of a previously received heartbeat message for the second target function and the current time is greater than a preset duration threshold, and to determine whether the second target function is running based on the feedback result of the function detection request.

[0014] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0015] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0016] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description

[0017] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0018] Figure 1 This illustration schematically depicts an application scenario of a function monitoring method according to an embodiment of this application;

[0019] Figure 2 A flowchart illustrating a function monitoring method according to an embodiment of this application is shown schematically;

[0020] Figure 3 This schematic diagram illustrates a structural block diagram of a function monitoring device according to an embodiment of this application;

[0021] Figure 4 A block diagram schematically illustrates an electronic device suitable for implementing a function monitoring method according to an embodiment of this application. Detailed Implementation

[0022] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0025] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0026] With the increasing prevalence of digitalization, the stability of business functions is receiving more and more attention. For example, for services provided locally on devices, it is often necessary to monitor the operational status of these services to take appropriate action in case of anomalies, thereby improving service stability. However, self-monitoring of local services often proves ineffective in the event of local device malfunctions or downtime, resulting in poor stability. For instance, if a local device malfunctions, the message indicating a service failure cannot be delivered to relevant personnel, rendering the service unusable and reducing its stability.

[0027] To address the aforementioned technical problems, embodiments of this application provide a function monitoring method.

[0028] This method employs a distributed system for function monitoring. The distributed system can contain multiple nodes, each running functions. Monitoring can be performed on these running functions, specifically monitoring their execution status. Furthermore, for any given function on any node, the distributed system as a whole can monitor it. Specifically, all nodes in the distributed system can jointly synchronize a set of function information to be monitored, containing information about the functions running on each node. Each node can then monitor the execution status of functions on other nodes, maintaining and updating its local function information set. This can be achieved through heartbeat messages from functions on other nodes.

[0029] Since a distributed system is used for function monitoring, it is understandable that even if any node malfunctions or crashes, the running status information of the function can be determined through other nodes, thereby improving the stability of function monitoring.

[0030] Furthermore, since this method monitors functions, it can improve the granularity of monitoring compared to solutions that monitor application instances or processes, making it easier to accurately locate abnormal functions and improving the accuracy, precision, and efficiency of anomaly detection.

[0031] It should be noted that the function monitoring method and apparatus provided in the embodiments of this application can be applied to the field of distributed technology and also to the field of fintech. For example, for functions in the financial or banking fields, the function monitoring method provided in the embodiments of this application can be used for monitoring. The embodiments of this application can also be applied to any field other than fintech, such as risk control, audio / video, or image processing. When function monitoring is required, the function monitoring method provided in the embodiments of this application can be used. The application fields of the function monitoring method and apparatus provided in the embodiments of this application are not limited.

[0032] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.

[0033] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.

[0034] Figure 1 The illustration shows an application scenario diagram of a function monitoring method according to an embodiment of this application.

[0035] like Figure 1 As shown, application scenario 100 according to this embodiment may include: a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0036] Users can use the first terminal device 101, the second terminal device 102, or the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0038] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, or the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0039] It should be noted that the function monitoring method provided in this application embodiment can generally be executed by server 105. Correspondingly, the function monitoring device provided in this application embodiment can generally be located in server 105. The function monitoring method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the function monitoring device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0040] The server 105 can be a server in a distributed system, which can execute a function monitoring method provided in this application embodiment to monitor the running status of functions in the distributed system.

[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0042] Figure 2 A flowchart illustrating a function monitoring method according to an embodiment of this application is shown schematically.

[0043] like Figure 2 As shown, the function monitoring method provided in this embodiment may include operations S210 to S220. The embodiments of this application do not limit the executing entity of the function monitoring method; optionally, it can be any device or any application, specifically any system in a distributed system.

[0044] Optionally, the function monitoring method can be applied to any node in a distributed node system. This distributed node system can contain multiple nodes that can communicate with each other. It is understood that this method flow describes the steps executed for any node in the distributed node system. For other nodes or individual nodes in the distributed node system, this method flow can be executed separately to achieve function monitoring between nodes. Optionally, the function monitoring method provided in the embodiments of this application can be used to monitor the running status of a target function in a distributed node system.

[0045] In operation S210, upon receiving a heartbeat message for the first objective function in another node, it is determined that the first objective function is in a running state.

[0046] In operation S220, for the second target function in other nodes, if the duration between the preset time of the previously received heartbeat message for the second target function and the current time is greater than the preset duration threshold, a function detection request for the second target function is sent to the other nodes to which the second target function belongs, and the second target function is determined to be in a running state based on the feedback result of the function detection request.

[0047] This method utilizes a distributed node system to monitor the function's runtime status, improving the stability of function monitoring. Even if any node in the distributed node system malfunctions or crashes, monitoring of the target function's runtime status can continue, as other nodes can be used to determine the target function's runtime status, further enhancing the stability of function monitoring.

[0048] This method can also monitor the running status of the target function. Compared with the solution of monitoring application instances or processes, it can improve the monitoring granularity, facilitate the accurate location of abnormal functions, and improve the accuracy, precision and efficiency of anomaly location.

[0049] This method can also send a function detection request to the corresponding node when the heartbeat message of the target function times out, in order to determine the current state of the target function and improve the accuracy of the function's running status.

[0050] The embodiments of this application do not limit the objective function. Optionally, the objective function can be any function in the distributed node system that needs to monitor its running status; for ease of description, it is referred to as the objective function. It is understood that a distributed node system may contain one or more objective functions, and the embodiments of this application do not limit the number of objective functions.

[0051] Furthermore, for ease of distinction, the objective functions for different situations are described using terms such as "first objective function," "second objective function," "third objective function," etc., and so on. It's understandable that these objective functions may or may not be the same. For example, the first objective function and the second objective function can be descriptions of the same objective function under different circumstances. Specifically, the objective function for receiving a heartbeat message can be referred to as the first objective function, while if the heartbeat message times out under this objective function, it can be referred to as the second objective function. These descriptions (first, second, third, and fourth) are used to distinguish and describe the objective function under different circumstances, and do not limit whether these objective functions are the same function.

[0052] The embodiments of this application do not limit the state of the objective function. Optionally, the state of the objective function may include a running state and an abnormal state. The running state indicates that the objective function is running normally and can perform the corresponding function. The abnormal state indicates that the objective function is running abnormally, specifically, the objective function may stop running and cannot perform the corresponding function. Of course, the state of the objective function may also include other states.

[0053] The embodiments of this application do not limit the distribution of the objective function in the distributed node system. Optionally, any node in the distributed node system may run one or more objective functions, or it may not run any objective function.

[0054] Optionally, for monitoring the objective function, any node in the distributed node system can contain information about one or more objective functions that need to be monitored in the distributed node system. This facilitates monitoring the running status of the objective functions. Specifically, it can contain information about all objective functions in the distributed node system, facilitating monitoring the running status of all objective functions. Specifically, it can record the information about the objective functions based on their monitored running status for easy subsequent querying.

[0055] The embodiments of this application do not limit the specific method of monitoring the running status of the target function. Optionally, monitoring can be performed through heartbeat messages of the target function. Specifically, for a running target function, heartbeat messages can be periodically synchronized in the distributed node system, which can involve periodically sending heartbeat messages to other nodes in the distributed node system. For a node that receives a heartbeat message for the target function, it can determine that the target function is currently running, and can record the information represented by the heartbeat message locally. Specifically, it can record that the target function is currently running, and record the sending or receiving time of the heartbeat message, facilitating subsequent determination of whether the heartbeat message has timed out.

[0056] Accordingly, for a target function in an abnormal state, it cannot send heartbeat messages, and other nodes cannot receive new heartbeat messages for the target function. Therefore, based on the sending or receiving time of the previous heartbeat message, it can be determined whether the heartbeat message has timed out, and thus it can be further determined whether the target function is in an abnormal state.

[0057] The embodiments of this application do not limit the format and sending process of the heartbeat message. Optionally, the target function may include a program for periodically sending heartbeat messages, so that heartbeat messages can be sent periodically during the normal operation of the target function; and heartbeat message sending can be stopped when the target function is not currently running. Alternatively, a thread or process for periodically sending heartbeat messages can be configured for the target function, which can monitor the running status of the target function, and periodically send heartbeat messages when it is determined that the target function is currently running; and stop sending heartbeat messages when it is determined that the target function is not currently running.

[0058] The embodiments of this application do not limit the specific data synchronization method in the distributed node system, nor do they limit the synchronization method of the heartbeat messages. Optionally, a node running a target function can actively synchronize heartbeat messages for the target function with the distributed node system. Specifically, this can be done periodically by synchronizing heartbeat messages, or by sending heartbeat messages for the target function based on requests from other nodes. A node running a target function can also send heartbeat messages for the target function to other nodes in the distributed node system, or it can send heartbeat messages for the target function to neighboring nodes. This allows neighboring nodes to continue forwarding heartbeat messages for the target function to other nodes until the heartbeat messages for the target function are synchronized in the distributed node system.

[0059] It is understandable that information about the objective function whose running status needs to be monitored in a distributed node system, as well as other information, can be synchronized in the distributed node system using the methods mentioned above or other means.

[0060] In one optional embodiment, heartbeat messages in the distributed node system can be sent periodically. The embodiments of this application do not limit the specific method and process of periodically sending heartbeat messages. Optionally, for a single objective function, heartbeat messages can be periodically synchronized in the distributed node system. Optionally, different types of operations can be periodically executed in the distributed node system, divided into different time periods. Different operations can be controlled to be executed by each node in different time periods, thereby facilitating the improvement of data consistency in the distributed node system.

[0061] In one specific embodiment, a cyclical period can be set in the distributed node system. Within this period, there may be a heartbeat message transmission period, during which heartbeat messages are transmitted between nodes for the monitored target function. During non-heartbeat message transmission periods, nodes may choose not to transmit or receive heartbeat messages, thus facilitating control over heartbeat message transmission. The period may also include a data synchronization period, during which data synchronization can be performed, specifically synchronizing the running status of the target function between nodes. Furthermore, the period may include a target function registration period, during which a new function running on a node can be registered as the target function whose running status needs to be monitored, thereby synchronizing the information of the new target function to the distributed node system. By performing different operations in different time periods, data consistency in the distributed node system can be easily improved.

[0062] Regarding the second objective function, the main explanation concerns the heartbeat message timeout situation. If no heartbeat message for the second objective function is received for an extended period, it can be directly determined that the second objective function is in an abnormal state. Further status determination can also be performed, specifically by querying the node running the second objective function for its status.

[0063] The embodiments of this application do not limit the specific handling method for the timeout of the heartbeat message for the second objective function. Optionally, if the duration between the preset time of the previously received heartbeat message for the second objective function and the current time is greater than a preset duration threshold, a function detection request for the second objective function can be sent to other nodes to which the second objective function belongs, and the second objective function can be determined to be in a running state based on the feedback result of the function detection request.

[0064] The embodiments of this application do not limit the specific method for determining the timeout of the heartbeat message for the second objective function, nor do they limit the preset time and preset duration threshold of the heartbeat message. Optionally, the preset time of the previously received heartbeat message for the second objective function can specifically be the reception time of the previously received heartbeat message for the second objective function, or it can be the transmission time of the previously received heartbeat message for the second objective function. Optionally, the heartbeat message can be sent periodically according to a preset period, and the preset duration threshold can be determined based on the duration of a single preset period, specifically it can be the sum of the durations of N preset periods. N can be a positive integer.

[0065] In this context, the other nodes belonging to the second objective function can be other nodes running the second objective function. The embodiments of this application do not limit the response method of any node in the distributed node system to a function monitoring request. Optionally, when any node in the distributed node system receives a function monitoring request for the second objective function from another node, it can determine the feedback operation based on the state of the second objective function. Specifically, if the second objective function is currently running, it can send a heartbeat message for the second objective function; if the second objective function is currently in an abnormal state, it can either not perform any operation or send back information indicating that the second objective function is currently in an abnormal state. Therefore, correspondingly, the node receiving the feedback information can determine the current state of the second objective function and thus perform appropriate processing. For example, upon receiving a heartbeat message for the second objective function, it can be determined that the second objective function is currently running, and the information of the second objective function recorded locally can be updated based on the received heartbeat message. If no heartbeat message for the second objective function is received for a long time, or if feedback information indicating that the second objective function is currently in an abnormal state is received, it can be determined that the second objective function is currently in an abnormal state, and the information of the second objective function recorded locally can be updated. In addition, the information indicating that the second objective function is currently in an abnormal state can also be synchronized in the distributed node system.

[0066] Therefore, optionally, the state of the second objective function can be determined based on whether the heartbeat message for the second objective function has timed out. Optionally, determining whether the second objective function is in a running state based on the feedback result of the function detection request can specifically include: determining that the second objective function is in an abnormal state if no heartbeat message for the second objective function is received within a preset waiting time; and determining that the second objective function is in a running state if a heartbeat message for the second objective function is received within the preset waiting time. This embodiment can determine the state of the second objective function based on whether a heartbeat message for the second objective function is received within a preset waiting time, and can improve the accuracy of the determined state of the second objective function by relying on the feedback results of other nodes to which the second objective function belongs.

[0067] The embodiments of this application do not limit the operations after determining the state of the second objective function. Optionally, the determined state of the second objective function can be synchronized in the distributed node system, specifically by sending it to other nodes. For example, information indicating that the second objective function is in an abnormal state can be sent to other nodes in the distributed node system, or the heartbeat message for the second objective function can be forwarded to other nodes in the distributed node system.

[0068] The embodiments of this application do not limit the specific method by which any node in the distributed node system sends heartbeat messages. Optionally, a synchronization heartbeat message can be sent proactively, a heartbeat message can be sent based on a request received from another node, or a heartbeat message can be sent by continuously or periodically monitoring the status of the locally running objective function.

[0069] Optionally, the above method may further include: if it is determined that the local third objective function is running, synchronizing the heartbeat message for the third objective function in the distributed node system. This embodiment allows nodes to proactively determine the status of the locally running third objective function and send heartbeat messages for synchronization, which can improve the efficiency and accuracy of objective function status synchronization.

[0070] It is understandable that if it is determined that the local third objective function is in an abnormal state or is not running, no operation may be performed, or the information indicating that the third objective function is in an abnormal state may be synchronized in the distributed node system. For details on the synchronization process, please refer to the explanations in other embodiments.

[0071] The embodiments of this application do not limit the source of the target function. Optionally, the target function can be a pre-specified function or a function determined by the distributed node system. The target function can be updated, specifically by deleting or adding a target function. For example, for any target function, the monitoring requirements can be updated to determine that it is no longer necessary to monitor the running status of a certain target function, thereby deleting the information of that target function. For any function, if it is determined that it is necessary to monitor the running status of that function, then the information of that function as a target function can be added.

[0072] Therefore, optionally, the above method flow may further include: if it is determined that any local function needs to be monitored for its running status, the determined function is identified as the fourth objective function, and the information of the fourth objective function is synchronized in the distributed node system. This embodiment can improve the comprehensiveness and flexibility of objective function monitoring by identifying the function that needs to be monitored as the new objective function.

[0073] The embodiments of this application do not limit the specific method for determining the functions whose runtime status needs to be monitored. Optionally, it can be based on business requirements to determine any function on any node that needs to have its runtime status monitored; or it can be based on the node itself to determine the functions whose runtime status needs to be monitored, specifically, the node can determine the functions whose runtime status needs to be monitored based on function requirements or function importance, so that the distributed node system can monitor the runtime status of functions locally on the node; or it can be based on the distributed node system to determine the functions whose runtime status needs to be monitored based on function call requirements. In this case, functions on any node that are called by other nodes or other devices often need to have improved stability. Therefore, functions that are frequently called by other nodes or other devices, or functions that are of high importance during the call process, can be identified as functions whose runtime status needs to be monitored, so as to facilitate monitoring of runtime status and improve the stability of function calls.

[0074] Therefore, optionally, functions whose call importance exceeds a preset importance threshold can be identified as target functions whose runtime status needs to be monitored, based on function call patterns. The importance of function calls can be determined based on the number of function calls and / or the frequency of function calls, and the importance of function calls can be positively correlated with the number of function calls and / or the frequency of function calls.

[0075] In one optional embodiment, to facilitate monitoring of the running status of the objective function in the distributed node system by each node, any node in the distributed node system can store a set of objective function information. Specifically, each node in the distributed node system can store a set of objective function information.

[0076] The embodiments of this application do not limit the format and content of the objective function information set. Optionally, the objective function information set may include information about the objective function in the distributed node system, and may also include the state information corresponding to the objective function. The state information may specifically include at least one of the following: the current state of the objective function, the preset time of the previous heartbeat message for the objective function, the heartbeat message for the objective function, etc.

[0077] The objective functions stored in the objective function information sets of different nodes can be the same or different. Optionally, the objective function information sets stored in different nodes can contain the same or different objective functions; the objective function information set stored in any node in the distributed node system can contain information about all objective functions in the distributed node system; or, the objective function information sets stored in each node in the distributed node system can contain information about all objective functions in the distributed node system.

[0078] It is understandable that, due to the information transmission situation in a distributed node system, the sets of objective function information stored in different nodes may differ. Specifically, the state information for the same objective function may be inconsistent, thus requiring synchronization operations to improve the consistency of objective function information sets between different nodes.

[0079] Of course, optionally, the objective function information set stored in each node of the distributed node system may contain information about the objective functions of each node in the distributed node system. Accordingly, the objective functions in the objective function information sets stored in different nodes should be consistent.

[0080] In one alternative embodiment, if the target function information sets of different nodes in a distributed node system are inconsistent, a synchronization operation can be performed on the target function information sets.

[0081] Therefore, optionally, any node in the distributed node system can store a set of objective function information; the above method flow may further include: if it is determined that the locally stored set of objective function information is inconsistent with the set of objective function information stored by other nodes, performing a synchronization operation on the set of objective function information stored between different nodes in the distributed node system. This embodiment can improve data consistency in the distributed node system by synchronizing the set of objective function information between different nodes.

[0082] The embodiments of this application do not limit the content of the objective function information set. Optionally, the objective function information set stored in any node may include information about the objective functions in the distributed node system.

[0083] The embodiments of this application do not limit the specific method of synchronizing the objective function information set. Optionally, the objective function information sets of different nodes in the distributed node system can be comprehensively analyzed to determine the current objective function information set. That is, for any objective function in the objective function information set, the corresponding state information is the current state information comprehensively determined by the distributed node system. Thus, the determined current objective function information set can be further synchronized to each node in the distributed node system to achieve data synchronization.

[0084] The embodiments of this application do not limit the specific circumstances and methods for determining inconsistencies in the sets of objective function information stored on different nodes. Optionally, if it is determined that the set of objective function information stored locally is inconsistent with the set of objective function information stored on any other node, a synchronization operation can be performed on the sets of objective function information stored on different nodes in the distributed node system. Alternatively, if it is determined that the sets of objective function information stored on any two nodes are inconsistent, a synchronization operation can be performed on the sets of objective function information stored on different nodes in the distributed node system. It is understood that, in cases of information inconsistency, a synchronization operation can be triggered in real time to improve data consistency in the distributed node system.

[0085] Therefore, optionally, the objective function information set includes the state information of the objective function and the corresponding state determination time information. Performing a synchronization operation on the objective function information sets stored between different nodes in the distributed node system can specifically include: for any objective function in the objective function information sets stored between different nodes in the distributed node system, determining the current state information of the targeted objective function based on the corresponding state determination time information, and synchronizing the determined current state information in the distributed node system. This embodiment can improve data consistency in the distributed node system by determining the current state information of any objective function based on the state determination time information and integrating the objective function information sets in the distributed node system, and then synchronizing it in the distributed node system.

[0086] The embodiments of this application do not limit the state determination time information; specifically, it can be the sending or receiving time of the previous heartbeat message for the corresponding objective function. It is understood that the state information of the objective function in the objective function information set can be either a running state or an abnormal state, thus enabling synchronization for objective functions in different states.

[0087] The embodiments of this application do not limit the method of determining whether the information of different nodes in a distributed node system is consistent, nor do they limit the method of determining whether the set of objective function information stored by different nodes in a distributed node system is consistent.

[0088] Alternatively, the objective function information sets stored on different nodes can be directly compared, or a hash value can be determined for the objective function information sets stored on different nodes, thereby determining whether the objective function information sets stored on different nodes are consistent based on the hash value.

[0089] Therefore, optionally, any node in the distributed node system stores a set of objective function information; the above method flow may further include: determining, based on the hash value of the objective function information set, whether the locally stored objective function information set is consistent with the objective function information sets stored by other nodes. This embodiment uses the hash value of the objective function information set for data consistency determination, which can improve the efficiency and accuracy of consistency determination for objective function information sets on different nodes.

[0090] The embodiments of this application do not limit the specific method and process for determining the hash value of the target function information set. Optionally, the hash value can be determined for the entire target function information set, or the target function information set can be grouped, and the hash value can be determined for each group. Then, by comparing the hash values ​​of the groups, it can be determined whether the data is consistent, thereby determining the data consistency at a finer granularity. Specifically, it can identify consistent group data and inconsistent group data.

[0091] The embodiments of this application do not limit the specific grouping method. Optionally, grouping can be based on the objective function, determining the information of a single objective function as a group and determining the hash value; alternatively, grouping can be based on the node to which the objective function belongs, determining the information of objective functions belonging to the same node as a group and determining the hash value; alternatively, grouping can be based on state information, specifically determining the objective functions whose state information represents the running state as one group and the objective functions whose state information represents the abnormal state as another group, and determining the hash value for each group.

[0092] Therefore, optionally, determining whether the locally stored target function information set is consistent with the target function information sets stored on other nodes based on the hash value of the target function information set can specifically include: determining whether the locally stored target function information set is consistent with the target function information sets stored on other nodes based on the hash value of a preset group in the target function information set; the preset group may contain information about target functions belonging to the same node; the above method flow may further include: when it is determined that the locally stored target function information set is inconsistent with the target function information sets stored on other nodes, performing a synchronization operation on the inconsistent preset groups between different nodes in the distributed node system. This embodiment can determine the consistency between preset groups in the target function information set by comparing the hash values ​​of preset groups, which can easily identify the cases of consistent and inconsistent preset groups, and further facilitate synchronization operations on inconsistent preset groups. For consistent groups, synchronization operations may not be performed, thereby improving the efficiency and accuracy of data consistency synchronization. It is understood that the preset group can also be determined based on other information.

[0093] The embodiments of this application are not limited to synchronization operations. Optionally, for inconsistent preset groups, the current state information of the objective function in the preset group can be determined in the distributed node system, the current preset group can be determined, and the determined current preset group can be synchronized in the distributed node system. For a more detailed explanation, please refer to the explanations of other embodiments.

[0094] In one optional embodiment, the judgment of data consistency and the data synchronization operation can be performed during the data synchronization period within the cycle of the distributed node system. During the data synchronization period, different nodes may not transmit heartbeat messages, and the current state of the objective function can remain unchanged. Thus, the current state of any objective function can be determined by combining the objective function information set between various nodes in the distributed node system, which is used to achieve data synchronization and improve data consistency.

[0095] It is understandable that for any objective function monitored in a distributed node system, the running status of the objective function can be determined through any node in the distributed node system. If any node crashes or malfunctions, the running status of the objective function can be determined through any other node, thereby improving the stability of determining the function status.

[0096] The embodiments of this application do not limit the subsequent operations based on the determined state of the target function. Optionally, it can be determined whether the target function needs to be called or whether services related to the target function need to be used based on the determined state of the target function. Specifically, for example, if the target function is determined to be running based on the distributed node system, the target function can be called or services related to the target function can be used; if the target function is determined to be in an abnormal state based on the distributed node system, calling the target function or using services related to the target function can be stopped.

[0097] For ease of understanding, in a specific example, the objective function could be a function that provides database access functionality, allowing the execution status of the objective function to be monitored through a distributed node system.

[0098] The objective function can run on a node in the distributed node system. Once the node determines that the objective function is currently running, it can synchronously send a heartbeat message for the objective function in the distributed node system, thereby synchronizing the information about the current running status of the objective function in the distributed node system.

[0099] For other devices outside the distributed node system, or other nodes within the distributed node system, the target function can be invoked to access the database based on the information about its current running status.

[0100] If the target function malfunctions, the node to which it belongs can stop sending heartbeat messages once it is determined that the target function is currently in an abnormal state.

[0101] Other nodes in the distributed node system can detect when the heartbeat message of the target function times out. That is, when the time elapsed since the last heartbeat message of the target function exceeds a threshold, they can directly determine that the target function is currently in an abnormal state, or they can initiate a function detection request for the target function.

[0102] When the node to which the objective function belongs receives a function detection request for that objective function, it can provide feedback based on the current state of the objective function. Specifically, if the objective function is currently running (or has recovered to running), it can send a heartbeat message for that objective function. If the objective function is currently in an abnormal state (or has stopped running), it may not send feedback or may send information indicating that the objective function is in an abnormal state.

[0103] For distributed node systems, the state of the objective function can be determined based on feedback information and synchronized, making it easy to determine whether the objective function can continue to be called to access data.

[0104] For ease of understanding, this application also provides an application embodiment.

[0105] As services continue to expand, the stability of batch task execution becomes increasingly important. Some functions and methods that require long-term operation need effective monitoring. Current technical solutions primarily focus on the overall survival of application instances. Therefore, this embodiment proposes a method capable of monitoring the survival of functions within an application, meeting the need for more granular monitoring.

[0106] This embodiment can establish a fine-grained monitoring chain to perceive the survival status (equivalent to running status) of individual long-running functions (such as batch tasks, asynchronous threads, etc.) within the application, thereby improving the granularity of monitoring. Furthermore, a decentralized availability monitoring mechanism can be implemented through a consensus mechanism.

[0107] This embodiment proposes a method for monitoring and detecting the availability of functions at the function level. Using this method, it is possible to monitor whether functions are alive during program execution, thereby improving the granularity of monitoring.

[0108] This embodiment includes: a monitoring and registration module, a heartbeat sending module, and a liveness detection module. It is understood that different servers or different nodes can each contain these three modules to achieve distributed information synchronization and status monitoring.

[0109] The monitoring and registration module is a service that registers function execution information into the monitoring information table.

[0110] The heartbeat sending module is a service that periodically updates the liveness information of functions to the monitoring information table.

[0111] The liveness detection module will periodically check whether the heartbeat update time of the functions in the monitoring information meets the maximum duration limit.

[0112] First is the monitoring registration module. This module can be a function, with the input parameter being a monitoring type enumeration class. The enumeration class determines the timeout for the corresponding monitoring type. When the function runs, it calls the monitoring registration module. The monitoring module obtains information such as the current system's network address, service instance name, server canary deployment, and service cluster by calling system functions. It then registers this information in the memory of the distributed registry center, storing it in the following structure: service address, function availability monitoring type, last liveness probe upload time, function liveness flag, canary deployment flag, monitoring record start time, and monitoring record last update time.

[0113] This module provides two ways to register availability information, which can be implemented through direct function calls or through aspect injection. (1) When the function is called, the business code needs to be modified and the availability type enumeration is passed in. The advantage is that the registration method is more flexible, and the monitoring registration information is only sent when certain judgment logic is met. (2) When the function is called, the monitoring information is automatically registered without any additional operation.

[0114] Heartbeat sending module. After the monitoring information registration is completed, this module is responsible for periodically sending the function running status monitoring to the monitoring information table. This module is the core module for monitoring whether the function is running normally. This method will automatically obtain the identifier of the running instance of the function, the monitoring type and other information through the system function, and send a broadcast message to the liveness detection module of the nearby nodes (selected by random algorithm and other methods). The broadcast method is as follows: 1. The sending node sends a dictionary table to the nearby nodes, which includes the node identifier and the registration information of all functions corresponding to the node identifier (the hash value of the function name and the last heartbeat sending time). 2. The receiving node obtains the list and compares whether the hash value of the registration information is consistent. If they are inconsistent, it requests the full registration information of the node identifier from the sending node. 3. The sending node accepts the request and sends the full registration information of a certain node identifier to the receiving node. This method can also be implemented by direct function call and aspect embedding. (1) Function call, by modifying the business code, the heartbeat information is sent directly. The sending method (time interval, whether to send, etc.) can be flexibly configured. (2) When the function is called, a child thread attached to the main thread is automatically started, and the child thread is closed when the main thread ends. The heartbeat sending interval of this child thread can be configured as required.

[0115] Liveness detection module. (1) Service heartbeat timeout detection. First, the liveness detection module will periodically receive heartbeat information from nearby nodes and compare the information with the existing heartbeat information in memory, and determine the priority of data coverage according to the order of update time. Then, the system will initialize the service availability query model, obtain monitoring records and current time from memory according to the instance type, and provide basic data for subsequent processing. Subsequently, the system will filter instances with heartbeat timeout in real time through streaming processing, specifically including calculating the interval between the most recent heartbeat time and the current time, and comparing it with the upper limit threshold allowed for the service type, so as to accurately identify abnormal nodes. When a function heartbeat timeout is detected, a request will be actively initiated to the server where the function is located. If the server does not return heartbeat information or returns heartbeat information that the function has timed out, the function will be recorded as timed out in memory and broadcast to other servers. (2) Message upload and call unregistration. For the detected heartbeat timeout instance, the system will perform a series of automated processing operations. If the instance is found to be down or invalid, the monitoring information status will be automatically marked as invalid in memory. At the same time, the system will construct detailed information, including key data such as service identifier, cluster type and offline duration, and send monitoring information asynchronously to ensure that the information is deduplicated at the hourly granular level. (3) Method for launching the liveness detection module. The liveness detection task is launched periodically. This can be achieved by setting a timed task or by launching it periodically in other ways. This module does not require modification of the original system code. It can be implemented by adding an additional code module. (4) User actively queries availability information. Users can access any node to know the current availability status.

[0116] This embodiment proposes a service governance scheme based on fine-grained function-level monitoring. By constructing a function operation monitoring system in a distributed environment, the following technical advantages are achieved: (1) Breaking through the traditional service monitoring granularity, it realizes real-time collection and anomaly detection of function-level operation status, and can accurately locate the execution status of a single function, effectively solving the problem of insufficient perception of the internal operation status of code in traditional schemes. (2) Adopting an intelligent monitoring resource recycling mechanism, it automatically cleans up the monitoring information of failed functions, avoids resource leakage caused by failed instances, and ensures the optimization of system resource utilization. (3) Based on aspect-oriented programming, it realizes function monitoring. Business systems can automatically complete the collection and reporting of function operation data by adding simple annotations, without modifying business logic code, which greatly reduces the system transformation cost. (4) Through a distributed consensus algorithm, it reduces the potential single point of failure risk of recording in the database.

[0117] Corresponding to the above method embodiments, this application also provides a function monitoring device. The following will be combined with... Figure 3 The device is described in detail.

[0118] Figure 3 The schematic diagram illustrates a structural block diagram of a function monitoring device according to an embodiment of this application.

[0119] like Figure 3 As shown, this embodiment provides a function monitoring device 300, which includes a first module 310 and a second module 320. This device can be applied to any node in a distributed node system and can be used to monitor the running status of a target function in the distributed node system.

[0120] The first module 310 is used to determine that the first objective function is in a running state upon receiving a heartbeat message for the first objective function in other nodes. In one embodiment, the first module 310 can be used to execute the operation S210 and related steps described above, which will not be repeated here.

[0121] The second module 320 is configured to, for a second objective function in other nodes, send a function detection request for the second objective function to the other nodes to which the second objective function belongs if the duration between the preset time of the previously received heartbeat message for the second objective function and the current time is greater than a preset duration threshold, and determine whether the second objective function is in a running state based on the feedback result of the function detection request. In one embodiment, the second module 320 may be used to execute the operation S220 and related steps described above, which will not be repeated here.

[0122] Optionally, the second module 320 is used to: determine that the second objective function is in an abnormal state if no heartbeat message for the second objective function is received within a preset waiting time; and determine that the second objective function is in a running state if a heartbeat message for the second objective function is received within a preset waiting time.

[0123] Optionally, the above apparatus further includes a third module for: synchronizing heartbeat messages for the third objective function in the distributed node system when it is determined that the local third objective function is in a running state.

[0124] Optionally, the above-mentioned device further includes a fourth module, used to: determine the determined function as the fourth objective function when it is determined that any local function needs to monitor its running status, and synchronize the information of the fourth objective function in the distributed node system.

[0125] Optionally, any node in the distributed node system stores a set of objective function information; the above device further includes a synchronization module, used to: perform a synchronization operation on the objective function information sets stored between different nodes in the distributed node system when it is determined that the locally stored objective function information set is inconsistent with the objective function information sets stored by other nodes.

[0126] Optionally, the objective function information set includes the state information of the objective function and the corresponding state determination time information; the synchronization module is used to: for any objective function in the objective function information set stored between different nodes in the distributed node system, determine the current state information of the target objective function based on the corresponding state determination time information, and synchronize the determined current state information in the distributed node system.

[0127] Optionally, any node in the distributed node system stores a set of objective function information; the synchronization module is also used to: determine, based on the hash value of the objective function information set, whether the locally stored objective function information set is consistent with the objective function information sets stored by other nodes.

[0128] Optionally, the synchronization module is used to: determine whether the locally stored target function information set is consistent with the target function information sets stored by other nodes based on the hash value of the preset group in the target function information set; the preset group contains information about target functions belonging to the same node; the synchronization module is also used to: perform synchronization operation on the preset group that is inconsistent between different nodes in the distributed node system when it is determined that the locally stored target function information set is inconsistent with the target function information sets stored by other nodes.

[0129] According to embodiments of this application, any plurality of modules among the first module 310, second module 320, third module, fourth module, and synchronization module can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the first module 310, second module 320, third module, fourth module, and synchronization module can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented by any other reasonable means of integrating or packaging circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these. Alternatively, at least one of the first module 310, second module 320, third module, fourth module, and synchronization module can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0130] For an explanation of this device embodiment, please refer to other embodiments. Each embodiment in the above method embodiment can be executed by the corresponding module in this device embodiment.

[0131] Figure 4A block diagram schematically illustrates an electronic device suitable for implementing a function monitoring method according to an embodiment of this application.

[0132] like Figure 4 As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0133] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that the programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0134] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0135] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0136] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.

[0137] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement any of the method embodiments provided in the embodiments of this application.

[0138] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0139] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0140] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0141] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0143] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A function monitoring method, characterized in that, Applied to any node in a distributed node system, the method is used to monitor the running status of an objective function in the distributed node system; the method includes: Upon receiving a heartbeat message for the first objective function in other nodes, it is determined that the first objective function is in a running state; For the second target function in other nodes, if the duration between the preset time of the previously received heartbeat message for the second target function and the current time is greater than a preset duration threshold, a function detection request for the second target function is sent to the other nodes to which the second target function belongs, and the second target function is determined to be in a running state based on the feedback result of the function detection request.

2. The method according to claim 1, characterized in that, Determining whether the second target function is in a running state based on the feedback result of the function detection request includes: If no heartbeat message for the second objective function is received within the preset waiting time, it is determined that the second objective function is in an abnormal state; If a heartbeat message for the second objective function is received within a preset waiting time, it is determined that the second objective function is in a running state.

3. The method according to claim 1, characterized in that, The method further includes: If it is determined that the local third objective function is running, the heartbeat message for the third objective function will be synchronized in the distributed node system.

4. The method according to claim 1 or 3, characterized in that, The method further includes: If it is determined that any local function needs to monitor its running status, the determined function is designated as the fourth objective function, and the information of the fourth objective function is synchronized in the distributed node system.

5. The method according to claim 1, characterized in that, Each node in the distributed node system stores a set of objective function information; the method further includes: If the target function information set stored locally is inconsistent with the target function information set stored on other nodes, a synchronization operation is performed on the target function information sets stored on different nodes in the distributed node system.

6. The method according to claim 5, characterized in that, The objective function information set includes the state information of the objective function and the corresponding state determination time information; The synchronization operation performed on the set of objective function information stored between different nodes in the distributed node system includes: For any objective function in the set of objective function information stored between different nodes in the distributed node system, the current state information of the objective function is determined based on the corresponding state determination time information, and the determined current state information is synchronized in the distributed node system.

7. The method according to claim 1, characterized in that, Each node in the distributed node system stores a set of objective function information; the method further includes: Based on the hash value of the objective function information set, determine whether the objective function information set stored locally is consistent with the objective function information sets stored on other nodes.

8. The method according to claim 7, characterized in that, The step of determining whether the locally stored objective function information set is consistent with the objective function information sets stored on other nodes, based on the hash value of the objective function information set, includes: Based on the hash value of a preset group in the objective function information set, it is determined whether the locally stored objective function information set is consistent with the objective function information sets stored by other nodes; the preset group contains information about objective functions belonging to the same node; The method further includes: If the target function information set stored locally is inconsistent with the target function information set stored on other nodes, a synchronization operation is performed on the pre-defined groups that are inconsistent between different nodes in the distributed node system.

9. A function monitoring device, characterized in that, Applied to any node in a distributed node system, the device is used to monitor the running status of the objective function in the distributed node system; the device includes: The first module is used to determine that the first objective function is in a running state when a heartbeat message for the first objective function in other nodes is received; The second module is used to send a function detection request for the second target function to other nodes to which the second target function belongs if the duration between the preset time of the previously received heartbeat message for the second target function and the current time is greater than a preset duration threshold, and to determine whether the second target function is in a running state based on the feedback result of the function detection request.

10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.