A software monitoring method and apparatus
Patent Information
- Application Number
- CN202311445091.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2043-11-01
AI Technical Summary
但是智慧机房系统主要是面向硬件设备,无法监视DCS平台软件的运行状态,亟需一种对NC-DCS平台软件监视的方法及装置
[0008] The above technical solution can obtain software failure modes characterizing NC-DCS platform software failures. It identifies software health parameters associated with these failure modes from all software parameters involved in the operation of the NC-DCS platform software, uses these parameters as the platform's health parameters, collects their values, and analyzes these values to determine the platform's operational status. Software failure modes characterize NC-DCS platform software failures, and associated health parameters can determine whether the platform has failed. Furthermore, when a failure occurs, the cause can be located and investigated, thus enabling monitoring of the NC-DCS platform's operational status, reducing the difficulty of failure localization, and shortening the maintenance cycle.
Smart Images

Figure CN117389832B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of product control technology, and in particular relates to a software monitoring method and apparatus. Background Technology
[0002] The stability and reliability of the Non-Safety Digital Instrument Control System (NC-DCS) platform software are crucial for nuclear power safety. Currently, equipment nodes integrating the NC-DCS platform software can be monitored through smart data center systems. These systems employ robots to inspect the locations of equipment nodes, capturing video footage from cameras on the robots. This video is then processed by the smart data center system's internal video processing model to obtain parameters characterizing the operational status of the equipment nodes, such as instrument data and switch status (e.g., transformer oil level gauge readings, disconnector switch open / closed status). However, smart data center systems primarily focus on hardware and cannot monitor the operational status of the DCS platform software. Therefore, a method and device for monitoring the NC-DCS platform software are urgently needed. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a software monitoring method and apparatus for monitoring the operating status of NC-DCS platform software. The technical solution is as follows:
[0004] In a first aspect, this application provides a software monitoring method, the method comprising: obtaining a failure mode list of a non-safety-grade digital instrumentation and control system platform software, wherein the failure mode list records software failure modes characterizing the failure of the non-safety-grade digital instrumentation and control system platform software, the software failure modes being obtained by analyzing at least one of the following: software failure modes of each level unit, on-site problems occurring when running the non-safety-grade digital instrumentation and control system platform software, and historical defect data of the non-safety-grade digital instrumentation and control system platform software; obtaining all software parameters involved in the operation of the non-safety-grade digital instrumentation and control system platform software; determining software health parameters associated with the software failure modes from all the software parameters; wherein the software health parameters associated with the software failure modes are used as software health parameters of the non-safety-grade digital instrumentation and control system platform software, and the operating status of the non-safety-grade digital instrumentation and control system platform software is determined by analyzing the parameter values of the software health parameters.
[0005] Secondly, this application provides a software operation and maintenance monitoring module, the module comprising: a list acquisition unit, used to acquire a failure mode list of non-safety-grade digital instrumentation and control system platform software, wherein the failure mode list records software failure modes characterizing the failure of the non-safety-grade digital instrumentation and control system platform software, the software failure modes being obtained by analyzing at least one of the following: software failure modes of each level unit, on-site problems occurring when running the non-safety-grade digital instrumentation and control system platform software, and historical defect data of the non-safety-grade digital instrumentation and control system platform software; each level unit including device nodes, tasks, threads, and functions involved in the operation of the non-safety-grade digital instrumentation and control system platform software; a parameter acquisition unit, used to acquire all software parameters involved in the operation of the non-safety-grade digital instrumentation and control system platform software; and a determination unit, used to determine software health parameters associated with the software failure modes from all the software parameters; wherein the software health parameters associated with the software failure modes are used as software health parameters of the non-safety-grade digital instrumentation and control system platform software, so as to determine the operating status of the non-safety-grade digital instrumentation and control system platform software by analyzing the parameter values of the software health parameters.
[0006] Thirdly, this application provides a computer-readable storage medium storing computer program code, which implements the above-described software monitoring method when the computer program code is executed.
[0007] Compared with the prior art, the technical solution provided in this application has the following advantages:
[0008] The above technical solution can obtain software failure modes characterizing NC-DCS platform software failures. It identifies software health parameters associated with these failure modes from all software parameters involved in the operation of the NC-DCS platform software, uses these parameters as the platform's health parameters, collects their values, and analyzes these values to determine the platform's operational status. Software failure modes characterize NC-DCS platform software failures, and associated health parameters can determine whether the platform has failed. Furthermore, when a failure occurs, the cause can be located and investigated, thus enabling monitoring of the NC-DCS platform's operational status, reducing the difficulty of failure localization, and shortening the maintenance cycle. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is an optional flowchart of the software monitoring method provided in the embodiments of this application;
[0011] Figure 2 This is a schematic diagram illustrating the determination of software failure modes provided in an embodiment of this application;
[0012] Figure 3 This is another optional flowchart of the software monitoring method provided in the embodiments of this application;
[0013] Figure 4 This is a schematic diagram of the deployment software operation and maintenance monitoring module provided in an embodiment of this application;
[0014] Figure 5 This is a schematic diagram of the design acquisition and storage strategies provided in the embodiments of this application;
[0015] Figure 6 This is a schematic diagram of an optional structure of the software operation and maintenance monitoring module provided in this application embodiment;
[0016] Figure 7 This is a schematic diagram of another optional structure of the software operation and maintenance monitoring module provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] Currently, the monitoring of NC-DCS platform software mainly focuses on monitoring the device nodes that integrate the NC-DCS platform software, such as monitoring the instrument data and switch status of the device nodes. Therefore, the current monitoring is mainly from the perspective of hardware devices. Monitoring from the perspective of hardware devices cannot characterize the operating status of the NC-DCS platform software. The operating status of the NC-DCS platform software can characterize whether the NC-DCS platform software has failed.
[0019] Typically, monitoring device nodes yields logs and dump files (memory images of processes). Engineers can examine the contents of these logs and dump files to determine if the NC-DCS platform software has failed and manually locate the failure. However, this process presents challenges in locating software failures and results in long maintenance cycles.
[0020] Some embodiments of this application provide a software monitoring method and apparatus. This method and apparatus can obtain software failure modes characterizing NC-DCS platform software failures, determine software health parameters associated with the software failure modes from all software parameters involved in the NC-DCS platform software, and use these associated software health parameters as the software health parameters of the NC-DCS platform software. The parameter values of the software health parameters are collected during NC-DCS platform software operation (including before operation), and the operating status of the NC-DCS platform software can be determined through these parameter values. Because the software failure mode is a failure mode of the NC-DCS platform software, the software health parameters can be parameters characterizing NC-DCS platform software failures. By using the software health parameters to determine whether the NC-DCS platform software has failed, and by using the software health parameters to locate and troubleshoot the cause of the failure, the operating status of the NC-DCS platform software can be monitored, reducing the difficulty of failure localization and shortening the maintenance cycle.
[0021] Please see Figure 1 This illustrates an optional flow of the software monitoring method provided in this application embodiment, which may include the following steps:
[0022] S101. Obtain the failure mode list of the NC-DCS platform software. The failure mode list records the software failure modes that characterize the failure of the NC-DCS platform software. The software failure modes are obtained by analyzing at least one of the following: the software failure modes of each level unit, the field problems that occur when running the NC-DCS platform software, and the historical defect data of the NC-DCS platform software. Each level unit includes the device nodes (or deployment nodes), tasks, threads, and functions involved in the operation of the NC-DCS platform software.
[0023] Among them, the device nodes involved in the operation of NC-DCS platform software can be devices that have deployed NC-DCS platform software, such as servers, operator stations, engineer stations, maintenance stations, and gateways that have deployed NC-DCS platform software; tasks can be tasks in NC-DCS platform software and / or tasks called by NC-DCS platform software; threads can be threads used during task execution, and functions can be functions used during thread execution, so as to analyze the software failure modes that characterize the failure of NC-DCS platform software from the hierarchical analysis of device nodes, tasks, threads, and functions.
[0024] Whether it's the software failure modes at each level of the unit, or the field problems that occur when running the NC-DCS platform software and the historical defects of the NC-DCS platform software, this data contains software failure modes that may characterize NC-DCS platform software failure. Analyzing this data helps identify the software failure modes that characterize NC-DCS platform software failure. Therefore, in some examples, the process of obtaining software failure modes is as follows:
[0025] The software failure modes of each level unit are obtained through forward analysis based on software design to identify the software failure modes that characterize the NC-DCS platform software failure; the field problems that occur when running the NC-DCS platform software are analyzed to identify the software failure modes that characterize the NC-DCS platform software failure from the field problems; the historical defect data of the NC-DCS platform software is analyzed to identify the software failure modes that characterize the NC-DCS platform software failure from the historical defect data, so as to analyze from multiple perspectives and improve the software failure modes as much as possible.
[0026] The software failure modes of each unit level are analyzed in a top-down order: device node -> task -> thread -> function. That is, the software failure modes of device nodes are analyzed first, followed by task failure modes, then thread failure modes, and finally function failure modes. From multiple software failure modes at different unit levels, the software failure modes characterizing the NC-DCS platform software failure are identified. The resulting failure mode list includes device-level failure modes, task-level failure modes, thread-level failure modes, and function-level failure modes, which can characterize the NC-DCS platform software failure.
[0027] The forward analysis based on software design derives software failure modes for each level of unit by inferring failure modes from a software design perspective. This perspective includes software runtime environment design, general software function design, and software-specific function design. General software function design refers to the general functional design of the NC-DCS platform software, while software-specific function design refers to the unique functional design of NC-DCS. By inferring software failure modes characterizing NC-DCS failures from multiple design perspectives, the inferred software failure modes are categorized into runtime environment failure modes, general function failure modes, and software-specific function failure modes. Runtime environment failure modes include operating system environment errors, IP address configuration errors, dependent task configuration errors, and insufficient runtime memory resources. General function failure modes include communication anomalies and load anomalies. Software-specific function failure modes are analyzed based on specific software functions; for example, software-specific function failure modes include computational function anomalies, real-time data function anomalies, and procedural function anomalies. Each type of software failure mode can be included in the failure mode list.
[0028] Field problems refer to actual software failures discovered during the operation of the NC-DCS platform software. Fault Tree Analysis (FTA) can be used to identify software failure modes characterizing these failures. Some field problems have low severity and occur infrequently, making them unlikely to cause NC-DCS platform software failure. To reduce the number of field problems and accelerate analysis efficiency, field problems are screened to select those suitable for FTA analysis, from which software failure modes are identified. For example, selected field problems can be frequent and significant; frequent means occurring many times, and significant means highly severe. In this embodiment, these can be determined by frequency and severity.
[0029] Fault tree analysis is based on constructing a software fault tree based on field problems. The field problem is the top event, which is decomposed into multiple intermediate events, and then further decomposed into basic events. These events represent various factors that may cause the software to fail due to the field problem. Each event chain may lead to the occurrence of the top event. For example, the event chain includes engineering configuration abnormalities, human-machine interface abnormalities, clock synchronization abnormalities, etc. These event chains are summarized as software failure modes.
[0030] Historical defect data for the NC-DCS platform software contains information about defects present in the software. This data can be obtained during the NC-DCS platform software development process. These defects cause the NC-DCS platform software to malfunction, such as runtime tasks, threads, and functions failing. Based on this historical defect data, logs and dump files corresponding to the historical defect data are obtained at the software code level. From these logs and dump files, software failure modes characterizing NC-DCS failures are derived, such as software parameter validation failures, thread initialization failures, and external interface call failures.
[0031] S102. Obtain all software parameters involved in the runtime of the NC-DCS platform software. These parameters include the NC-DCS platform software's metrics, configuration parameters, logs, and dump files.
[0032] Among them, the indicator parameters are used to provide feedback on the operating status of the device nodes running the NC-DCS platform software. For example, the indicator parameters may include server heartbeat, network load, central processing unit (CPU) load, and memory load, so as to characterize the changes in the operating status of the device nodes caused by the software during the operation of the NC-DCS platform software.
[0033] Configuration parameters are used to locate failures caused by configuration errors in the NC-DCS platform software. For example, configuration parameters include environment configuration parameters and service configuration parameters. Environment configuration parameters are used to check for failures caused by environment configuration problems and runtime environment errors, while service configuration parameters are used to check for failures caused by service configuration problems.
[0034] Logs are used to record the operational information of the NC-DCS platform software, and the location of failures can be traced by analyzing the operational records in the logs. Dump files are used to reconstruct the state of the NC-DCS platform software at the time of failure, allowing for quick location and troubleshooting of the failure.
[0035] S103. Identify software health parameters that are associated with software failure modes from all software parameters.
[0036] In this embodiment, determining the software health parameters associated with the software failure mode can be achieved by analyzing the software failure mode to establish a correspondence between the software failure mode and software parameters. The software parameters in this correspondence can be those characterizing the NC-DCS platform software as being in that software failure mode. Therefore, the software parameters in this correspondence are the software health parameters associated with the software failure mode.
[0037] When identifying software health parameters associated with software failure modes, the existing software parameters in the NC-DCS platform software can be reviewed in the order of device node -> task -> thread -> function. Utilizing the inherent characteristics of the software parameters, they can be categorized into environment parameters, common characteristic parameters, specific characteristic parameters, thread parameters, and function parameters. Within the environment, common characteristic, and specific characteristic parameters, software health parameters associated with device-level and task-level failure modes are identified. Similarly, within the thread and function parameters, software health parameters associated with thread-level and function-level failure modes are identified.
[0038] Figure 2 A schematic diagram of the software health parameter design roadmap is presented. This roadmap includes failure mode analysis (FMA) routes for each level of unit and failure detection parameter design routes for each level of unit. The FMA routes for each level of unit are used to obtain the software failure modes in the failure mode list. This mainly includes forward analysis based on software design, reverse analysis based on field problems, and supplementation of software failure modes based on historical defects. Forward analysis based on software design involves systematically reviewing software failure modes at each level of unit, identifying software failure modes characterizing NC-DCS platform software failures from multiple failure modes at different levels to obtain the failure mode list. Reverse analysis based on field problems involves performing FAT analysis on field problems to supplement the software failure modes. Supplementation of software failure modes based on historical defect data further identifies software failure modes characterizing NC-DCS platform software failures from historical defect data, supplementing the software failure modes. The specific process is detailed in step S101.
[0039] The failure detection parameter design route for each level of unit is used to obtain software health parameters that are associated with software failure modes. The process can be as follows: First, sort out the existing software parameters and classify them. Then, sort out the association between the existing software parameters and software failure modes. The main purpose is to establish the correspondence between software failure modes and each type of software parameter. Use this correspondence to obtain the software health parameters that are associated with software failure modes. For details, please refer to step S103.
[0040] Because software failure modes characterize the failure patterns of NC-DCS platform software, software health parameters associated with these failure modes can also characterize NC-DCS platform software failures. These software health parameters, collected during NC-DCS platform software runtime, can be used to determine the platform's operational status. For example, the values of these parameters can be used to determine if the NC-DCS platform software has failed, and if a failure occurs, the cause can be located and investigated.
[0041] Figure 3 The illustration shows another optional flow of the software monitoring method provided in the embodiments of this application, which may include the following steps:
[0042] S101. Obtain the failure mode list of the NC-DCS platform software. The failure mode list records the software failure modes that characterize the failure of the NC-DCS platform software. The software failure modes are obtained by analyzing at least one of the following: the software failure modes of each level unit, the field problems that occur when running the NC-DCS platform software, and the historical defect data of the NC-DCS platform software. Each level unit includes the device nodes (or deployment nodes), tasks, threads, and functions involved in the operation of the NC-DCS platform software.
[0043] S102. Obtain all software parameters involved in the runtime of the NC-DCS platform software.
[0044] S103. Identify software health parameters that are associated with software failure modes from all software parameters.
[0045] Software health parameters associated with software failure modes include metric parameters, configuration parameters, logs, and dump files. Configuration parameters include environment configuration parameters and service configuration parameters. These parameters are collected by the software operation and maintenance monitoring module (SpeAgent), which is deployed on device nodes running the NC-DCS platform software. The module is designed to be non-intrusive and deployed as a software plug-in on the device. It is used to collect and store these software health parameters. The module transmits the software health parameters associated with each software failure mode through a separate Health System Network (HSNET), while other data from the device nodes is transmitted through the Management Network (MNET). Figure 4As shown, a software operation and maintenance monitoring module is deployed on operator stations, engineer stations, maintenance stations, gateways, and various servers. Because the software operation and maintenance monitoring module is designed to be non-intrusive and deployed as a software plug-in on the device, it does not affect the functionality of the original device nodes. Furthermore, the software operation and maintenance monitoring module uses a dedicated HSNET to transmit software health parameters associated with each software failure mode, separate from the original management network used by the device nodes, thus not affecting the data transmission of the original device nodes.
[0046] When designing the software operation and maintenance monitoring module, the sensing requirements for different types of software parameters are considered. Different collection strategies, storage strategies, and storage paths are set for environment configuration parameters, indicator parameters, and service configuration parameters, such as... Figure 5 As shown. Furthermore, the collection and saving strategies for logs and dump files differ from those of other types of software parameters. For details, please refer to steps S104 to S106.
[0047] S104. Before starting the NC-DCS platform software, collect and save environment configuration parameters. Environment configuration parameters may include operating system version information, software configuration information, and network communication configuration information, which can effectively reduce NC-DCS platform software failures caused by environment configuration problems and operating environment issues. Furthermore, when the NC-DCS platform software fails, the environment configuration parameters can be used to determine whether the failure was caused by these parameters. In this embodiment, the environment configuration parameters can be collected and saved in the environment configuration parameter folder before the NC-DCS platform software starts, and the environment configuration parameters are only saved once before the NC-DCS platform software starts, reducing the amount of data saved.
[0048] S105. During the operation of the NC-DCS platform software, collect and save indicator parameters and service configuration parameters.
[0049] In some examples, the metric parameters can include device node-level and process-level metrics. Device node-level metrics primarily reflect the operating status of the device, such as network load, CPU load, and memory load. These metrics reflect the impact of the NC-DCS platform software on the device node. Process-level metrics primarily reflect the operating status of each process within the NC-DCS platform software. These metrics can include process network load, CPU load, memory load, process handles, and process Graphics Device Interface (GDI).
[0050] Service configuration parameters are mainly used to check for failures caused by service configuration issues. Service configuration parameters can include the process name, process identifier (PID), service start list, and service stop list. The service start list and service stop list are used to provide feedback on the start and stop status of each service in the NC-DCS platform software.
[0051] The collection and storage strategies for indicator parameters and service configuration parameters are as follows:
[0052] When using the NC-DCS platform software, the indicator parameters are collected according to the preset collection cycle. The indicator parameters are saved once after each collection and can be saved in the indicator parameter folder. It is determined whether the indicator parameters meet the preset cycle setting conditions. If it is determined that the indicator parameters do not meet the preset cycle setting conditions, the preset collection cycle is shortened and the collection frequency is increased. If it is determined that the indicator parameters meet the preset cycle setting conditions, the preset collection cycle is maintained unchanged.
[0053] The preset period setting conditions can include the parameter values (which can be a range or a single value) of the indicator parameters when collecting and saving the indicator parameters according to the preset collection period. For example, setting the threshold for each parameter in the device node-level indicator parameters, collecting data according to the preset collection period when the parameter value of a certain parameter is less than or equal to its corresponding threshold, and reducing the preset collection period when the parameter value of a certain parameter is greater than its corresponding threshold. For example, when the CPU load is greater than its corresponding threshold, the preset collection period is changed from 5s / time to 0.5s / time to increase the collection frequency and capture instantaneous abnormal increases in resource usage. This embodiment does not limit the value of the preset collection period or how to reduce the preset collection period.
[0054] When the NC-DCS platform software is running, if the collected service configuration parameters do not meet the preset saving conditions, the service configuration parameters are saved once and can be stored in the service configuration parameter folder. The preset saving conditions can be a triggered saving mechanism. The service configuration parameters are recorded in the preset saving conditions. If the collected service configuration parameters are different from those recorded in the preset saving conditions, the currently collected service configuration parameters are saved to the service configuration parameter folder, and the service configuration parameters recorded in the preset saving conditions are updated to the currently collected service configuration parameters.
[0055] S106. During the operation of the NC-DCS platform software, logs associated with software failure modes are retrieved and saved online from the NC-DCS platform software's log library, and dump files associated with software failure modes are retrieved and saved online from the NC-DCS platform software's dump file library. The logs and dump files associated with software failure modes can be retrieved periodically, and the periods for the logs and dump files can be the same or different.
[0056] In this embodiment, different types of software parameters employ different acquisition strategies, storage strategies, and storage paths to meet the needs of different types of software parameters. Furthermore, the software operation and maintenance monitoring module can be designed with a configuration library, allowing operators to modify the monitoring objects (such as selecting monitoring objects from the aforementioned software parameters), periods, thresholds, etc., thereby changing the granularity of software monitoring to suit different operating environments.
[0057] In summary, the software monitoring method provided in this application has the following advantages:
[0058] 1. Software monitoring methods can collect metrics, configuration parameters, logs, and dump files related to software failure modes. This allows for the collection and aggregation of various software parameters across all device nodes running the NC-DCS platform software. The real-time operational status of the NC-DCS platform software on each device is obtained, enabling effective integration and unified display of various software parameters. After obtaining this data, maintenance engineers can analyze the NC-DCS platform software running on different device nodes to achieve failure early warning, alerts, location, and provide software repair solutions. This approach monitors the operational status of the NC-DCS platform software while reducing the difficulty of failure location, improving operational efficiency, and thus shortening the maintenance cycle.
[0059] Taking memory load in process-level metrics as an example, when a process's memory usage reaches a threshold, a process memory growth alarm is triggered. Operations engineers can then locate the problem of process-level software resource usage and detect and handle it in a timely manner before the process crashes by actively restarting and releasing resources.
[0060] In general, the software failure mode of the lower-level unit may be propagated to the upper-level unit, i.e., function-level failure mode -> thread-level failure mode -> task-level failure mode -> device node-level failure mode. The failure is located by the propagation relationship between the units.
[0061] 2. A software health parameter design roadmap was designed. Through this roadmap, software health parameters that are related to software failure modes were analyzed. This allowed for in-depth collection of software health parameters related to software failure modes during software operation. These parameters can serve both production and operation and maintenance. By "governing" this data, the operation and maintenance quality and efficiency of the NC-DCS platform software were greatly improved.
[0062] 3. Online operation and maintenance of NC-DCS platform software can be achieved by using the collected health software parameters that are associated with software failure modes; using software health parameters that are associated with software failure modes as a data basis for predictive research on software failure can significantly save operation and maintenance costs and bring considerable economic benefits.
[0063] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0064] This application also provides a software operation and maintenance monitoring module, the optional structure of which is as follows: Figure 6 As shown, it may include: a list acquisition unit 10, a parameter acquisition unit 20, and a determination unit 30.
[0065] The list acquisition unit 10 is used to obtain a list of failure modes of the NC-DCS platform software. The failure mode list records the software failure modes that characterize the failure of the NC-DCS platform software. The software failure modes are obtained by analyzing at least one of the following: the software failure modes of each level unit, the field problems that occur when running the NC-DCS platform software, and the historical defect data of the NC-DCS platform software. Each level unit includes the device nodes, tasks, threads, and functions involved in the operation of the NC-DCS platform software.
[0066] In some examples, the list obtains Unit 10, which is specifically used to infer the software failure modes of each level of unit from a software design perspective, in order to identify the software failure modes that characterize the failure of the NC-DCS platform software. The software design perspective includes software operating environment design, software general function design, and software special function design. The on-site problem is taken as the top event of the software fault tree. The top event is decomposed into multiple intermediate events and each intermediate event is decomposed into multiple basic events. An event chain leading to the occurrence of the top event is formed from the top event to the basic events. The event chain is summarized as the software failure mode that characterizes the failure of the NC-DCS platform software. The logs and dump files corresponding to the historical defect data of the NC-DCS platform software are obtained. The software failure modes that characterize the failure of the NC-DCS platform software are obtained from the logs and dump files.
[0067] The parameter acquisition unit 20 is used to acquire all software parameters involved in the operation of the NC-DCS platform software.
[0068] The determination unit 30 is used to determine the software health parameters that are associated with the software failure mode from all software parameters; wherein, the software health parameters that are associated with the software failure mode are used as the software health parameters of the NC-DCS platform software, so as to determine the operating status of the NC-DCS platform software by analyzing the parameter values of the software health parameters.
[0069] In some examples, the determination unit 30 is specifically used to analyze software failure modes to establish a correspondence between software failure modes and software parameters. The software parameters in this correspondence are software health parameters associated with the software failure modes. These associated software health parameters include: indicator parameters, configuration parameters, logs, and dump files. Indicator parameters are used to report the operating status of device nodes running the NC-DCS platform software; configuration parameters are used to locate failures in the configuration representation of the NC-DCS platform software; logs are used to record the operating information of the NC-DCS platform software; and dump files are used to reconstruct the state of the NC-DCS platform software at the time of failure.
[0070] The software operation and maintenance monitoring module may also include: a data acquisition and storage unit 40, such as... Figure 7 As shown. The data acquisition and storage unit 40 is used to acquire and store environmental configuration parameters related to software failure modes before the NC-DCS platform software starts; to acquire and store indicator parameters and service configuration parameters related to software failure modes during the operation of the NC-DCS platform software; to retrieve logs related to software failure modes online from the NC-DCS platform software's log library and dump files related to software failure modes online from the NC-DCS platform software's dump file library during the operation of the NC-DCS platform software; and to store the logs and dump files.
[0071] The data acquisition and storage unit 40 collects indicator parameters and service configuration parameters according to a preset collection cycle during the operation of the NC-DCS platform software, and saves the indicator parameters once each time they are collected. It determines whether the indicator parameters meet the preset cycle setting conditions. If the indicator parameters meet the preset cycle setting conditions, the preset collection cycle remains unchanged. If the indicator parameters do not meet the preset cycle setting conditions, the preset collection cycle is shortened. During the operation of the NC-DCS platform software, if the collected service configuration parameters do not meet the preset storage conditions, the service configuration parameters are saved once.
[0072] In this embodiment, the software operation and maintenance monitoring module can be deployed on the device nodes of the NC-DCS platform software. The software operation and maintenance monitoring module is designed to be non-intrusive and deployed on the device nodes as a software plug-in. The software operation and maintenance monitoring module is used to collect and save software health parameters. The software operation and maintenance monitoring module is connected to HSNET, and HSNET only transmits software health parameters.
[0073] Furthermore, embodiments of this application also provide a computer-readable storage medium storing computer program code, which implements the above-described software monitoring method when executed.
[0074] It should be noted that the various embodiments in this specification can be described in a progressive manner, and the features described in the various embodiments can be substituted for or combined with each other. Each embodiment focuses on describing the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to mutually. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0075] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0076] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0077] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A software monitoring method, characterized in that, The method includes: A failure mode list of the non-safety-grade digital instrumentation and control system platform software is obtained. The failure mode list records software failure modes that characterize the failure of the non-safety-grade digital instrumentation and control system platform software. The software failure modes are obtained by analyzing at least one of the following: software failure modes of each level unit, field problems that occur when running the non-safety-grade digital instrumentation and control system platform software, and historical defect data of the non-safety-grade digital instrumentation and control system platform software. Each level unit includes the device nodes, tasks, threads, and functions involved in the operation of the non-safety-grade digital instrumentation and control system platform software. Obtain all software parameters involved in the runtime of the non-safety-grade digital instrumentation and control system platform software; From all the software parameters, determine the software health parameters that are associated with the software failure mode; wherein, the software health parameters associated with the software failure mode are used as the software health parameters of the non-safety-grade digital instrumentation and control system platform software, so as to determine the operating status of the non-safety-grade digital instrumentation and control system platform software by analyzing the parameter values of the software health parameters; The software failure modes are obtained by analyzing at least one of the following data: software failure modes of each level unit, field problems occurring when running the non-safety-grade digital instrumentation and control system platform software, and historical defect data of the non-safety-grade digital instrumentation and control system platform software. This includes: inferring software failure modes of each level unit from a software design perspective to identify software failure modes characterizing the failure of the non-safety-grade digital instrumentation and control system platform software; using the field problems as the top event in the software fault tree, decomposing the top event into multiple intermediate events and each intermediate event into multiple basic events, forming an event chain from the top event to the basic events leading to the top event, and summarizing the event chain as the software failure mode characterizing the failure of the non-safety-grade digital instrumentation and control system platform software; obtaining logs and dump files corresponding to the historical defect data of the non-safety-grade digital instrumentation and control system platform software, and obtaining the software failure modes characterizing the failure of the non-safety-grade digital instrumentation and control system platform software from the logs and dump files.
2. The method according to claim 1, characterized in that, The step of determining the software health parameters associated with the software failure mode from all the software parameters includes: analyzing the software failure mode to establish a correspondence between the software failure mode and the software parameters, wherein the software parameters in the correspondence are software health parameters associated with the software failure mode.
3. The method according to claim 2, characterized in that, The software health parameters associated with the software failure mode include: indicator parameters, configuration parameters, logs, and dump files; the indicator parameters are used to provide feedback on the operating status of the device nodes running the non-safety-grade digital instrumentation and control system platform software, the configuration parameters are used to locate failures caused by configuration errors in the non-safety-grade digital instrumentation and control system platform software, the logs are used to record the operating information of the non-safety-grade digital instrumentation and control system platform software, and the dump files are used to restore the state of the non-safety-grade digital instrumentation and control system platform software at the time of failure.
4. The method according to claim 1, characterized in that, The method further includes: Before the non-safety-grade digital instrumentation and control system platform software is started, environmental configuration parameters that are associated with the software failure mode are collected and saved. During the operation of the non-safety-grade digital instrumentation and control system platform software, indicator parameters and service configuration parameters that are associated with the software failure modes are collected and saved. When the non-safety-grade digital instrumentation and control system platform software is running, logs associated with the software failure mode are obtained online from the log library of the non-safety-grade digital instrumentation and control system platform software, and dump files associated with the software failure mode are obtained online from the dump file library of the non-safety-grade digital instrumentation and control system platform software. Save the log and the dump file; Software health parameters associated with the software failure mode include metric parameters, configuration parameters, logs, and dump files. The configuration parameters include the environment configuration parameters and the service configuration parameters.
5. The method according to claim 4, characterized in that, During the operation of the non-safety-grade digital instrumentation and control system platform software, the system collects and saves indicator parameters and service configuration parameters that are associated with the software failure modes, including: When the non-safety-grade digital instrumentation and control system platform software is running, the indicator parameters are collected according to a preset collection cycle, and the indicator parameters are saved once each time they are collected. Determine whether the indicator parameter meets the preset period setting conditions. If the indicator parameter meets the preset period setting conditions, maintain the preset collection period unchanged. If the indicator parameter does not meet the preset period setting conditions, shorten the preset collection period. When the non-safety-grade digital instrumentation and control system platform software is running, if the collected service configuration parameters do not meet the preset saving conditions, the service configuration parameters are saved once.
6. The method according to claim 4, characterized in that, A software operation and maintenance monitoring module is deployed in the device nodes of the non-safety-grade digital instrumentation and control system platform software. The software operation and maintenance monitoring module is designed to be non-intrusive and is deployed as a software plug-in on the device nodes. The software operation and maintenance monitoring module is used to collect and save the software health parameters. The software operation and maintenance monitoring module is connected to the health system network, which only transmits the software health parameters.
7. A software operation and maintenance monitoring module, characterized in that, The module includes: The list acquisition unit is used to obtain a failure mode list of the non-safety-grade digital instrumentation and control system platform software. The failure mode list records software failure modes that characterize the failure of the non-safety-grade digital instrumentation and control system platform software. The software failure modes are obtained by analyzing at least one of the following: software failure modes of each level unit, field problems that occur when running the non-safety-grade digital instrumentation and control system platform software, and historical defect data of the non-safety-grade digital instrumentation and control system platform software. Each level unit includes the device nodes, tasks, threads, and functions involved in the operation of the non-safety-grade digital instrumentation and control system platform software. The parameter acquisition unit is used to acquire all software parameters involved in the runtime of the non-safety-grade digital instrumentation and control system platform software. A determining unit is configured to determine software health parameters associated with the software failure mode from all software parameters; wherein the software health parameters associated with the software failure mode are used as software health parameters of the non-safety-grade digital instrumentation and control system platform software, so as to determine the operating status of the non-safety-grade digital instrumentation and control system platform software by analyzing the parameter values of the software health parameters. The list acquisition unit is specifically used to infer the software failure modes of each level of unit from a software design perspective, so as to identify the software failure modes characterizing the failure of the non-safety-grade digital instrumentation and control system platform software. The software design perspective includes software operating environment design, software general function design, and software special function design. The on-site problem is taken as the top event of the software fault tree. The top event is decomposed into multiple intermediate events and each intermediate event is decomposed into multiple basic events. An event chain leading to the occurrence of the top event is formed from the top event to the basic events. The event chain is summarized as the software failure mode characterizing the failure of the non-safety-grade digital instrumentation and control system platform software. Logs and dump files corresponding to the historical defect data of the non-safety-grade digital instrumentation and control system platform software are obtained from the logs and dump files.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program code, which, when executed, implements the software monitoring method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and system for refinery enterprise equipment management informatization
CN113537681A
Software health parameter identification method and device, equipment and medium
CN115437814A