Maintenance method and device for configuration management database

Through the NLP model and graph neural network automated processing of configuration management database, the problem of data errors caused by manual entry and difficult to capture multi-layer dependencies is solved, and efficient configuration item management and fault prediction are achieved.

CN120371813AActive Publication Date: 2025-07-25CHANGSHA YULIAN INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510421899.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-25
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the existing configuration management database, configuration item management relies on manual and static scripts, which can easily lead to data errors and make it difficult to capture multi-layer dependencies in complex environments.

Method used

The NLP model is used to extract CI information, generate standardized configuration items, and establish a CI relationship model in combination with the graph neural network. The failure risk is predicted through time series analysis, the interactive module is used for human-computer interaction, and the integrated asset automatic discovery module is used for dynamic updates.

Benefits of technology

It realizes the replacement of manual entry of automated processes, improves the convenience and accuracy of system updates, reduces the failure rate, and accurately captures dependencies and predicts the scope of the failure impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371813A_ABST
    Figure CN120371813A_ABST
Patent Text Reader

Abstract

The invention relates to the field of database management, in particular to a maintenance method and device for a configuration management database, and discloses the maintenance device for the configuration management database, which comprises an NLP model, a time sequence analysis model, a graph neural network, a database and an interaction module, the time sequence analysis model is used for predicting a resource use trend or a fault risk, the graph neural network is used for establishing a CI relation model and analyzing a dependency relation and an influence range, the database is used for storing user data, and the interaction module is used for supporting human-computer interaction between a user and the device. According to the method, the text information is extracted through the NLP model, the CI information in the data is obtained, the dependency relationship in the information is captured for the graph neural network in combination with semantic understanding, and the limitation that a traditional tool can only process structured data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of database management, and particularly to a method and device for maintaining a configuration management database. Background Art

[0002] The Configuration Management Database, abbreviated as CMDB, is a core system for storing and managing all configuration items and their relationships in an IT system. It provides a centralized database for describing the structure, status, and dependencies of an IT infrastructure, usually including configuration items, attributes of configuration items, and dependencies of configuration items.

[0003] In the maintenance of existing configuration management databases, the management of configuration items mainly relies on manual, static script collection, or automated discovery tools based on certain rules. However, with the change of enterprise architecture, manual participation is prone to data errors due to operation oversights or inconsistent standards, and in a large database, it is difficult to capture multi-layer dependencies in a complex environment by describing CI associations through predefined relationship types. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides the following technical solutions:

[0005] A maintenance device for a configuration management database, comprising: an NLP model, a time series analysis model, a graph neural network, a database, and an interaction module.

[0006] Specifically, the NLP model is used to extract CI information from data and generate standardized configuration items. The time series analysis model is used to predict resource usage trends or failure risks. The graph neural network establishes a CI relationship model to analyze dependency relationships and influence scopes. The database is used to store user data, and the interaction module is used to support human-computer interaction between the user and the device.

[0007] As an improvement of the above technical solution, the steps for the NLP model to extract CI information from data and generate standardized configuration items are as follows:

[0008] Perform text splitting on the classified and marked data by calling the spaCy library;

[0009] Perform fine-grained entity extraction on the split problem through the BERT model, locate the CI information fields in the text, and generate standardized configuration items based on the CI information.

[0010] As an improvement of the above technical solution, the database is integrated with an asset automatic discovery module. The asset automatic discovery module discovers asset information by obtaining the traffic of data packets and performs dynamic updates by establishing the CI information of assets.

[0011] A maintenance method for a configuration management database, applied to the maintenance device of the configuration management database described in any one of the foregoing technical solutions, includes the following steps:

[0012] S10: Data preprocessing.

[0013] S20: Extract the processed data through an NLP model to obtain CI information, and generate standardized configuration items based on the CI information.

[0014] S30: Based on the standardized configuration items, establish a CI relationship model through a graph neural network to analyze the dependency relationship and influence scope.

[0015] S40: Analyze the resource metric data through a time series analysis model, and combine it with the CI relationship model to predict the resource usage trend or failure risk.

[0016] As an improvement of the above technical solution, the step S10 includes the following steps:

[0017] S11: Pre-collection of data.

[0018] S12: Clean the redundant, missing or abnormal data in the collected data.

[0019] S13: Classify and label the processed data to provide supervised data for model training.

[0020] As an improvement of the above technical solution, the pre-collection of the data includes the following steps:

[0021] S111: Collect time series data and resource metric data through the Prometheus monitoring system.

[0022] S112: Collect system alarm record information through the Nagios model.

[0023] S113: Collect system log information through the ELK tool.

[0024] As an improvement of the above technical solution, the step S20 includes the following steps:

[0025] S21: Input the processed data into the NLP model, split the classified and labeled data through the spaCy library, and perform fine-grained entity extraction on the split problems through the BERT model to locate the CI information fields in the text.

[0026] S22: Through configuration item standardization processing, organize and match the corresponding CI information fields to generate standardized configuration items.

[0027] As an improvement of the above technical solution, step S30 includes the following steps:

[0028] S31: Take each configuration item in the standardized configuration items as the node data of the graph, and define the relationship type between CIs as the edge data.

[0029] S32: Establish a structure diagram according to the node data, edge data, and implicit relationships between the data.

[0030] S33: Establish a GraphSAGE graph neural network and learn the importance of neighbors through the attention mechanism.

[0031] S34: Calculate the node similarity through GNN embedding, discover potential dependencies, and analyze directly connected nodes and indirectly dependent nodes.

[0032] S35: Identify the CI list affecting the scope and the topological path in the structure diagram according to the directly connected nodes and indirectly dependent nodes, so as to obtain the dependency relationship and influence scope of the faulty node.

[0033] Advantages of the present invention:

[0034] Based on the NLP model, accurately extract CI attributes (such as version numbers, dependency descriptions) from unstructured texts such as operation and maintenance documents and logs. Extract text information through the NLP model to obtain CI information in the data, and combine semantic understanding to capture the dependency relationships in the information by the graph neural network, solving the limitation that traditional tools can only process structured data. The automated process replaces the non-automated process update methods such as manual entry, making the update of the entire system more convenient and the failure rate of system update lower. Description of the Drawings

[0035] Figure 1 It is the principle block diagram of a maintenance device for a configuration management database of the present invention;

[0036] Figure 2 It is the flow block diagram of a maintenance method for a configuration management database of the present invention. Detailed Embodiments

[0037] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0038] In the maintenance of the existing configuration management database, the management of configuration items mainly relies on manual, static script collection or automated discovery tools based on certain rules. However, with the changes in enterprise architecture, manual participation is prone to data errors due to operational omissions or inconsistent standards. In addition, in a large database, it is difficult to capture multi-layer dependencies in a complex environment by describing CI associations through predefined relationship types.

[0039] In order to solve the above problems, the following embodiments are provided:

[0040] Embodiment 1

[0041] Provided is a maintenance device for a configuration management database, comprising: an NLP model, a time series analysis model, a graph neural network, a database, and an interaction module.

[0042] Specifically, the NLP model is used to extract CI information from the data and generate standardized configuration items. The time series analysis model is used to predict resource usage trends or failure risks. The graph neural network establishes a CI relationship model and analyzes dependencies and impact ranges. The database is used to store user data. The interaction module is used to support human-computer interaction between users and devices.

[0043] The NLP model is used to analyze and process the pre-prepared data, extract the CI information, generate standardized configuration items based on the CI information, and establish a CI relationship model through the graph neural network. The CI relationship model includes: physical connection relationships between devices (such as network cable connections between servers and switches), logical dependencies (such as application dependencies on database services), and deployment relationships (such as virtual machines deployed on physical hosts). Based on these relationships, the dependencies and impact range of faulty nodes can be analyzed.

[0044] When a fault occurs, the time series analysis model can analyze the predicted resource indicators (such as disk usage). If the predicted indicator of a device exceeds the threshold, the linkage graph neural network quickly generates a list to achieve rapid alarm prompts. The interactive module is further implemented through the NLP model. The NLP model performs semantic understanding of the text information entered by the user, analyzes the user's needs, and then generates and executes corresponding action instructions based on their needs. The database stores the extracted CI information, which usually includes configuration items, their attributes, and their dependencies.

[0045] In one embodiment, the NLP model extracts CI information from the data and generates standardized configuration items, including the following steps:

[0046] The spaCy library is used to perform text segmentation on the classified and labeled data.

[0047] The spaCy library is a high-performance natural language processing library that can split long texts into sentence-level fragment texts, which can be used to split long text information and reduce the length requirements for semantic recognition.

[0048] Through the BERT model, fine-grained entity extraction is performed on the split questions to locate the CI information fields in the text, and standardized configuration items are generated based on the CI information.

[0049] As the "semantic understanding layer", the BERT model solves the problems of context dependence and fuzzy matching. It can cooperate with the spaCy library to accurately identify the CI information fields.

[0050] In one embodiment, the database is integrated with an asset automatic discovery module. The asset automatic discovery module discovers asset information by obtaining the traffic of data packets and performs dynamic updates by establishing the CI information of the assets.

[0051] Specifically, when new data information appears in the database, the type and IP of the assets are discovered by obtaining the traffic of data packets, and the CI information is established by sorting out this information. Then, the CI information is updated to the system through an update tool. Among them, the ServiceNowDiscovery module can be used to implement the selection of the update tool. This module is one of the core modules of the ServiceNow platform and can automatically discover the configuration items (Configuration Items, CI) in the enterprise IT environment and their association relationships, and synchronize this data to the Configuration Management Database (CMDB) in real time.

[0052] Embodiment 2

[0053] To cooperate with Embodiment 1, a method for maintaining a configuration management database is provided, which is applied to the maintenance device of the configuration management database described in any one of Embodiment 1, and includes the following steps:

[0054] S10: Data preprocessing.

[0055] Specifically, the step S10 includes the following steps:

[0056] S11: Pre-collection of data.

[0057] Among them, the pre-collection of data includes the following steps:

[0058] S111: Collect time series data and resource metric data through the Prometheus monitoring system.

[0059] The resource metric data collected by the Prometheus monitoring system usually includes data such as CPU utilization rate and network latency.

[0060] S112: Collect system alarm record information through the Nagios model.

[0061] Alarm record information is the alarm information temporarily generated when there are faults in the system. These alarm information are usually used for storage and analysis. When the same fault problem occurs next time, alarm feedback can be provided more quickly.

[0062] S113: Collect system log information through the ELK tool.

[0063] In the ELK tool, Elasticsearch is used as a dedicated log database and may be transferred to Hadoop HDFS during long-term archiving. This tool usually collects the log information of the system's daily work, and these information can be used to help the neural network for training.

[0064] S12: Clean the redundant, missing or abnormal data in the collected data.

[0065] In the previous data collection process, it is inevitable to collect a large amount of useless data, missing data, abnormal data, etc. If all these data are used for training and analysis, it will not only waste a lot of time, but also affect the training results. Therefore, cleaning is still needed. Usually, tools such as Pandas (for small-scale data), Dask (for distributed processing), and PySpark (for large amounts of data) in the Python ecosystem can be used for processing.

[0066] S13: Classify and label the processed data to provide supervised data for model training.

[0067] After the data cleaning is completed, the data needs to be classified and labeled. Here, methods such as knowledge graphs and semi-automatic annotation can usually be used to assist. These are all related solutions disclosed in the prior art, so they will not be elaborated too much in this embodiment.

[0068] S20: Extract the processed data through the NLP model to obtain CI information, and generate standardized configuration items according to the CI information.

[0069] Specifically, the step S20 includes the following steps:

[0070] S21: Input the processed data into the NLP model, split the classified and labeled data through calling the spaCy library, and perform fine-grained entity extraction on the split questions through the BERT model to locate the CI information fields in the text.

[0071] The spaCy library can quickly extract fixed entities and perform associated matching on text. Therefore, long texts can be quickly decomposed into short texts. The BERT model, on the other hand, can identify context-related complex entities and deeply understand semantic structures, enabling it to quickly extract CI information fields from related short texts. In this step, the extracted information also includes the corresponding implicit relationships.

[0072] S22: Through the standardization process of configuration items, the corresponding CI information fields are sorted and matched to generate standardized configuration items.

[0073] The standardization process of configuration items usually includes the unification of naming formats, pattern verification, and error correction. The purpose of unifying the naming format is to preset a rule engine to ensure the regularity of data and reduce the complexity of subsequent retrieval. Pattern verification is usually carried out through methods such as dictionary constraints and logical verification. The purpose of matching and error correction is to correct the spelling of misinput words, etc. For example, changing WindwosServer to WindowsServer.

[0074] Finally, the CI information fields processed by the standardization of configuration items are sorted and stored as standardized configuration items in the database.

[0075] S30: Based on the standardized configuration items, a CI relationship model is established through a graph neural network to analyze the dependency relationships and influence scopes.

[0076] Specifically, step S30 includes the following steps:

[0077] S31: Each configuration item in the standardized configuration item is used as the node data of the graph, and the relationship type between CIs is defined as the edge data.

[0078] S32: A structure diagram is established based on the node data, edge data, and implicit relationships between the data.

[0079] The structure diagram constructed through steps S31 and S32 contains configuration items and the relationships between configuration items. These data correspond to CI information, thus converting the CI information into a topological structure diagram linked by relationships.

[0080] S33: Establish a GraphSAGE graph neural network and learn the importance of neighbors through an attention mechanism.

[0081] Analyze and learn the structure diagram through the GraphSAGE graph neural network, understand the relationships between nodes, and be able to establish the recognition of dependency relationships.

[0082] S34: Calculate the node similarity through GNN embedding, discover potential dependencies, and analyze directly connected nodes and indirectly dependent nodes.

[0083] By calculating node similarity through GNN embeddings, the implicit dependencies between some nodes can be obtained. In this way, not only can directly connected nodes be obtained, but also nodes with indirect dependencies can be obtained.

[0084] S35: Identify the CI list of the impact scope and the topological paths in the structure diagram based on directly connected nodes and indirectly dependent nodes to obtain the dependency relationship and impact scope of the faulty node.

[0085] When a fault occurs, while locating the faulty node, directly connected nodes and indirectly dependent nodes can be immediately analyzed. Generally, it is considered that the synchronous impact rate of directly connected nodes is relatively high, while the synchronous impact rate of indirectly dependent nodes needs to be calculated and predicted based on GNN. After all impact calculations are completed, the CI information and list corresponding to the nodes and edges included in these affected scopes are output, and the topological paths in the structure diagram are sorted out for display.

[0086] S40: Analyze the resource metric data through a time series analysis model, and combine it with the CI relationship model to predict the resource usage trend or failure risk.

[0087] The time series model (such as LSTM or Prophet) analyzes historical and real-time resource data (such as CPU utilization, disk I / O) to capture periodic patterns and abnormal fluctuations; at the same time, the CI relationship model based on the graph neural network (GNN) correlates and models resource metrics with topological dependencies (such as the deployment relationship between servers and applications). Thus, when predicting the resource trend of a single device, the cascading impact on upstream and downstream systems (such as database service interruption that may be triggered by insufficient storage space) can be evaluated synchronously. This dual-model collaboration mechanism can not only accurately predict resource bottlenecks but also dynamically generate fault impact paths, providing a decision-making basis for capacity planning and fault self-healing.

[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A maintenance device for a configuration management database, characterized in that, It includes: An NLP model for extracting CI information from data and generating standardized configuration items; A time series analysis model for predicting resource usage trends or failure risks; A graph neural network for establishing a CI relationship model and analyzing dependency relationships and influence scopes; A database for storing user data; An interaction module for supporting human-computer interaction between the user and the device.

2. The maintenance device for a configuration management database according to claim 1, characterized in that: The steps for the NLP model to extract CI information from data and generate standardized configuration items are as follows: Perform text splitting on the classified and labeled data by calling the spaCy library; Perform fine-grained entity extraction on the split questions through the BERT model, locate the CI information fields in the text, and generate standardized configuration items based on the CI information.

3. The maintenance method of the configuration management database according to claim 1, wherein: The database is integrated with an asset automatic discovery module. The asset automatic discovery module discovers asset information by obtaining the traffic of data packets and performs dynamic updates by establishing the CI information of the assets.

4. A method for maintaining a configuration management database, applied to the maintenance device of the configuration management database according to any one of claims 1 to 3, characterized in that It includes the following steps: S10: Data preprocessing; S20: Extract the processed data through the NLP model, obtain CI information, and generate standardized configuration items based on the CI information; S30: Based on the standardized configuration items, establish a CI relationship model through a graph neural network and analyze dependency relationships and influence scopes; S40: Analyze the resource metric data through the time series analysis model, and combine it with the CI relationship model to predict resource usage trends or failure risks.

5. The maintenance method of the configuration management database according to claim 4, wherein: The step S10 includes the following steps: S11: Pre-collection of data; S12: Clean the redundant, missing, or abnormal data in the collected data; S13: Classify and label the processed data to provide supervised data for model training.

6. The maintenance method of the configuration management database according to claim 5, characterized in that: The pre-collection of the data includes the following steps: S111: Collect time series data and resource metric data (such as CPU utilization, network latency) through the Prometheus monitoring system; S112: Collect system alarm record information through the Nagios model; S113: Collect system log information through the ELK tool.

7. The maintenance method of the configuration management database according to claim 4, characterized in that: The step S20 includes the following steps: S21: Input the processed data into the NLP model, perform text splitting on the classified and labeled data by calling the spaCy library, and perform fine-grained entity extraction on the split questions through the BERT model to locate the CI information fields in the text; S22: Through configuration item standardization processing, organize and match the corresponding CI information fields to generate standardized configuration items.

8. The maintenance method of the configuration management database according to claim 4, characterized in that: The step S30 includes the following steps: S31: Use each configuration item in the standardized configuration item as the node data of the graph, and define the relationship type between CIs as the edge data; S32: Establish a structure diagram based on the node data, edge data, and implicit relationships between the data; S33: Establish a GraphSAGE graph neural network and learn the importance of neighbors through the attention mechanism; S34: Calculate node similarity through GNN embedding, discover potential dependencies, and analyze directly connected nodes and indirectly dependent nodes; S35: Identify the CI list of the influence scope and the topological path in the structure diagram based on the directly connected nodes and indirectly dependent nodes, so as to obtain the dependency relationship and influence scope of the faulty node.

Citation Information

Patent Citations

  • CI model definition method based on CMDB

    CN110717726A

  • Database change auditing management method and related equipment

    CN114218200A

  • Data query method and device of CMDB and electronic equipment

    CN117633002A

  • System and method for building knowledgebase of dependencies between configuration items

    US20230259368A1