Multi-strategy driven configuration item automatic association method

Through a multi-strategy driven automatic configuration item association method, using multi-level association strategies and relationship edge evaluation, the problems of tedious and error-prone manual operations in operation and maintenance data association are solved, and more efficient and accurate data association and system management are achieved.

CN117272135BActive Publication Date: 2025-09-23SHANGHAI ZHONGYI TURING DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311150969.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-06
Publication Date
2025-09-23
Estimated Expiration
2043-09-06

AI Technical Summary

Technical Problem

In existing technologies, operation and maintenance data correlation analysis relies on manually defined rules, which leads to heavy workload, prone to errors, difficulty in discovering unknown correlations, lack of intelligence, and difficulty in handling complex systems and network architectures.

Method used

A multi-strategy driven automatic association method for configuration items is adopted. The configuration item instance attributes are initially screened according to scenario requirements. A topology diagram is constructed using multi-level association strategies. The reliability of association relationships is evaluated through relationship edge evaluation strategies, thereby improving the maintainability and scalability of the data system.

Benefits of technology

It reduces manual operations, improves the accuracy and efficiency of operation and maintenance data association, can better handle complex systems and network architectures, and provides a global view and flexible association management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117272135B_ABST
    Figure CN117272135B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-strategy driven automatic configuration item association method, comprising: based on scenario requirements, initially screening the attributes of the marked configuration item instances to obtain the configuration item instances to be analyzed and their attributes to be analyzed; based on a multi-level association strategy, performing association analysis on the attributes to be analyzed of a first configuration item instance and a second configuration item instance belonging to different categories of configuration items, determining and marking the relationship edges between the two categories of configuration item instances; constructing a configuration item topology graph based on the relationship edges, the first configuration item instance and the second configuration item instance; determining an evaluation score of the configuration item topology graph based on a relationship edge evaluation strategy and a node evaluation strategy; and determining whether it is necessary to redetermine the relationship edge between the first and second configuration item instances based on whether the evaluation score of the configuration item topology graph meets a preset trust threshold. The method provided by the present invention can realize the automatic association of configuration items of different categories and improve the maintainability and scalability of the data system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of operation and maintenance data technology, and in particular to a multi-strategy driven configuration item automatic association method. Background Art

[0002] Massive amounts of operations and maintenance data contain a wealth of knowledge and information. However, extracting meaningful information from this data and establishing relationships between data is a challenging task. Traditional association analysis methods rely primarily on manually defined association rules. However, with the expansion of data volumes and rapidly increasing user demands, manual processing has become cumbersome and error-prone. Consequently, automated tools have emerged, such as network topology discovery tools, which can automatically discover relationships between network devices and draw network topology maps. These tools can reduce manual errors and workload, improving efficiency. However, they still require manual rule definition, lack intelligence, and can make it difficult to discover unknown relationships. Therefore, intelligent association has emerged as a new solution. However, due to the complex relationships between operations and maintenance data, such as the complex dependencies and relationships between multiple components such as servers, network devices, and applications, existing intelligent association technologies may struggle to accurately capture and understand these complex relationships, resulting in incomplete or erroneous association results.

[0003] Traditional operations and automated data association tools typically require manual data entry or updates, pre-defined fixed association rules, or manual feature selection to establish associations. This creates a heavy workload for subsequent maintenance and lacks intelligence, making it difficult to discover unknown associations. With the massive growth of systems and network architectures, manual operations have become increasingly cumbersome and error-prone.

[0004] Therefore, the above-mentioned defects existing in the related art have become technical problems that need to be solved urgently in this field. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention provides a multi-strategy driven configuration item automatic association method.

[0006] In a first aspect, the present invention provides a multi-policy driven configuration item automatic association method, comprising:

[0007] Based on scenario requirements, the attributes of each marked configuration item instance are preliminarily screened to obtain the configuration item instance to be analyzed and the attributes to be analyzed of the configuration item instance to be analyzed; the configuration item instance to be analyzed includes multiple attributes and attribute values ​​of the multiple attributes;

[0008] Based on a multi-level association strategy, an association analysis is performed on the attributes to be analyzed of the first configuration item instance and the attributes to be analyzed of the second configuration item instance, and a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge are determined; the multi-level association strategy is used to define a rule for associating the attributes to be analyzed in a manner of prioritizing attribute values ​​and then attribute names; the first configuration item instance and the second configuration item instance are any one of the configuration item instances to be analyzed, and belong to different categories of configuration items.

[0009] Constructing a configuration item topology graph based on the relationship edge between the first configuration item instance and the second configuration item instance, the first configuration item instance, and the second configuration item instance;

[0010] Determining an evaluation score of the configuration item topology graph based on a relationship edge evaluation strategy and a node evaluation strategy;

[0011] Based on whether the evaluation score of the configuration item topology graph meets a preset trust threshold, it is determined whether it is necessary to redetermine the relationship edge between the first configuration item instance and the second configuration item instance.

[0012] Optionally, based on scenario requirements, the attributes of each marked configuration item instance are preliminarily screened to obtain the attributes to be analyzed of the configuration item instance to be analyzed, including:

[0013] Based on the prior library, the attributes of each configuration item instance are marked according to the attribute name type and attribute value type;

[0014] Based on the scenario requirements, the attribute name type and / or attribute value type that needs to be associated is determined, and the attributes of each configuration item instance are preliminarily screened to obtain the attributes to be analyzed of the configuration item instance to be analyzed.

[0015] Optionally, based on scenario requirements, after initially screening the attributes of each marked configuration item instance to obtain the attributes to be analyzed of the configuration item instance to be analyzed, the following steps may be performed:

[0016] Compare the attribute to be analyzed of the first configuration item instance with the attribute to be analyzed of the second target configuration item, and select all attribute pairs that satisfy equal attribute values ​​as a first candidate set; the attribute pair consists of the first configuration item instance, the first attribute, the second configuration item instance, and the second attribute, the first attribute is any attribute to be analyzed of the first configuration item instance, and the second attribute is one of the attributes to be analyzed of the second configuration item instance that satisfies the attribute value equal to that of the first attribute.

[0017] Optionally, performing association analysis on the to-be-analyzed attribute of the first configuration item instance and the to-be-analyzed attribute of the second configuration item instance based on the multi-level association strategy, and determining a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge includes:

[0018] In the first candidate set, based on the attribute name alignment strategy included in the multi-level association strategy, attribute pairs with the same attribute name are selected as the second candidate set;

[0019] Marking a first relationship edge between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the second candidate set;

[0020] In the third candidate set, based on the attribute name matching strategy included in the multi-level association strategy, attribute pairs that satisfy the attribute name mapping to the namespace of the same attribute category are screened as the fourth candidate set; the third candidate set is the intersection of the complement of the second candidate set and the first candidate set;

[0021] A second relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the fourth candidate set.

[0022] Optionally, the performing association analysis on the to-be-analyzed attribute of the first configuration item instance and the to-be-analyzed attribute of the second configuration item instance based on the multi-level association strategy to determine a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge further includes:

[0023] Determine a complement of the union of the second candidate set and the fourth candidate set, and an intersection of the complement and the first candidate set as a fifth candidate set;

[0024] In the fifth candidate set, based on the multi-level association strategy including a sequential pattern matching strategy, the attribute pairs whose sequence similarity satisfies a first threshold are screened as a sixth candidate set; the sequence similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair;

[0025] Alternatively, in the fifth candidate set, based on the multi-level association strategy including a semantic matching strategy, the attribute pairs whose semantic similarity meets the second threshold are screened as the seventh candidate set; the semantic similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair, which are respectively mapped to the first word vector and the second word vector in the pre-trained word vector space;

[0026] Acquire the attribute pair included in the sixth candidate set or the attribute pair included in the seventh candidate set as the target attribute pair;

[0027] A third relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the target attribute pair.

[0028] Optionally, determining the evaluation score of the relationship edge in the configuration item topology graph based on the relationship edge evaluation strategy and the node evaluation strategy includes:

[0029] Determine, based on the priority order corresponding to the multi-level association strategy for determining the relationship edge between the first configuration item instance and the second configuration item instance, a first evaluation score, a second evaluation score, and a third evaluation score; the first evaluation score is the evaluation score of the first relationship edge, the second evaluation score is the evaluation score of the second relationship edge, and the third evaluation score is the evaluation score of the third relationship edge, the first evaluation score being the highest, the second evaluation score being the second highest, and the third evaluation score being the lowest;

[0030] determining, based on the number of relationship edges between the first configuration item instance and the second configuration item instance, a fourth evaluation score of the relationship edges in the configuration item topology graph;

[0031] An evaluation score of a relationship edge in the configuration item topology graph is determined based on a weighted sum algorithm and the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score.

[0032] Optionally, determining whether it is necessary to redetermine the relationship edge between the first configuration item instance and the second configuration item instance based on whether the evaluation score of the configuration item topology graph meets a preset trust threshold includes:

[0033] If the evaluation score of the configuration item topology graph meets a preset trust threshold, the relationship edge is marked as reliable;

[0034] If the evaluation score of the configuration item topology graph does not meet the preset trust threshold, the relationship edge is marked as unreliable and output for manual analysis.

[0035] Optionally, in the third candidate set, based on the attribute name matching strategy included in the multi-level association strategy, attribute pairs that satisfy the mapping of attribute names to the namespace of the same attribute category are screened as the fourth candidate set, including:

[0036] Construct a unified namespace based on all synonyms, derivatives, strongly related words, abbreviations and affixes included in each attribute category;

[0037] In the third candidate set, the first configuration item instance and the second configuration item instance corresponding to the attribute pair are obtained as the first target configuration item instance and the second target configuration item instance respectively;

[0038] Mapping the attribute name of the first target configuration item instance to the unified namespace to obtain a first attribute name corresponding to the attribute name of the first target configuration item instance;

[0039] Mapping the attribute name of the second target configuration item instance to the unified namespace to obtain a second attribute name corresponding to the attribute name of the second target configuration item instance;

[0040] If the first attribute name and the second attribute name are the same, it is determined that the first target configuration item instance and the second target configuration item instance meet the attribute name matching strategy, and are added to the fourth candidate set.

[0041] Optionally, the sequence pattern matching strategy includes SequenceMatcher, Cosine similarity, Jaccard similarity, Levenshtein distance, and Jaro-Winkler distance.

[0042] Optionally, the pre-trained word vector space is constructed based on a SentenceTransformer model.

[0043] The multi-strategy driven configuration item automatic association method provided by the present invention preliminarily screens the original configuration item instances according to scenario requirements, and defines multi-level association strategies. The configuration item instances that pass the preliminary screening are screened according to the association strategies of each level, and the relationship edge between any two configuration item instances is determined. The relationship edge is used to represent the association relationship between the attributes of the two configuration item instances. Then, based on the association relationship between the configuration item instance and the attributes of any two configuration item instances, a configuration item topology graph is constructed to further determine whether the above-mentioned determined association relationship is reliable, thereby improving the maintainability and scalability of the data system. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 This is one of the flow diagrams of the multi-policy driven configuration item automatic association method provided by an embodiment of the present invention;

[0046] Figure 2 This is the second flow chart of the multi-policy driven configuration item automatic association method provided by an embodiment of the present invention;

[0047] Figure 3 is a schematic diagram of an attribute name alignment strategy provided by an embodiment of the present invention;

[0048] Figure 4 is a schematic diagram of an attribute name matching strategy provided by an embodiment of the present invention;

[0049] Figure 5 It is a schematic diagram of a special case analysis strategy provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0051] Existing operation and maintenance and data association methods have the following shortcomings:

[0052] 1. Large manual workload: Linked data requires manual design, entry, and updating, which takes a lot of time and effort, especially when the data volume is large or the association relationships are complex.

[0053] 2. Error-prone: Since associated data involves relationships between multiple tables, manual errors may occur, such as incorrectly establishing associations, missing or incorrectly entering data, etc.

[0054] 3. Data accuracy needs to be confirmed: Errors, omissions, or outdated data in the Configuration Management Database (CMDB) can cause problems in establishing relationships. CMDB data often contains multiple data sources and data collection points, which can lead to data consistency issues. Differences in naming conventions and data formats between different data sources can also make establishing relationships difficult.

[0055] 4. Inflexibility: Traditional linked data methods are usually static. Once the association relationship is established, modifying or adjusting the association relationship may require a lot of work and risks.

[0056] 5. Poor scalability: When the scale of linked data expands or the data model changes, traditional linked data methods may be unable to meet the needs and require complex adjustments and reconstruction.

[0057] 6. Difficulty in handling complex relationships: Real-world data often has complex relationships, including multi-level relationships, time series relationships, and graph structure relationships. Traditional association analysis methods often have difficulty handling these complex relationships.

[0058] 7. Lack of a global view: CMDBs are often maintained by different teams and departments. Each team or department may focus only on the data and relationships relevant to them, lacking a global view. This can make it difficult to establish relationships between data across teams and departments, leading to a lack of a global view.

[0059] These drawbacks limit the efficiency and flexibility of traditional linked data approaches. Consequently, the O&M field is constantly exploring new approaches, moving towards intelligent association. While existing technologies have made some progress in this area, challenges and shortcomings remain. The present invention addresses these challenges in related technologies and proposes a technical solution.

[0060] The specific implementation process of the method provided by this invention can be abstracted into three stages: a priori filtering, association analysis, and a posteriori evaluation. The process of establishing associations is an analytical process ranging from simple to complex. Adopting a divide-and-conquer approach, specific strategies are applied to different situations, focusing on objectives, solving problems, establishing boundaries, and effectively avoiding conflicts between strategies. This framework provides strong maintainability and scalability for the use and management of strategies. Figure 1 This is one of the flow charts of the multi-strategy driven configuration item automatic association method provided by the embodiment of the present invention, such as Figure 1 As shown, including:

[0061] 1. Prior Identification

[0062] The a priori strategy provided by this invention performs a preliminary screening of the original configuration item instances. Specifically, it labels the attributes of the configuration item instances according to the a priori library. Based on the scenario requirements, it determines which attributes require analysis and which do not. This preliminary screening identifies the configuration item instances that require analysis and their attributes.

[0063] Then, the attribute values ​​of the attributes of any two configuration item instances belonging to the two categories of configuration items are compared, and the attributes of the two configuration item instances with equal attribute values ​​are determined as attribute pairs, that is, the initial candidate set is determined.

[0064] 2. Correlation Analysis

[0065] According to the association strategy, for the initial candidate set, the association relationships of the attribute pairs in the initial candidate set are determined in a nested manner according to the multiple-level association strategies associated with the attribute names, and the corresponding relationship edges between the two configuration item instances in the attribute pairs are marked;

[0066] For example, when the attribute values ​​and attribute names are equal, or when the attribute values ​​are equal but the attribute names are different, it is necessary to continue matching the attribute names based on other related words in the attribute names. If no match is found, the attribute name matching process can also be performed based on sequence matching or semantic matching to filter out all associated relationship edges as much as possible.

[0067] 3. Posterior Evaluation

[0068] Based on the evaluation strategy, the scores of relationship edges are evaluated for association strategies at different levels. Here, a relationship edge refers to the relationship edge corresponding to the labels of two matching configuration item instances. This relationship edge represents the association relationship between the attributes of the two configuration item instances. For example, the higher the evaluation strategy is in the hierarchy, that is, the earlier the matching strategy is performed, the simpler the relationship strategy is, and the higher the corresponding relationship edge score is. The lower the evaluation strategy is in the hierarchy, that is, the later the matching strategy is, the more complex the relationship strategy is, and the lower the corresponding relationship edge score is. Alternatively, the score can be evaluated based on the number of relationship edges between any two configuration item instances. These scores are combined and compared with a trust threshold determined by historical data. If the score falls below the threshold, these relationship edges are marked as unreliable and output for manual analysis.

[0069] Figure 2 This is a second flow chart of the multi-strategy driven configuration item automatic association method provided by an embodiment of the present invention; Figure 2 As shown, the method includes:

[0070] Step 201: Based on scenario requirements, preliminarily screen the attributes of each marked configuration item instance to obtain a configuration item instance to be analyzed and its attributes to be analyzed; the configuration item instance to be analyzed includes multiple attributes and attribute values ​​of the multiple attributes;

[0071] Step 202: Based on a multi-level association strategy, an association analysis is performed on the attributes to be analyzed of the first configuration item instance and the attributes to be analyzed of the second configuration item instance, and a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge are determined; the multi-level association strategy is used to define a rule for associating the attributes to be analyzed in a manner of prioritizing attribute values ​​and then attribute names; the first configuration item instance and the second configuration item instance are any one of the configuration item instances to be analyzed, and belong to different categories of configuration items.

[0072] Step 203: Construct a configuration item topology graph based on the relationship edge between the first configuration item instance and the second configuration item instance, the first configuration item instance, and the second configuration item instance;

[0073] Step 204: Determine an evaluation score of the configuration item topology graph based on the relationship edge evaluation strategy and the node evaluation strategy;

[0074] Step 205: Based on whether the evaluation score of the configuration item topology graph meets a preset credibility threshold, determine whether it is necessary to redetermine the relationship edge between the first configuration item instance and the second configuration item instance.

[0075] Specifically, we obtain various categories of configuration items requiring association analysis. Based on the scenario requirements, we determine which marked attributes in the configuration item instances require further analysis and which do not. Specifically, we determine which configuration item instances require analysis and which attributes in each configuration item instance are the attributes to be analyzed. Configuration item instances here include multiple attributes, the attribute value type corresponding to each attribute, and possibly the attribute value range for each attribute.

[0076] Different scenarios require different attributes within the corresponding configuration item instances that require further analysis. For example, in a scenario focused on the network environment, the corresponding network device name, IP address, and port number all require further analysis. For example, in a mobile banking transaction scenario, the attributes of the configuration item instances involved, such as the mobile banking system, virtual machine, minicomputer LPAR server (minipartition), middleware (such as WAS and Nginx), storage device, and network type, all require further analysis.

[0077] After determining the attributes to be analyzed of the configuration item instances to be analyzed, the multi-level association strategy provided by the present invention is used to perform association analysis on the attributes to be analyzed of the above configuration item instances to be analyzed. For example, association analysis is performed based on attribute values ​​and / or association analysis is performed based on attribute names to determine the association relationship between the attributes of any two configuration item instances, and the corresponding relationship edge is marked between the two configuration item instances.

[0078] After determining the relationship edge between any two configuration item instances, a configuration item topology graph can be constructed by combining the two configuration item instances.

[0079] The edge and node evaluation strategies are further used to estimate the scores of the edges in the configuration item topology graph and determine whether the evaluation scores meet the preset trust threshold. This determines whether the edge identified in the previous step is reliable. If it is unreliable, the edge between the two configuration item instances needs to be re-determined. Alternatively, the previously marked edge can be deleted and the association between the attributes of the two configuration item instances can be re-determined using the multi-level association strategy. If it is reliable, it is retained. Alternatively, the association between the attributes of the two configuration item instances can be re-determined based on subsequent new data, that is, the edge between the two configuration item instances can be re-determined.

[0080] The multi-strategy driven configuration item automatic association method provided by the present invention preliminarily screens the original configuration item instances according to scenario requirements, and defines multi-level association strategies. The configuration item instances that pass the preliminary screening are screened according to the association strategies of each level, and the relationship edge between any two configuration item instances is determined. The relationship edge is used to represent the association relationship between the attributes of the two configuration item instances. Then, based on the association relationship between the configuration item instance and the attributes of any two configuration item instances, a configuration item topology graph is constructed to further determine whether the above-mentioned determined association relationship is reliable, thereby improving the maintainability and scalability of the data system.

[0081] Optionally, based on scenario requirements, the attributes of each marked configuration item instance are preliminarily screened to obtain the configuration item instance to be analyzed and the attributes to be analyzed, including:

[0082] Based on the prior library, the attributes of each configuration item instance are marked according to the attribute name type and attribute value type;

[0083] Based on the scenario requirements, the attribute name type and / or attribute value type that needs to be associated is determined, and the attributes of each configuration item instance are preliminarily screened to obtain the configuration item instance to be analyzed and the attributes to be analyzed.

[0084] Specifically, the present invention analyzes configuration item instances, focusing on specific attributes and specific attribute value types. The specific attributes and specific attribute classifications herein refer to the attributes and their classifications that require further analysis. The attribute types of general enterprise configuration items are limited. Therefore, an a priori library is established based on the attribute name types and attribute value types included in the attributes of the configuration item instances. The attribute classifications in the a priori library can include multiple attribute name types and attribute value types, for example, as shown below:

[0085] Classification illustrate Attributes are organization-related Such as administrator, maintenance personnel, contact information, department, etc. Attributes are related to costs Such as cost, finance, order, fees, payment, etc. Attributes are development related Such as development tools, development environment, protocols, etc. Attributes are security related Such as permissions, authorization, security policies, identity authentication, etc. Attributes are position-dependent Such as geographic information, region, location, computer room, venue, etc. Attribute values ​​are time-dependent Such as date, time, etc. Attribute values ​​associated with enumerations Such as Boolean type, status field, etc. Attribute values ​​are related to numerical values Such as percentages, decimals, memory capacity, hard disk capacity, etc.

[0086] Based on the prior library, the attributes of each configuration item instance are marked according to the attribute name type and attribute value type.

[0087] Based on the scenario requirements, determine which attributes in the configuration item instance need to be associated, which is conducive to subsequent analysis, and which attributes do not need to be associated, that is, those attributes have no practical application significance when they are associated.

[0088] From an application perspective, attributes that don't need to be associated fall into two categories: First, attributes with generally meaningless associations, such as personnel, organization, expense, and development, are useless in specific scenarios, such as intelligent operations and maintenance. Second, attributes that require user intervention to establish relationships. For example, location-related attributes have multiple definitions. Establishing relationships for all of them could lead to an explosion in the number of relationships, complicating the topology and hindering intelligent management. Therefore, a balance needs to be struck when selecting attributes to associate.

[0089] Optionally, based on scenario requirements, after initially screening the attributes of each marked configuration item instance to obtain the attributes to be analyzed of the configuration item instance to be analyzed, the following steps may be performed:

[0090] Compare the attribute to be analyzed of the first configuration item instance with the attribute to be analyzed of the second target configuration item, and select all attribute pairs that satisfy equal attribute values ​​as a first candidate set; the attribute pair consists of the first configuration item instance, the first attribute, the second configuration item instance, and the second attribute, the first attribute is any attribute to be analyzed of the first configuration item instance, and the second attribute is one of the attributes to be analyzed of the second configuration item instance that satisfies the attribute value equal to that of the first attribute.

[0091] Specifically, through the above method, after determining which attributes of the configuration item instances need to be associated, that is, determining which attributes of any two configuration item instances correspond to the relationship edge, it is also necessary to align the attribute values ​​of any two configuration item instances. Specifically, obtain any two configuration item instances among the attributes to be analyzed of the above-mentioned configuration item instance to be analyzed, such as the first configuration item instance and the second configuration item instance. Generally, these two configuration item instances are different. Compare the attributes of the first configuration item instance with the attributes of the second configuration item instance. When the attribute values ​​of the two attributes are determined to be equal, they are regarded as an attribute pair, which can be represented as: configuration item instance-attribute-attribute-configuration item instance. The attribute values ​​of the attributes of all configuration item instances are determined in such a way that the attribute values ​​are equal, that is, the first candidate set is determined, which can be represented as candidate set C1. The equality of attribute values ​​indicates that the attributes of the two configuration item instances may be associated with each other.

[0092] Optionally, performing association analysis on the to-be-analyzed attribute of the first configuration item instance and the to-be-analyzed attribute of the second configuration item instance based on the multi-level association strategy, and determining a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge includes:

[0093] In the first candidate set, based on the attribute name alignment strategy included in the multi-level association strategy, attribute pairs with the same attribute name are selected as the second candidate set;

[0094] Marking a first relationship edge between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the second candidate set;

[0095] In the third candidate set, based on the attribute name matching strategy included in the multi-level association strategy, attribute pairs that satisfy the attribute name mapping to the namespace of the same attribute category are screened as the fourth candidate set; the third candidate set is the intersection of the complement of the second candidate set and the first candidate set;

[0096] A second relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the fourth candidate set.

[0097] Specifically, multi-level association strategies include attribute name alignment, attribute name matching, sequence pattern matching, and semantic matching, and are arranged from high to low, or from top to bottom. The terms "high" and "low" here represent the order in which the corresponding association strategies are executed. Association strategies at higher levels, or those at higher levels, are executed earlier; association strategies at lower levels, or those at lower levels, are executed later. This means that the association strategies at the later levels must be executed after the previous ones have been executed.

[0098] In the first candidate set, attribute name matching is first performed on the first and second configuration item instances based on the attribute name alignment strategy. Specifically, attribute name matching is performed on all attribute pairs included in the first candidate set. Specifically, if two attribute names are identical, a relationship edge is formed between the first and second configuration item instances, labeled as a first relationship edge. This first relationship edge indicates an association between an attribute of the first configuration item instance and an attribute of the second configuration item, also known as an attribute name alignment relationship edge. Furthermore, the labeling of this first relationship edge is used for a posteriori evaluation. This also determines the second candidate set, consisting of all attribute pairs that satisfy the same attribute name. Figure 3 Schematic diagram of the attribute name alignment strategy provided by an embodiment of the present invention. If two attribute names are different, then such attribute pairs in the first candidate set proceed to the next step, which is attribute name matching. Figure 4 This is a schematic diagram of an attribute name matching strategy provided by an embodiment of the present invention.

[0099] The second candidate set is removed from the first candidate set to obtain the third candidate set. In the third candidate set, the association relationship between the first configuration item instance and the second configuration item instance is determined according to the attribute name matching strategy. Specifically, according to the unified namespace, the attribute name of the first configuration item instance is mapped to the namespace of the same attribute category to which the attribute name belongs. For example, the unified namespace includes multiple different attribute categories, such as db, ip, floor, device, app, etc., where each attribute category corresponds to multiple different attribute names. Take an attribute category ip (or attribute category A) as an example, which includes attribute names such as severIP, NodesIP, and RoutIP. The attribute name of the first configuration item instance is severIP, and similarly, the attribute name of the second configuration item instance is NodesIP. Therefore, both the first configuration item instance and the second configuration item instance can be mapped to the attribute category ip. This means that the attribute name severIP of the first configuration item instance and the attribute name NodesIP of the second configuration item instance satisfy the attribute name matching strategy. A relationship edge is formed between the first configuration item instance and the second configuration item instance, which is marked as the second relationship edge. The second relationship edge is also called a keyword matching and attribute category relationship edge, which is used to indicate that there is an association between the attribute name severIP of the first configuration item instance and the attribute name NodesIP of the second configuration item instance. In addition, the label is also used for a posteriori evaluation. At the same time, the fourth candidate set consisting of attribute pairs that satisfy the attribute name matching procedure is determined.

[0100] In addition, the attribute names belonging to the same attribute category in the unified namespace may also include different attribute names expressed in the form of synonyms, derivatives, strongly related words, abbreviations and affixes. Among them, abbreviations include db, vm, etc., and affixes include IP, ip, etc. For example, if the attribute name of the first configuration item instance is IP_a and the attribute name of the second configuration item instance is IP_b, then the attribute name IP_a of the first configuration item instance and the attribute name IP_b of the second configuration item instance satisfy the affix matching in the attribute name matching strategy. Affix matching is a type of keyword matching, that is, it satisfies the attribute name matching strategy. A relationship edge is formed between the first configuration item instance and the second configuration item instance, marked as a second relationship edge. The second relationship edge is also called a keyword matching and attribute category relationship edge, which is used to indicate that there is an association between the attribute name IP_a of the first configuration item instance and the attribute name IP_b of the second configuration item instance.

[0101] Alternatively, in the unified namespace, the attribute names of an attribute category B include device, serialnumber, and assetserial. The attribute name of the first configuration item instance is assetserial, and the attribute name of the second configuration item instance is device. Both the attribute name of the first configuration item instance and the attribute name of the second configuration item instance can be mapped to device. Then, the attribute name assetserial of the first configuration item instance and the attribute name device of the second configuration item instance satisfy the strong related word matching in the attribute name matching strategy. Strong related word matching is a type of keyword matching, that is, it satisfies the attribute name matching strategy. A relationship edge is formed between the first configuration item instance and the second configuration item instance, marked as a second relationship edge. This second relationship edge is also called a keyword matching and attribute category relationship edge, and is used to indicate that there is an association between the attribute name assetserial of the first configuration item instance and the attribute name device of the second configuration item instance.

[0102] Optionally, the performing association analysis on the to-be-analyzed attribute of the first configuration item instance and the to-be-analyzed attribute of the second configuration item instance based on the multi-level association strategy to determine a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge further includes:

[0103] Determine a complement of the union of the second candidate set and the fourth candidate set, and an intersection of the complement and the first candidate set as a fifth candidate set;

[0104] In the fifth candidate set, based on the multi-level association strategy including a sequential pattern matching strategy, the attribute pairs whose sequence similarity satisfies a first threshold are screened as a sixth candidate set; the sequence similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair;

[0105] Alternatively, in the fifth candidate set, based on the multi-level association strategy including a semantic matching strategy, the attribute pairs whose semantic similarity meets the second threshold are screened as the seventh candidate set; the semantic similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair, which are respectively mapped to the first word vector and the second word vector in the pre-trained word vector space;

[0106] Acquire the attribute pair included in the sixth candidate set or the attribute pair included in the seventh candidate set as the target attribute pair;

[0107] A third relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the target attribute pair.

[0108] Specifically, through the above-mentioned attribute name alignment and attribute name matching, it is possible that the attributes of the first configuration item instance and the attributes of the second configuration item instance are not yet associated. It is necessary to further associate these attributes according to the sequence pattern matching strategy or semantic matching strategy provided by the present invention.

[0109] Intersecting the complement of the union of the second candidate set and the fourth candidate set with the first candidate set, and obtaining the intersection as the fifth candidate set;

[0110] In the fifth candidate set, according to the sequential pattern matching strategy, the sequence similarity between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance is determined, and whether this sequence similarity meets the first threshold is determined. If so, an edge is formed between the first configuration item instance and the second configuration item instance, and is marked as a third relationship edge, which is used to indicate that there is an association between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance. This third relationship edge is also called a sequence match and an attribute category, and this label is used for a posteriori evaluation. Specifically, the similarity between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance can be determined using SequenceMatcher, Cosine similarity, Jaccard similarity, Levenshtein distance, Jaro-Winkler distance, etc.

[0111] Alternatively, in the fifth candidate set, according to the semantic matching strategy, the semantic similarity between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance is determined, and whether the semantic similarity meets the second threshold is determined. If so, an edge is formed between the first configuration item instance and the second configuration item instance, and is marked as a third relationship edge, which is used to indicate that there is an association between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance. The third relationship edge represents sequence matching and attribute category, and the label is used for posterior evaluation. In the semantic matching strategy, the semantic similarity between the attribute name A of the first configuration item instance and the attribute name B of the second configuration item instance is determined, specifically by mapping the attribute name A of the first configuration item instance to the pre-trained word vector space to obtain a first word vector corresponding to the attribute name A. Similarly, the attribute name B of the second configuration item instance is mapped to the pre-trained word vector space to obtain a second word vector corresponding to the attribute name B. The similarity between the first word vector and the second word vector is determined. If the similarity meets the second threshold, an edge is formed between the first configuration item instance and the second configuration item instance, and is marked as a third relationship edge. There's an association between attribute name A, representing the first configuration item instance, and attribute name B, representing the second configuration item instance. This third edge represents sequence matching and attribute category, and this tag is used for posterior evaluation. The pretrained word vector space is built based on the SentenceTransformer model and trained using a large-scale corpus. It maps each word to a word vector in a high-dimensional vector space. These word vectors typically contain semantic information and can therefore be used to calculate similarity between words.

[0112] After establishing an association relationship between the attributes of the first configuration item instance and the attributes of the second configuration item instance through the aforementioned multi-level association strategies, including attribute name alignment, attribute name matching, sequence pattern matching, and semantic matching, the relationship between some attributes may still be unclear. This can be accomplished through special case analysis, analyzing the attributes of the first configuration item instance and the attributes of the second configuration item instance to establish an association relationship. Of course, special case analysis is also included in the aforementioned multi-level association strategy and is generally performed last. For example, some attribute names cannot be parsed solely from the attribute name. Further analysis, including the configuration item name, attribute value type, and attribute value, is required to establish specific rules and algorithm models. Figure 5 This is a schematic diagram of the special case analysis strategy provided by an embodiment of the present invention. Special case analysis is a necessary step in a multi-level association strategy, targeting long-tail scenarios. However, this special case analysis strategy has a very low input-output ratio. Therefore, in actual implementation, this step can be performed or skipped based on the attribute association requirements in the scenario.

[0113] Optionally, determining the evaluation score of the relationship edge in the configuration item topology graph based on the relationship edge evaluation strategy and the node evaluation strategy includes:

[0114] Determine, based on the order corresponding to the multi-level association strategies, a first evaluation score of the first relationship edge, a second evaluation score of the second relationship edge, and a third evaluation score of the third relationship edge; the first evaluation score is the highest, the second evaluation score is the second highest, and the third evaluation score is the lowest;

[0115] determining, based on the number of relationship edges between the first configuration item instance and the second configuration item instance, a fourth evaluation score of the relationship edges in the configuration item topology graph;

[0116] An evaluation score of a relationship edge in the configuration item topology graph is determined based on a weighted sum algorithm and the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score.

[0117] Specifically, through the above multi-level association strategy, the attributes of the first configuration item instance and the attributes of the second configuration item instance are sequentially subjected to different levels of association strategies. After the relationship edges between the first configuration item instance and the second configuration item instance are determined, the reliability of these relationship edges needs to be further evaluated. Specifically, this can be done through two aspects:

[0118] 1. Relationship Edge Evaluation Strategy

[0119] Based on the relationship edge between the first configuration item instance and the second configuration item instance determined in the above steps, and the order corresponding to the multi-level association strategies adopted to determine these relationship edges, the relationship edges determined by the multi-level association strategies of different orders are evaluated and obtained with different scores. For example, the higher the level or the higher the level of the association strategy, that is, the earlier the association strategy is executed, the simpler the association strategy to be used for the association is, and the higher the evaluation score of the corresponding relationship edge is; the lower the level or the lower the level of the association strategy, that is, the later the association strategy is executed, the more complex the association strategy to be used for the association is, and the lower the evaluation score of the corresponding relationship edge is. For example, the first evaluation score of the first relationship edge, the second evaluation score of the second relationship edge, and the third evaluation score of the third relationship edge are determined; the first evaluation score is the highest, the second evaluation score is the second, and the third evaluation score is the lowest.

[0120] The order corresponding to the multi-level association strategy here is the order in which the attributes to be analyzed are associated, with attribute values ​​taking precedence and attribute names taking second place. The association strategy for attribute names includes attribute name alignment and attribute name matching; attribute name matching specifically includes keyword matching and sequence matching; sequence matching includes sequence pattern matching and semantic matching.

[0121] 2. Node evaluation strategy

[0122] A relationship edge connects two configuration item instances, where the configuration item instances are nodes. After establishing a configuration item topology graph based on the relationship edges corresponding to the attributes between all the configuration item instances mentioned above, that is, after determining the configuration item topology graph between any two categories of configuration item instances as nodes, a score can be assigned based on the number of links, with a lower score corresponding to a higher number of links. A link here is the relationship edge between two configuration item instances.

[0123] Based on the number of relationship edges between the first configuration item instance and the second configuration item instance, a fourth evaluation score of the relationship edge in the configuration item topology graph is determined.

[0124] All the above evaluation scores are combined and compared with the preset trust threshold. If it is determined that all evaluation scores meet the preset trust threshold, that is, all evaluation scores are greater than or equal to the preset trust threshold, then the relationship edge is marked as reliable; if it is determined that all evaluation scores do not meet the preset trust threshold, that is, all evaluation scores are less than the preset trust threshold, then the relationship edge is marked as unreliable and output for manual analysis. The relationship edge can also be deleted and the association relationship between the first configuration item instance and the second configuration item instance is rebuilt, that is, the relationship edge between the first configuration item instance and the second configuration item instance is re-determined.

[0125] The multi-strategy driven configuration item automatic association method provided by the present invention preliminarily screens the original configuration item instances according to scenario requirements, and defines multi-level association strategies. The configuration item instances that pass the preliminary screening are screened according to the association strategies of each level, and the relationship edge between any two configuration item instances is determined. The relationship edge is used to represent the association relationship between the attributes of the two configuration item instances. Then, based on the association relationship between the configuration item instance and the attributes of any two configuration item instances, a configuration item topology graph is constructed to further determine whether the above-mentioned determined association relationship is reliable, thereby improving the maintainability and scalability of the data system.

[0126] The advantages of the multi-policy driven configuration item automatic association method provided by the present invention include:

[0127] 1. Prior filtering: Filter objects with potential for correlation and pass them on to downstream tasks for correlation analysis. Similar to information search, this approach uses a simple and unpretentious strategy to quickly focus on target objects.

[0128] 2. Association analysis: We established an analysis process from simple to complex, using a divide-and-conquer approach. We established different strategies for alignment calculations for different situations.

[0129] 3. Posteriori evaluation: Evaluate the validity of the relationship edge. Here, the process of establishing the relationship edge is traced back and analyzed, and the evaluation score of the relationship edge is calculated based on the topological structure.

[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-strategy driven configuration item automatic association method, characterized in that: include: Based on scenario requirements, the attributes of each marked configuration item instance are initially screened to obtain the configuration item instances to be analyzed and their attributes to be analyzed; The configuration item instance to be analyzed includes a plurality of attributes and attribute values ​​of the plurality of attributes; Based on a multi-level association strategy, an association analysis is performed on the attribute to be analyzed of the first configuration item instance and the attribute to be analyzed of the second configuration item instance, and a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge are determined; the multi-level association strategy is used to define a rule for associating the attributes to be analyzed in a manner of prioritizing attribute values ​​and ranking attributes; the first configuration item instance and the second configuration item instance are any one of the configuration item instances to be analyzed, and belong to different categories of configuration items; Constructing a configuration item topology graph based on the relationship edge between the first configuration item instance and the second configuration item instance, the first configuration item instance, and the second configuration item instance; Determining an evaluation score of the configuration item topology graph based on a relationship edge evaluation strategy and a node evaluation strategy; Based on whether the evaluation score of the configuration item topology graph meets a preset trust threshold, it is determined whether it is necessary to redetermine the relationship edge between the first configuration item instance and the second configuration item instance.

2. The multi-strategy driven configuration item automatic association method according to claim 1, characterized in that: Based on the scenario requirements, the attributes of each marked configuration item instance are initially screened to obtain the configuration item instance to be analyzed and its attributes to be analyzed, including: Based on the prior library, the attributes of each configuration item instance are marked according to the attribute name type and attribute value type; Based on the scenario requirements, the attribute name type and / or attribute value type that needs to be associated is determined, and the attributes of each configuration item instance are preliminarily screened to obtain the configuration item instance to be analyzed and its attributes to be analyzed.

3. The multi-strategy driven configuration item automatic association method according to claim 2, characterized in that: Based on the scenario requirements, the attributes of each marked configuration item instance are initially screened to obtain the configuration item instance to be analyzed and its attributes to be analyzed, including: Compare the attribute to be analyzed of the first configuration item instance with the attribute to be analyzed of the second target configuration item, and select all attribute pairs that satisfy equal attribute values ​​as a first candidate set; the attribute pair consists of the first configuration item instance, the first attribute, the second configuration item instance, and the second attribute, the first attribute is any attribute to be analyzed of the first configuration item instance, and the second attribute is one of the attributes to be analyzed of the second configuration item instance that satisfies the attribute value equal to that of the first attribute.

4. The multi-strategy driven configuration item automatic association method according to claim 3, characterized in that: The step of performing association analysis on the attribute to be analyzed of the first configuration item instance and the attribute to be analyzed of the second configuration item instance based on the multi-level association strategy, and determining a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge, includes: In the first candidate set, based on the attribute name alignment strategy included in the multi-level association strategy, attribute pairs with the same attribute name are selected as the second candidate set; Marking a first relationship edge between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the second candidate set; In the third candidate set, based on the attribute name matching strategy included in the multi-level association strategy, attribute pairs that satisfy the attribute name mapping to the namespace of the same attribute category are screened as the fourth candidate set; the third candidate set is the intersection of the complement of the second candidate set and the first candidate set; A second relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the attribute pair in the fourth candidate set.

5. The multi-strategy driven configuration item automatic association method according to claim 4, characterized in that: The method further includes performing association analysis on the attribute to be analyzed of the first configuration item instance and the attribute to be analyzed of the second configuration item instance based on the multi-level association strategy, determining a relationship edge between the first configuration item instance and the second configuration item instance and a label corresponding to the relationship edge, and: Determine a complement of the union of the second candidate set and the fourth candidate set, and an intersection of the complement and the first candidate set as a fifth candidate set; In the fifth candidate set, based on the multi-level association strategy including a sequential pattern matching strategy, the attribute pairs whose sequence similarity satisfies a first threshold are screened as a sixth candidate set; the sequence similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair; Alternatively, in the fifth candidate set, based on the multi-level association strategy including a semantic matching strategy, the attribute pairs whose semantic similarity meets the second threshold are screened as the seventh candidate set; the semantic similarity is the similarity between the attribute name of the first configuration item instance and the attribute name of the second configuration item instance corresponding to the attribute pair, which are respectively mapped to the first word vector and the second word vector in the pre-trained word vector space; Acquire the attribute pair included in the sixth candidate set or the attribute pair included in the seventh candidate set as the target attribute pair; A third relationship edge is marked between the first configuration item instance and the second configuration item instance corresponding to the target attribute pair.

6. The multi-strategy driven configuration item automatic association method according to claim 5, characterized in that: Determining the evaluation score of the relationship edge in the configuration item topology graph based on the relationship edge evaluation strategy and the node evaluation strategy includes: Determine, based on the order corresponding to the multi-level association strategies, a first evaluation score of the first relationship edge, a second evaluation score of the second relationship edge, and a third evaluation score of the third relationship edge; the first evaluation score is the highest, the second evaluation score is the second highest, and the third evaluation score is the lowest; determining, based on the number of relationship edges between the first configuration item instance and the second configuration item instance, a fourth evaluation score of the relationship edges in the configuration item topology graph; An evaluation score of a relationship edge in the configuration item topology graph is determined based on a weighted sum algorithm and the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score.

7. The multi-strategy driven configuration item automatic association method according to claim 6, characterized in that: The determining whether it is necessary to redetermine the relationship edge between the first configuration item instance and the second configuration item instance based on whether the evaluation score of the configuration item topology graph meets a preset trust threshold includes: If the evaluation score of the configuration item topology graph meets a preset trust threshold, the relationship edge is marked as reliable; If the evaluation score of the configuration item topology graph does not meet the preset trust threshold, the relationship edge is marked as unreliable and output for manual analysis.

8. The multi-strategy driven configuration item automatic association method according to claim 4, characterized in that: In the third candidate set, based on the attribute name matching strategy included in the multi-level association strategy, attribute pairs that satisfy the attribute name mapping to the namespace of the same attribute category are screened as the fourth candidate set, including: Construct a unified namespace based on all synonyms, derivatives, strongly related words, abbreviations and affixes included in each attribute category; In the third candidate set, the first configuration item instance and the second configuration item instance corresponding to the attribute pair are obtained as the first target configuration item instance and the second target configuration item instance respectively; Mapping the attribute name of the first target configuration item instance to the unified namespace to obtain a first attribute name corresponding to the attribute name of the first target configuration item instance; Mapping the attribute name of the second target configuration item instance to the unified namespace to obtain a second attribute name corresponding to the attribute name of the second target configuration item instance; If the first attribute name and the second attribute name are the same, it is determined that the first target configuration item instance and the second target configuration item instance meet the attribute name matching strategy, and are added to the fourth candidate set.

9. The multi-strategy driven configuration item automatic association method according to claim 5, characterized in that: The sequence pattern matching strategies include SequenceMatcher, Cosine similarity, Jaccard similarity, Levenshtein distance, and Jaro-Winkler distance.

10. The multi-strategy driven configuration item automatic association method according to claim 5, characterized in that: The pre-trained word vector space is constructed based on the SentenceTransformer model.

Citation Information

Patent Citations

  • Fast topology method and apparatus of IT infrastructure

    CN106375120A

  • Creating a correlation rule defining a relationship between event types

    WO2012138319A1