Knowledge Graph Construction Method, Device, Electronic Device, and Readable Storage Medium

By dividing the analysis fields into strings in the power communication network and counting them in different levels and security partitions to determine their complete indicators and comprehensive importance, the problem of poor quality of the security knowledge graph of the power communication network is solved, and the accuracy and reliability of the knowledge graph are improved.

CN120146174BActive Publication Date: 2025-07-29国网思极网安科技(北京)有限公司 +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510632705.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-29
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

When the existing methods build the power communication network security knowledge graph, due to the quality of data model design and unclear business needs, the data quality in the high-quality database is not high, which in turn makes the power communication network security knowledge graph poor.

Method used

The obtained analysis fields are divided into strings, and statistics are performed in different levels and security partitions to determine the complete indicators and comprehensive importance of the strings, build a knowledge graph through quality indicators, and give priority to strings with quality indicators above the threshold.

Benefits of technology

It improves the accuracy and reliability of the knowledge graph, provides a strong data foundation for the management and application of knowledge, and promotes the in-depth mining of information and the improvement of application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146174B_ABST
    Figure CN120146174B_ABST
Patent Text Reader

Abstract

The present application provides a method, an apparatus, an electronic device, and a readable storage medium for constructing a knowledge graph. The method includes: dividing each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; determining a complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; determining a comprehensive importance degree of the string according to the semantic vector corresponding to the string; determining a quality index of the string based on the complete index and the comprehensive importance degree; and constructing a knowledge graph based on the quality index of the string. The embodiments of the present application improve the accuracy and reliability of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of knowledge graphs, and particularly to a method, device, electronic device and readable storage medium for constructing a knowledge graph. Background Art

[0002] In the wave of digital transformation, the security of the power communication network is directly related to the development of grid intelligence and the stable operation of the power system. With the continuous evolution and complexity of network attack means, traditional network security defense measures are difficult to cope with the increasingly severe network security challenges. In order to improve the network security protection ability of the power system and the ability to respond to network security incidents, constructing a power communication network security knowledge graph provides artificial intelligence services for power communication network management, risk monitoring, etc.

[0003] The quality of graph data is a key factor in evaluating the success of constructing a power communication network security knowledge graph. Existing methods screen data manually and then construct a high-quality database, and construct a knowledge graph by obtaining power communication network security data from the high-quality database. However, due to problems such as the quality of data model design and unclear business requirements, the data quality in the high-quality database is not high, and thus the quality of the power communication network security knowledge graph is not good. Summary of the Invention

[0004] In view of this, the purpose of this application is to propose a method, device, electronic device and readable storage medium for constructing a knowledge graph.

[0005] Based on the above purpose, this application provides a method for constructing a knowledge graph, including:

[0006] Dividing each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions;

[0007] Determining the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition;

[0008] Determining the comprehensive importance of the string according to the semantic vector corresponding to the string;

[0009] Determining the quality index of the string based on the complete index and the comprehensive importance;

[0010] Constructing a knowledge graph based on the quality index of the string.

[0011] In a possible implementation, determining the complete index of the string according to the number of times the string appears in each hierarchical text set and the number of times the string appears in the partition text set includes:

[0012] Determining the hierarchical identity index of the string according to the discrete index of the number of times the string appears in each hierarchical text set and the number of times the string appears in all hierarchical text sets;

[0013] Determining the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference degree of the string in all partition text sets corresponding to the level;

[0014] Determining the complete index of the string based on the hierarchical identity index and the hierarchical isolation index.

[0015] In a possible implementation, determining the hierarchical identity index of the string according to the discrete index of the number of times the string appears in each hierarchical text set and the number of times the string appears in all hierarchical text sets includes:

[0016] Determining the quantity proportion according to the ratio of the number of times the string appears in each hierarchical text set to the total number of strings in the hierarchical text set;

[0017] Calculating the cumulative sum of the quantity proportions of the string in all hierarchical text sets to obtain the hierarchical importance;

[0018] Calculating the hierarchical dispersion according to the discrete index of the number of times the string appears in all hierarchical text sets;

[0019] Determining the hierarchical identity index according to the hierarchical importance and the hierarchical dispersion.

[0020] In a possible implementation, determining the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference degree of the string in all partition text sets corresponding to the level includes:

[0021] Calculating the partition dispersion according to the discrete index of the number of times the string appears in the partition text set;

[0022] Calculating the cumulative sum of the difference degrees between the number of times the string appears in a certain partition text set and the number of times the string appears in the remaining partition text sets other than the certain partition text set in the corresponding hierarchical text set to obtain the appearance difference degree;

[0023] Calculate the cumulative sum of the occurrence difference degrees of the said string in all partition text sets of the corresponding hierarchical text set to obtain the partition difference degree;

[0024] Based on the said partition dispersion degree and the said partition difference degree, determine the said hierarchical isolation index.

[0025] In a possible implementation manner, the determining the comprehensive importance degree of the said string according to the semantic vector corresponding to the said string includes:

[0026] Determine the local importance degree according to the similarity between the semantic vector corresponding to the said string and the target vectors of the said string at all levels;

[0027] Determine the comprehensive importance degree based on the said local importance degree.

[0028] In a possible implementation manner, the determining the quality index of the said string based on the said complete index and the said comprehensive importance degree includes:

[0029] Perform a normalization calculation on the said complete index and the said comprehensive importance degree to obtain the said quality index.

[0030] In a possible implementation manner, the constructing a knowledge graph based on the quality index of the said string includes:

[0031] Use the said strings corresponding to the quality index higher than the preset threshold to construct the said knowledge graph.

[0032] Based on the same inventive concept, an embodiment of the present application further provides a knowledge graph construction device, including:

[0033] A division module, configured to divide each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions;

[0034] A complete index determination module, configured to determine the complete index of the said string according to the number of times the said string appears in each hierarchical text set and the number of times the said string appears in the partition text set; the hierarchical text set is a set composed of strings of all the said security partitions within each level; the partition text set is a set composed of strings of each said security partition; [[ID=3,5]]

[0035] A comprehensive importance degree determination module, configured to determine the comprehensive importance degree of the said string according to the semantic vector corresponding to the said string;

[0036] A quality index determination module, configured to determine the quality index of the said string based on the said complete index and the said comprehensive importance degree;

[0037] A building module configured to build a knowledge graph based on the quality metrics of the string.

[0038] In a possible implementation, the complete metric determination module is further configured to:

[0039] A same-level metric determination unit configured to determine the same-level metric of the string according to the discrete metric of the number of times the string appears in each hierarchical text set and the number of times the string appears in all hierarchical text sets;

[0040] A hierarchical isolation metric determination unit configured to determine the hierarchical isolation metric of the string according to the discrete metric of the number of times the string appears in the partitioned text set and the appearance difference degree of the string in all partitioned text sets corresponding to the level;

[0041] A complete metric determination unit configured to determine the complete metric of the string based on the same-level metric and the hierarchical isolation metric.

[0042] In a possible implementation, the same-level metric determination unit is further configured to:

[0043] Determine the quantity proportion according to the ratio of the number of times the string appears in each hierarchical text set to the total number of strings in the hierarchical text set;

[0044] Calculate the cumulative sum of the quantity proportions of the string in all hierarchical text sets to obtain the hierarchical importance;

[0045] Calculate the hierarchical dispersion according to the discrete metric of the number of times the string appears in all hierarchical text sets;

[0046] Determine the same-level metric according to the hierarchical importance and the hierarchical dispersion.

[0047] In a possible implementation, the hierarchical isolation metric determination unit is further configured to:

[0048] Calculate the partition dispersion according to the discrete metric of the number of times the string appears in the partitioned text set;

[0049] Calculate the cumulative sum of the difference degrees between the number of times the string appears in a certain partitioned text set and the number of times the string appears in the remaining partitioned text sets other than the certain partitioned text set in the corresponding hierarchical text set to obtain the appearance difference degree;

[0050] Calculate the cumulative sum of the appearance difference degrees of all partitioned text sets in the corresponding hierarchical text set to obtain the partition difference degree;

[0051] Determine the hierarchical isolation index based on the partition discreteness and the partition difference degree.

[0052] In a possible implementation, the comprehensive importance determination module is further configured to:

[0053] Determine the local importance according to the similarities between the semantic vectors corresponding to the strings and the target vectors of the strings at all levels.

[0054] Determine the comprehensive importance based on the local importance.

[0055] In a possible implementation, the quality index determination module is further configured to:

[0056] Perform a normalization calculation on the complete index and the comprehensive importance to obtain the quality index.

[0057] In a possible implementation, the construction module is further configured to:

[0058] Construct the knowledge graph by using the strings corresponding to the quality index higher than a preset threshold.

[0059] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the knowledge graph construction method described in any one of the above is implemented.

[0060] Based on the same inventive concept, an embodiment of the present application further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the knowledge graph construction method described in any one of the above.

[0061] As can be seen from the above, the knowledge graph construction method, device, electronic device, and readable storage medium provided by this application divide each acquired analysis field into strings; the analysis fields are located in different levels and different security partitions; according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set, determine the complete index of the string; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; determine the comprehensive importance of the string according to the semantic vector corresponding to the string; determine the quality index of the string based on the complete index and the comprehensive importance; construct a knowledge graph based on the quality index of the string. In the embodiments of this application, each acquired analysis field is split into strings and statistically analyzed in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string in each level text set, but also its performance in a specific security partition. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality evaluation. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By normalizing the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that during the knowledge graph construction process, strings with quality indexes higher than the set threshold are preferentially selected. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data foundation for knowledge management and application, ultimately promoting the in-depth mining of information and the improvement of application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in this application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0063] Figure 1 It is a schematic flowchart of the knowledge graph construction method according to the embodiment of this application;

[0064] Figure 2 It is a schematic structural diagram of the knowledge graph construction device according to the embodiment of this application;

[0065] Figure 3 It is a schematic structural diagram of the electronic device according to the embodiment of this application. Detailed implementation manners

[0066] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.

[0067] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the field to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "include" or "comprise" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connect" or "be connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0068] It can be understood that, before using the technical solutions of the various embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0069] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solutions of the present application according to the prompt message.

[0070] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0071] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present application, and other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present application.

[0072] As described in the background art section, in the wave of digital transformation, the security of the power communication network is directly related to the development of grid intelligence and the stable operation of the power system. With the continuous evolution and complexity of network attack means, traditional network security defense measures are difficult to cope with the increasingly severe network security challenges. In order to improve the network security protection ability of the power system and the ability to respond to network security incidents, a power communication network security knowledge graph is constructed to provide artificial intelligence services for power communication network management, risk monitoring, etc.

[0073] The quality of the graph data is a key factor in evaluating the success of constructing a power communication network security knowledge graph. Existing methods screen data manually and then construct a high-quality database, and construct a knowledge graph by obtaining power communication network security data from the high-quality database. However, due to problems such as the quality of data model design and unclear business requirements, the data quality in the high-quality database is not high, which in turn leads to poor quality of the power communication network security knowledge graph.

[0074] In view of the above considerations, an embodiment of the present application proposes a method for constructing a knowledge graph. Each obtained analysis field is divided into strings. The analysis fields are located in different levels and different security partitions. According to the number of times the string appears in each level text set and the number of times the string appears in the partition text set, the complete index of the string is determined. The level text set is a set composed of strings of all the security partitions within each level. The partition text set is a set composed of strings of each security partition. According to the semantic vector corresponding to the string, the comprehensive importance of the string is determined. Based on the complete index and the comprehensive importance, the quality index of the string is determined. Based on the quality index of the string, a knowledge graph is constructed. In the embodiment of the present application, each obtained analysis field is split into strings and statistically analyzed in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string appearing in each level text set, but also considers its performance in a specific security partition. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality evaluation. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By normalizing the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that in the process of constructing the knowledge graph, strings with quality indexes higher than the set threshold are preferentially selected. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data basis for the management and application of knowledge, ultimately promoting the in-depth mining of information and the improvement of application value.

[0075] The following will detail the technical solutions of the embodiments of the present application through specific embodiments.

[0076] Reference Figure 1 , the method for constructing a knowledge graph according to the embodiment of the present application includes the following steps:

[0077] Step S101: Divide each obtained analysis field into strings. The analysis fields are located in different levels and different security partitions.

[0078] Step S102: Determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set. The level text set is a set composed of strings of all the security partitions within each level. The partition text set is a set composed of strings of each security partition.

[0079] Step S103, determine the comprehensive importance of the string according to the semantic vector corresponding to the string;

[0080] Step S104, determine the quality index of the string based on the complete index and the comprehensive importance;

[0081] Step S105, construct a knowledge graph based on the quality index of the string.

[0082] For step S101, first, it is necessary to obtain the analysis fields of different levels and different security partitions, and then divide the obtained analysis fields into strings.

[0083] In this embodiment, obtain the analysis fields of different security partitions within each level of the power communication network security system, and denote the set composed of the analysis fields within all levels as the overall text set; divide each analysis field into different strings; denote the set composed of the strings corresponding to the analysis fields of all security partitions within each level as the level text set, and denote the set composed of the strings corresponding to the analysis fields of each security partition as the partition text set; randomly select a string from the strings corresponding to the analysis fields in the overall text set for subsequent description, and denote this string as the target string in the following embodiments.

[0084] Specifically, in the embodiment of the present application, the power communication network security system is divided into three levels: the boundary layer, the core layer, and the access layer, and at the same time, each level is divided into an intranet security area and an extranet security area.

[0085] The intranet security area of the boundary layer includes: all devices, systems, and resources within the power communication network, such as power plants, transmission systems, communication devices, etc., and text files such as core system logs, device configuration files, and traffic analysis reports can be obtained; the extranet security area includes: security devices such as external entry points, firewalls, intrusion detection systems, and anti-virus gateways, and text files such as connection request logs from the outside and anti-virus scan logs can be obtained.

[0086] The intranet security area of the core layer includes: core routers, switches, and important servers and storage devices connected to these devices, and more sensitive and important data such as power production data, real-time sensor data, and transmission system status information can be obtained; the extranet security area includes: advanced firewalls, intrusion prevention systems, traffic analysis tools, etc., and performance statistics of core routers and switches, traffic analysis data, and operation logs of core devices can be obtained.

[0087] The internal network security area of the access layer includes various terminal devices, such as end-user computers, smart meters, etc., which can obtain information such as security audit logs, access control lists, and terminal device configurations of end-user devices; the external network security area involves the security control of user access points, such as network access control systems, authentication and authorization systems, and security gateway devices, which may obtain access request records from user terminals, user authentication information, audit logs of network access control systems, etc.

[0088] Obtain the text files of each security partition within each layer of the power communication network security system, and use the Jieba algorithm to perform word segmentation on the text files of each security partition by combining dictionary matching and dynamic programming to obtain the keywords of each security partition. Select any one security partition as the partition to be tested, select any one keyword within the partition to be tested as the target word, and match the target word with the keywords of the remaining security partitions except the partition to be tested through the brute-force matching algorithm. If there is a keyword in the keywords of the remaining security partitions that matches the target word, then record the target word as the analysis field of the partition to be tested, and traverse the keywords of the partition to be tested to obtain the analysis field of the partition to be tested. According to the above method, obtain the analysis fields of each security partition within each layer of the power communication network security system.

[0089] Denote the set composed of the analysis fields within all layers as the overall text set; to accurately judge the importance of the analysis fields, randomly divide each analysis field into strings of different lengths. For the convenience of subsequent description, denote the set composed of the strings corresponding to the analysis fields of all security partitions within each layer as the layer text set, and denote the set composed of the strings corresponding to the analysis fields of each security partition as the partition text set; select any one string from the strings corresponding to the analysis fields in the overall text set as the target string.

[0090] Further, for step S102, determine the integrity index of the string according to the number of times the string appears in each layer text set and the number of times the string appears in the partition text set; the layer text set is the set composed of the strings of all the security partitions within each layer; the partition text set is the set composed of the strings of each security partition.

[0091] In some embodiments, determining the complete index of the string according to the number of times the string appears in each hierarchical text set and the number of times the string appears in the partitioned text set includes: determining the hierarchical identity index of the string according to the discrete index of the number of times the string appears in each hierarchical text set and the number of times the string appears in all hierarchical text sets; determining the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partitioned text set and the appearance difference degree of the string in all partitioned text sets corresponding to the hierarchy; determining the complete index of the string based on the hierarchical identity index and the hierarchical isolation index.

[0092] In this embodiment, the core layer of the power domain knowledge graph application architecture bears the capabilities of natural language processing, knowledge extraction, knowledge fusion, and knowledge processing. The power communication network security knowledge data of each hierarchical architecture usually uses a relational database to manage files, and the corresponding relationship of the data is a one-to-many relationship. In the network security knowledge scheduling of power communication, its security pre-plan behavior is carried out according to the requirements of relevant security management regulations, security response manuals and other documents; its security knowledge builds a basic architecture system from top to bottom. The power communication network security knowledge data is relatively complete within the architecture system. The more complete the data is used, the more perfect the connections and nodes of the knowledge graph construction will be; there are many data source channels, and different levels have different requirements for security data. The higher the level of the data, the higher the authority and the higher the data quality. Based on this, the data quality is judged.

[0093] Each level of the terminal and system of the power communication network security system is only allowed to be used within the corresponding level. It is dedicated to a private network, and the equipment designed to connect to the lower-level network cannot be randomly connected to the upper-level network. The network operations in the corresponding network knowledge security manual have corresponding and strict specification requirements. Moreover, in the power communication network security system, there is a horizontal security isolation protection between different security partitions, so that the terminal systems of different security partitions cannot directly access each other. Isolation is for security protection and to strengthen the transmission control of data; because the network data of different security partitions at the same level are not connected to each other, the network security-related knowledge requirements of different security partitions are also different. Based on this, the integrity of the target string is calculated.

[0094] In the process of constructing a knowledge graph, two key principles of entity identity and entity isolation need to be followed. Since the security requirements and protection measures of the power communication network security system are common among different levels, there is identity between the strings in the keyword fields of different levels, that is, the dispersion degree of the number of occurrences of the target string in the text sets of all levels is low. Since there is security isolation protection between the terminal systems of different security zones at the same level of the power communication network security system, resulting in differences in the data of different security zones at the same level, there are differences between the number of occurrences of the target string in the text sets of different corresponding zones at the same level. The more the target string conforms to entity identity and entity isolation, the more suitable the analysis field where the target string is located is for constructing a knowledge graph, and the higher the integrity of the meaning expressed by the target string in the power communication network security knowledge; analyzing the dispersion degree of the number of occurrences of the target string in the text sets of all levels and the differences between the number of occurrences of the target string in the text sets of different corresponding zones at each level to improve the accuracy of the integrity index of the target string.

[0095] In some embodiments, determining the hierarchical identity index of the string according to the number of occurrences of the string in the text sets of each level and the dispersion index of the number of occurrences of the string in the text sets of all levels includes: determining the quantity proportion according to the ratio of the number of occurrences of the string in each of the text sets of the level and the total number of strings in the text set of the level; calculating the cumulative sum of the quantity proportions of the string in all the text sets of the levels to obtain the hierarchical importance; calculating the hierarchical dispersion degree according to the dispersion index of the number of occurrences of the string in the text sets of all levels; and determining the hierarchical identity index according to the hierarchical importance and the hierarchical dispersion degree.

[0096] Since the security requirements and protection measures of the power communication network security system are common among different levels, there is identity between the strings in the keyword fields of different levels, that is, the number of identical strings in the text sets of different levels is relatively close. The identity of the string reflects the importance degree of the data or information in the system or network. The higher the identity of the string, the more significant the core value and role of the information in the specific context, and it should be in the core position of the knowledge graph of the power communication network security. If the dispersion degree of the number of occurrences of the target string in the text sets of all levels is smaller and the number of occurrences of the target string in the text sets of all levels is larger, it indicates that the target string conforms more to entity identity and the entity identity is more important, so as to obtain the hierarchical identity index of the target string.

[0097] It should be noted that standard deviation, variance, range, and coefficient of variation are statistical indicators used to measure the dispersion degree of data. In the embodiments of the present application, variance is selected as the dispersion indicator, that is, the variance of the number of occurrences of the target string in all hierarchical text sets is used as the hierarchical dispersion degree of the target string; the smaller the hierarchical dispersion degree, the more the target string conforms to entity identity. The hierarchical importance represents the proportion of the character string of the target string in all hierarchical text sets, and is used to measure the importance of the target string in all hierarchical text sets. The greater the hierarchical importance, the more important the entity identity of the target string. Therefore, the hierarchical importance and the hierarchical identity index are in a positive correlation relationship, and the hierarchical dispersion degree and the hierarchical identity index are in a negative correlation relationship. In the embodiments of the present application, the product of the hierarchical importance and the hierarchical dispersion degree of the target string is normalized to obtain the hierarchical identity index of the target string. In the embodiments of the present application, the correlation relationship between the hierarchical importance, the hierarchical dispersion degree, and the hierarchical identity index can also be constructed through other basic mathematical operations, which will not be limited and elaborated herein.

[0098] It should be noted that in the embodiments of the present application, the Norm function is used for normalization processing. In the embodiments of the present application, other normalization methods can also be selected, such as function transformation, Sigmoid function and other normalization methods, which will not be limited herein.

[0099] In some embodiments, the hierarchical identity index is calculated by the following formula:

[0100]

[0101] In the formula, Q is the hierarchical identity index of the target string; is the hierarchical dispersion degree of the target string; A is the total number of hierarchical text sets; is the number of occurrences of the target string in the a-th hierarchical text set; is the total number of character strings in the a-th hierarchical text set; is the quantity proportion of the target string in the a-th hierarchical text set; is the hierarchical importance of the target string; exp is the exponential function with the natural constant e as the base; Norm is the normalization function. It should be noted that when the hierarchical identity index Q is larger, the target string conforms to entity identity more and the entity identity is more important, indicating that the analysis field where the target string is located is more suitable for constructing a knowledge graph, and the integrity of the meaning expressed by the target string in the power communication network security knowledge is higher.

[0102] In some embodiments, determining the hierarchical isolation index of the string according to the discrete index of the number of occurrences of the string in the partitioned text set and the difference degree of the occurrences of the string in all partitioned text sets at the corresponding level includes: calculating the partition discreteness according to the discrete index of the number of occurrences of the string in the partitioned text set; calculating the sum of the difference degrees between the number of occurrences of the string in a certain partitioned text set and the number of occurrences of the string in the remaining partitioned text sets other than the certain partitioned text set in the corresponding hierarchical text set to obtain the occurrence difference degree; calculating the sum of the occurrence difference degrees of all partitioned text sets in the corresponding hierarchical text set to obtain the partition difference degree; and determining the hierarchical isolation index based on the partition discreteness and the partition difference degree.

[0103] There is security isolation protection between terminal systems in different security partitions at the same level of the power communication network security system, so that data in different security partitions cannot be transmitted to each other. The specific reason is that the power communication network security protection systems at the same level are different, the system encryption is different, or different transmission protocols are used, etc., resulting in differences in network security knowledge in different security partitions at the same level. Then, the number of occurrences of the target string in the partitioned text sets corresponding to the same level is in a discrete state.

[0104] The discreteness of the number of occurrences of the target string in all partitioned text sets corresponding to each level focuses on comparing the differences and significance between the above-mentioned number of occurrences; the differences between the number of occurrences of the target string in different partitioned text sets corresponding to each level focus on describing the fluctuation magnitude of the above-mentioned number of occurrences; comprehensively analyzing the discrete state of the number of occurrences of the target string in the partitioned text sets corresponding to the same level from the above two aspects improves the accuracy of the hierarchical isolation index.

[0105] It should be noted that standard deviation, variance, range, and coefficient of variation are statistical indicators used to measure the discreteness of data. In the embodiments of the present application, variance is selected as the discrete index, that is, the variance of the number of occurrences of the target string in all partitioned text sets corresponding to each level is used as the partition discreteness of the target string at each level.

[0106] If the partition dispersion degree is greater, the difference and significance between the number of occurrences of the target string in the partition text sets at each level are greater; if the partition difference degree is greater, the fluctuation of the number of occurrences of the target string in the partition text sets at each level is greater, the partition isolation degree of the target string is greater, and the target string better conforms to entity isolation; then both the partition dispersion degree and the partition difference degree are positively correlated with the hierarchical isolation index. In the embodiments of the present application, the product of the partition dispersion degree and the partition difference degree of the target string at each level is normalized to obtain the hierarchical isolation index of the target string. In the embodiments of the present application, the correlation relationship between the partition dispersion degree, the partition difference degree, and the hierarchical isolation index can also be constructed through other basic mathematical operations, which will not be limited and elaborated herein.

[0107] In some embodiments, the hierarchical isolation index is calculated by the following formula:

[0108]

[0109]

[0110] In the formula, E is the hierarchical isolation index of the target string; A is the total number of hierarchical text sets; is the local isolation index of the target string at the a-th level; is the partition dispersion degree of the target string at the a-th level; is the total number of strings in the hierarchical text set at the a-th level; is the number of occurrences of the target string in the n1-th partition text set at the a-th level; is the number of occurrences of the target string in the n2-th partition text set at the a-th level; is the occurrence difference degree of the target string in the n1-th partition text set at the a-th level; is the partition difference degree of the target string at the a-th level; is the absolute value function; Norm is the normalization function.

[0111] Furthermore, according to the hierarchical identity index and the hierarchical isolation index, the complete index of the target string is obtained.

[0112] In the process of constructing a knowledge graph, two key principles need to be followed: entity identity and entity isolation. The hierarchical identity index reflects the entity identity of the target string, and the hierarchical isolation index reflects the entity isolation of the target string. If both the hierarchical identity index and the hierarchical isolation index are larger, it indicates that the analysis field where the target string is located is more suitable for constructing a knowledge graph, and the integrity of the meaning expressed by the target string in the power communication network security knowledge is higher, then the integrity index of the target string is larger. Therefore, both the hierarchical identity index and the hierarchical isolation index are positively correlated with the integrity index. In the embodiment of the present application, the product of the hierarchical identity index and the hierarchical isolation index of the target string is normalized to obtain the integrity index of the target string.

[0113] In the embodiment of the present application, the correlation relationship between the hierarchical identity index, the hierarchical isolation index, and the integrity index can also be constructed through other basic mathematical operations, which will not be limited and elaborated here.

[0114] It should be noted that in the embodiment of the present application, the Sigmoid function is used for normalization processing. In the embodiment of the present application, other normalization methods can also be selected, such as function transformation and other normalization methods, which will not be limited here.

[0115] For step S103, determine the comprehensive importance of the string according to the semantic vector corresponding to the string.

[0116] In some embodiments, the determining the comprehensive importance of the string according to the semantic vector corresponding to the string includes: determining the local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string at all levels; determining the comprehensive importance based on the local importance.

[0117] In this embodiment, the power communication network security knowledge is multi-source, heterogeneous, and fragmented. At the same time, a large number of useful information fragments such as security knowledge bases and information bases are scattered everywhere in the power communication network, and the credibility of the data changes due to different data sources. For example, for the same relatively complete data, the information data from CNKI has higher authority and better data quality compared to the information data from Baidu web pages.

[0118] In the power communication network security system, analyzing the authority of a string is a relatively complex task. Generally, it is considered that the higher the level of the subject to which the information source belongs, the higher the data authority; if the data appears more times in relevant laws, regulations, and regulatory documents, making the data more reliable, then the data authority is higher.

[0119] Since the target string may have different but same-meaning expressions in different files due to the reasons of the writers, it is easy to have errors when only using the similarity between strings to analyze the authority of keywords. Therefore, this application takes into account the context semantic features of the keywords where the string is located, and considers the authority of the target string by analyzing the semantic similarity degree between the analysis field where the target string is located and the analysis fields where the corresponding strings of the target string in each hierarchical text set are located, so as to ensure that the entities and relationships in the knowledge graph can accurately reflect the complex relationships and structures in the real world. The higher the authority of the target string, the higher its integrity in the power communication network security system, and integrity is an important indicator to measure the quality of the knowledge graph, which is directly related to whether the knowledge graph can comprehensively cover the knowledge in the relevant fields. Therefore, by combining the semantic similarity degree between the analysis field where the target string is located and the analysis fields where the corresponding strings of the target string in each hierarchical text set are located, and the integrity index for analysis, the accuracy of the quality index of the target string is improved.

[0120] It should be noted that in the embodiments of this application, the analysis field where the string is located is vectorized by using the bag-of-words model (BoW) to obtain the semantic vector of the analysis field where each string is located; the word embedding algorithm, bidirectional encoder representations from transformers (BERT), etc. can also be selected. Since the calculation of the semantic vector involves the semantic understanding of the text, the semantic vectors of the analysis field where the target string is located and the corresponding analysis fields in the field set at each level are different due to context, grammar, or intonation factors.

[0121] The method for obtaining the comprehensive importance is as follows: The sum of the similarity indexes between the semantic vector of the analysis field where the target string is located and the target vectors at each level of the target string is used as the comprehensive similarity of the target string at each level; the product of the number of target vectors at each level of the target string and the comprehensive similarity is used as the local importance of the target string at each level; the sum of the local importances of the target string at all levels is recorded as the comprehensive importance of the target string.

[0122] Further, for step S104, the quality index of the string is determined based on the integrity index and the comprehensive importance.

[0123] In some embodiments, the quality index is calculated by the following formula:

[0124]

[0125] In the formula, T is the quality index of the target string; Z is the integrity index of the target string; A is the total number of hierarchical text sets, which is equal to the total number of levels of the power communication network security system; is the total number of target vectors of the target string at the a-th level; is the semantic vector of the analysis field where the target string is located; is the b-th target vector of the target string at the a-th level; is the comprehensive similarity of the target string at the a-th level; is the local importance of the target string at the a-th level; is the comprehensive importance of the target string; cos is the cosine function; Sigmoid is the normalization function.

[0126] It should be noted that represents the level of the subject to which the target string belongs. When is larger, the target string has stronger relevance and importance in the data, so the level of the target string in the knowledge graph will be correspondingly higher, the authority of the analysis field where the target string is located is higher, and thus the quality of the analysis field where the target string is located is higher. When is larger, the similarity between the meaning of the analysis field where the target string is located and the semantics of the analysis field at the a-th level is higher, indicating that the analysis field where the target string is located can ensure that the entities and relationships in the knowledge graph can accurately reflect the complex relationships and structures in the real world, the authority of the analysis field where the target string is located is higher, and the quality of the analysis field where the target string is located is higher. When the complete index is larger, it means that the reliability of the target string is higher, that is, the energy level of the target string is higher, the analysis field where the target string is located should be in a more core position in the process of constructing the knowledge graph of power communication network security, and the quality index of the target string is larger.

[0127] The method for obtaining the quality index of each string in the set of text at all levels is the same as the method for obtaining the quality index of the target string.

[0128] Further, for step S105, a knowledge graph is constructed based on the quality index of the string.

[0129] In some embodiments, the constructing a knowledge graph based on the quality index of the string includes: using the corresponding strings with quality indexes higher than a preset threshold to construct the knowledge graph.

[0130] In this embodiment, for the quality index of each string in the analysis field of the overall text set, the analysis field where the string whose quality index is greater than the preset quality threshold is located is recorded as the target field; the target field is high-quality data, and the target fields with larger quality indexes should be in the core position of the knowledge graph of power communication network security. A knowledge graph of the power communication network is constructed using the high-quality target fields, making the knowledge graph more accurate and reliable, which is beneficial to improving the threat intelligence analysis ability, promoting the sharing and inheritance of security knowledge, improving the emergency response efficiency, helping maintenance personnel complete tasks such as decision-making guidance and instruction verification, and reducing the workload and operation risk.

[0131] It should be noted that in the embodiment of the present application, the preset quality threshold takes an empirical value of 0.7, and the implementer can set it according to specific circumstances. In the embodiment of the present application, the Neo4j graph database is selected to construct the knowledge graph, and community discovery algorithms, shortest path algorithms, node embedding algorithms, etc. can also be selected to construct the knowledge graph.

[0132] As can be seen from the above embodiments, the knowledge graph construction method described in the embodiments of the present application divides each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; according to the number of times the string appears in the text set of each level and the number of times the string appears in the text set of the partition, the complete index of the string is determined; the text set of the level is a set composed of strings of all the security partitions within each level; the text set of the partition is a set composed of strings of each security partition; the comprehensive importance of the string is determined according to the semantic vector corresponding to the string; based on the complete index and the comprehensive importance, the quality index of the string is determined; based on the quality index of the string, a knowledge graph is constructed. In the embodiment of the present application, each obtained analysis field is split into strings and counted in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string in the text sets of each level, but also considers its performance in specific security partitions. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality evaluation. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By performing a normalization calculation on the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that strings with quality indexes higher than the set threshold are preferentially selected during the knowledge graph construction process. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data foundation for the management and application of knowledge, ultimately promoting the in-depth mining of information and the improvement of application value.

[0133] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.

[0134] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0135] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a knowledge graph construction device.

[0136] Referring to Figure 2 , the knowledge graph construction device includes:

[0137] A division module 21, configured to divide each acquired analysis field into strings; the analysis fields are located in different levels and different security partitions;

[0138] A complete index determination module 22, configured to determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition;

[0139] A comprehensive importance determination module 23, configured to determine the comprehensive importance of the string according to the semantic vector corresponding to the string;

[0140] A quality index determination module 24, configured to determine the quality index of the string based on the complete index and the comprehensive importance;

[0141] A construction module 25, configured to construct a knowledge graph based on the quality index of the string.

[0142] In a possible implementation manner, the complete index determination module 22 is further configured to:

[0143] The same-level index determination unit is configured to determine the same-level index of the string according to the discrete index of the number of times the string appears in each level text set and the number of times the string appears in all level text sets;

[0144] The level isolation index determination unit is configured to determine the level isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference degree of the string in all partition text sets corresponding to the level;

[0145] The complete index determination unit is configured to determine the complete index of the string based on the same-level index and the level isolation index.

[0146] In a possible implementation manner, the same-level index determination unit is further configured to:

[0147] Determine the quantity proportion according to the ratio of the number of times the string appears in each level text set to the total number of strings in the level text set;

[0148] Calculate the cumulative sum of the quantity proportions of the string in all level text sets to obtain the level importance;

[0149] Calculate the level dispersion according to the discrete index of the number of times the string appears in all level text sets;

[0150] Determine the same-level index according to the level importance and the level dispersion.

[0151] In a possible implementation manner, the level isolation index determination unit is further configured to:

[0152] Calculate the partition dispersion according to the discrete index of the number of times the string appears in the partition text set;

[0153] Calculate the cumulative sum of the difference degrees between the number of times the string appears in a certain partition text set and the number of times the string appears in the remaining partition text sets other than the certain partition text set in the corresponding level text set to obtain the appearance difference degree;

[0154] Calculate the cumulative sum of the appearance difference degrees of all partition text sets in the corresponding level text set to obtain the partition difference degree;

[0155] Determine the level isolation index based on the partition dispersion and the partition difference degree.

[0156] In a possible implementation, the comprehensive importance determination module 23 is further configured to:

[0157] Determine the local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string at all levels;

[0158] Determine the comprehensive importance based on the local importance.

[0159] In a possible implementation, the quality index determination module 24 is further configured to:

[0160] Perform a normalization calculation on the complete index and the comprehensive importance to obtain the quality index.

[0161] In a possible implementation, the construction module 25 is further configured to:

[0162] Construct the knowledge graph by using the strings corresponding to the quality indexes higher than the preset threshold.

[0163] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0164] The device in the above embodiment is used to implement the corresponding knowledge graph construction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be described in detail here.

[0165] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the knowledge graph construction method described in any of the above embodiments.

[0166] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0167] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0168] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0169] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0170] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can also achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0171] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0172] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification and does not necessarily include all the components shown in the figure.

[0173] The electronic device of the above embodiment is used to implement the corresponding knowledge graph construction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0174] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the knowledge graph construction method as described in any of the above embodiments.

[0175] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0176] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the knowledge graph construction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0177] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

[0178] In addition, for simplicity of explanation and discussion, and so as not to make the embodiments of the present application difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions should be considered illustrative rather than restrictive.

[0179] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0180] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.

Claims

1. A method for constructing a knowledge graph, characterized in that Including: Dividing each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; The different levels are the boundary layer, core layer, and access layer of the power communication network security system; the different security partitions include dividing each of the levels into an intranet security area and an extranet security area; Determining the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set, including: determining the level identity index of the string according to the discrete index of the number of times the string appears in each level text set and the number of times the string appears in all level text sets; determining the level isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference degree of the string in all partition text sets corresponding to the level; determining the complete index of the string based on the level identity index and the level isolation index; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; Determining the comprehensive importance of the string according to the semantic vector corresponding to the string, including: determining the local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string in all levels; determining the comprehensive importance based on the local importance; Determining the quality index of the string based on the complete index and the comprehensive importance; Constructing a knowledge graph based on the quality index of the string.

2. The method according to claim 1, wherein The determining the level identity index of the string according to the discrete index of the number of times the string appears in each level text set and the number of times the string appears in all level text sets includes: Determining the quantity proportion according to the ratio of the number of times the string appears in each level text set to the total number of strings in the level text set; Calculating the cumulative sum of the quantity proportions of the string in all level text sets to obtain the level importance; Calculating the level dispersion according to the discrete index of the number of times the string appears in all level text sets; Determining the level identity index according to the level importance and the level dispersion.

3. The method according to claim 1, wherein The determining the level isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference degree of the string in all partition text sets corresponding to the level includes: Calculating the partition dispersion according to the discrete index of the number of times the string appears in the partition text set; Calculating the cumulative sum of the difference degrees between the number of times the string appears in a certain partition text set and the number of times the string appears in the remaining partition text sets except the certain partition text set in the corresponding level text set to obtain the appearance difference degree; Calculate the cumulative sum of the occurrence difference degrees of the said string in all partition text sets of the corresponding hierarchical text set to obtain the partition difference degree. Based on the said partition dispersion and the said partition difference degree, determine the said hierarchical isolation index.

4. The method according to claim 1, wherein The determining of the quality index of the said string based on the said complete index and the said comprehensive importance includes: Perform a normalization calculation on the said complete index and the said comprehensive importance to obtain the said quality index.

5. The method according to claim 1, wherein The constructing of the knowledge graph based on the quality index of the said string includes: Use the said strings corresponding to the quality index higher than the preset threshold to construct the said knowledge graph.

6. A knowledge graph construction device, characterized in that Includes: A partitioning module, configured to partition each obtained analysis field into strings; the analysis fields are located in different hierarchies and different security partitions. The said different hierarchies are the boundary layer, core layer, and access layer of the power communication network security system; the said different security partitions include partitioning each of the said hierarchies into an intranet security area and an extranet security area. A complete index determination module, configured to determine the complete index of the said string according to the number of times the said string appears in each hierarchical text set and the number of times the said string appears in the partition text set; the hierarchical text set is a set composed of strings of all the said security partitions within each hierarchy; the partition text set is a set composed of strings of each of the said security partitions. A comprehensive importance determination module, configured to determine the comprehensive importance of the said string according to the semantic vector corresponding to the said string, including: determining the local importance according to the similarity between the semantic vector corresponding to the said string and the target vectors of the said string in all hierarchies; determining the comprehensive importance based on the said local importance. A quality index determination module, configured to determine the quality index of the said string based on the said complete index and the said comprehensive importance. A construction module, configured to construct a knowledge graph based on the quality index of the said string. The said complete index determination module is further configured to: A hierarchical same index determination unit, configured to determine the hierarchical same index of the said string according to the discrete index of the number of times the said string appears in each hierarchical text set and the number of times the said string appears in all hierarchical text sets. A hierarchical isolation index determination unit, configured to determine the hierarchical isolation index of the said string according to the discrete index of the number of times the said string appears in the partition text set and the occurrence difference degree of the said string in all partition text sets of the corresponding said hierarchy. A complete index determination unit, configured to determine the complete index of the said string based on the said hierarchical same index and the said hierarchical isolation index.

7. The device according to claim 6, wherein The said hierarchical same index determination unit is further configured to: Determine the quantity proportion according to the ratio of the number of times the said string appears in each of the said hierarchical text sets and the total number of strings in the hierarchical text set. Calculate the cumulative sum of the said quantity proportions of the said string in all of the said hierarchical text sets to obtain the hierarchical importance. Calculate the hierarchical dispersion according to the discrete index of the number of times the said string appears in all hierarchical text sets. Determine the same index of the level according to the level importance and the level dispersion degree.

8. The device according to claim 6, wherein The level isolation index determination unit is further configured to: Calculate the partition dispersion degree according to the dispersion index of the number of occurrences of the string in the partition text set; Calculate the sum of the differences between the number of occurrences of the string in a certain partition text set and the number of occurrences of the string in the remaining partition text sets other than the certain partition text set in the corresponding level text set, to obtain the occurrence difference degree; Calculate the sum of the occurrence difference degrees of all the partition text sets in the corresponding level text set of the string, to obtain the partition difference degree; Determine the level isolation index based on the partition dispersion degree and the partition difference degree.

9. The device according to claim 6, wherein The quality index determination module is further configured to: Perform a normalization calculation on the complete index and the comprehensive importance to obtain the quality index.

10. The device according to claim 6, characterized in that, The construction module is further configured to: Construct the knowledge graph by using the corresponding strings whose quality index is higher than a preset threshold.

11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Knowledge network construction method and device, equipment and storage medium

    CN114706991A

  • Text theme segmentation method and device based on knowledge graph and electronic equipment

    CN116340525A