Knowledge graph construction method and device, electronic equipment and readable storage medium
By dividing the analysis fields into strings in the construction of power communication network security knowledge graphs and counting them at different levels and security partitions, the complete indicators and comprehensive importance of strings are determined, and the problem of low data quality in the existing technology is solved and high-quality knowledge graph construction is achieved.
Patent Information
- Application Number
- CN202510632705.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When building a power communication network security knowledge graph, the existing methods are not high in data model design quality problems and unclear business needs, which leads to poor quality of the knowledge graph.
By dividing each acquired analysis field into strings and counting it in different levels and security partitions, the complete indicators and comprehensive importance of the string are determined, and a knowledge graph is constructed based on these indicators.
This method not only considers the frequency of occurrence of strings in text sets at various levels and their performance in specific security partitions, but also incorporates deep semantic understanding, improving the accuracy and reliability of the knowledge graph.
Smart Images

Figure CN120146174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graphs, and in particular, to a method, apparatus, electronic device, and readable storage medium for constructing a knowledge graph. Background Art
[0002] In the wave of digital transformation, the security of the power communication network is directly related to the development of grid intelligence and the stable operation of the power system. With the continuous evolution and complexity of network attack means, traditional network security defense measures have been difficult to cope with the increasingly severe network security challenges. In order to improve the network security protection ability of the power system and the ability to respond to network security incidents, constructing a power communication network security knowledge graph provides artificial intelligence services for power communication network management, risk monitoring, etc.
[0003] The quality of graph data is a key factor in evaluating the success of constructing a power communication network security knowledge graph. Existing methods screen data manually and then construct a high-quality database, and construct a knowledge graph by obtaining power communication network security data from the high-quality database. However, due to problems such as the quality of data model design and unclear business requirements, the data quality in the high-quality database is not high, which in turn leads to poor quality of the power communication network security knowledge graph. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a method, apparatus, electronic device, and readable storage medium for constructing a knowledge graph.
[0005] Based on the above purpose, this application provides a method for constructing a knowledge graph, including: Dividing each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; Determining the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; Determining the comprehensive importance of the string according to the semantic vector corresponding to the string; Determining the quality index of the string based on the complete index and the comprehensive importance; Constructing a knowledge graph based on the quality index of the string.
[0006] In a possible implementation, the determining the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set includes: Determine the hierarchical identity index of the string according to the discrete index of the number of occurrences of the string in each hierarchical text set and the number of occurrences of the string in all hierarchical text sets; Determine the hierarchical isolation index of the string according to the discrete index of the number of occurrences of the string in the partition text set and the occurrence difference degree of the string in all partition text sets at the corresponding level; Determine the complete index of the string based on the hierarchical identity index and the hierarchical isolation index.
[0007] In a possible implementation manner, the determining the hierarchical identity index of the string according to the discrete index of the number of occurrences of the string in each hierarchical text set and the number of occurrences of the string in all hierarchical text sets includes: Determine the quantity proportion according to the ratio of the number of occurrences of the string in each hierarchical text set to the total number of strings in the hierarchical text set; Calculate the cumulative sum of the quantity proportions of the string in all hierarchical text sets to obtain the hierarchical importance; Calculate the hierarchical dispersion according to the discrete index of the number of occurrences of the string in all hierarchical text sets; Determine the hierarchical identity index according to the hierarchical importance and the hierarchical dispersion.
[0008] In a possible implementation manner, the determining the hierarchical isolation index of the string according to the discrete index of the number of occurrences of the string in the partition text set and the occurrence difference degree of the string in all partition text sets at the corresponding level includes: Calculate the partition dispersion according to the discrete index of the number of occurrences of the string in the partition text set; Calculate the cumulative sum of the difference degrees between the number of occurrences of the string in a certain partition text set and the number of occurrences of the string in the remaining partition text sets other than the certain partition text set in the corresponding hierarchical text set to obtain the occurrence difference degree; Calculate the cumulative sum of the occurrence difference degrees of the string in all partition text sets in the corresponding hierarchical text set to obtain the partition difference degree; Determine the hierarchical isolation index based on the partition dispersion and the partition difference degree.
[0009] In a possible implementation manner, the determining the comprehensive importance of the string according to the semantic vector corresponding to the string includes: Determine the local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string at all levels. Determine the comprehensive importance based on the local importance.
[0010] In a possible implementation manner, the determining the quality index of the string based on the complete index and the comprehensive importance includes: Perform a normalization calculation on the complete index and the comprehensive importance to obtain the quality index.
[0011] In a possible implementation manner, the constructing a knowledge graph based on the quality index of the string includes: Use the strings corresponding to the quality index higher than the preset threshold to construct the knowledge graph.
[0012] Based on the same inventive concept, an embodiment of the present application further provides a knowledge graph construction device, including: A partitioning module, configured to partition each acquired analysis field into strings; the analysis fields are located in different levels and different security partitions; A complete index determination module, configured to determine the complete index of the string according to the number of times the string appears in the text set of each level and the number of times the string appears in the text set of the partition; the text set of each level is a set composed of strings of all the security partitions within each level; the text set of the partition is a set composed of strings of each security partition; A comprehensive importance determination module, configured to determine the comprehensive importance of the string according to the semantic vector corresponding to the string; A quality index determination module, configured to determine the quality index of the string based on the complete index and the comprehensive importance; A construction module, configured to construct a knowledge graph based on the quality index of the string.
[0013] In a possible implementation manner, the complete index determination module is further configured to: A same-level index determination unit, configured to determine the same-level index of the string according to the discrete index of the number of times the string appears in the text set of each level and the number of times the string appears in the text sets of all levels; A level isolation index determination unit, configured to determine the level isolation index of the string according to the discrete index of the number of times the string appears in the text set of the partition and the appearance difference degree of the string in all the text sets of the corresponding level; A complete metric determination unit, configured to determine a complete metric of the string based on the same metric at the level and the isolation metric at the level.
[0014] In a possible implementation, the same metric determination unit at the level is further configured to: Determine a quantity proportion according to a ratio of the number of times the string appears in each of the text sets at the level to the total number of strings in the text set at the level; Calculate a cumulative sum of the quantity proportions of the string in all of the text sets at the level to obtain a level importance; Calculate a level dispersion based on a discrete metric of the number of times the string appears in all text sets at all levels; Determine the same metric at the level according to the level importance and the level dispersion.
[0015] In a possible implementation, the isolation metric determination unit at the level is further configured to: Calculate a partition dispersion based on a discrete metric of the number of times the string appears in the partition text set; Calculate a cumulative sum of the difference degrees between the number of times the string appears in a certain partition text set and the number of times the string appears in the remaining partition text sets other than the certain partition text set in the corresponding text set at the level to obtain an appearance difference degree; Calculate a cumulative sum of the appearance difference degrees of all partition text sets in the corresponding text set at the level to obtain a partition difference degree; Determine the isolation metric at the level based on the partition dispersion and the partition difference degree.
[0016] In a possible implementation, the comprehensive importance determination module is further configured to: Determine a local importance according to the similarity between the semantic vector corresponding to the string and the target vectors at all levels of the string; Determine the comprehensive importance based on the local importance.
[0017] In a possible implementation, the quality metric determination module is further configured to: Perform a normalization calculation on the complete metric and the comprehensive importance to obtain the quality metric.
[0018] In a possible implementation, the construction module is further configured to: Construct the knowledge graph by using the corresponding strings whose quality metrics are higher than a preset threshold.
[0019] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the knowledge graph construction method described in any one of the above is implemented.
[0020] Based on the same inventive concept, an embodiment of the present application further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the knowledge graph construction method described in any one of the above.
[0021] As can be seen from the above, the knowledge graph construction method, device, electronic device, and readable storage medium provided by the present application divide each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set, a complete index of the string is determined; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; the comprehensive importance of the string is determined according to the semantic vector corresponding to the string; based on the complete index and the comprehensive importance, the quality index of the string is determined; based on the quality index of the string, a knowledge graph is constructed. In the embodiment of the present application, each obtained analysis field is split into strings and statistically analyzed in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string in each level text set, but also considers its performance in a specific security partition. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality evaluation. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By performing normalization calculation on the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that during the knowledge graph construction process, strings with quality indexes higher than the set threshold are preferentially selected. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data basis for knowledge management and application, ultimately promoting the in-depth mining of information and the improvement of application value. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 Schematic flowchart of the knowledge graph construction method according to an embodiment of the present application; Figure 2 Schematic structural diagram of the knowledge graph construction device according to an embodiment of the present application; Figure 3 Schematic structural diagram of the electronic device according to an embodiment of the present application. Detailed implementation manners
[0024] To make the objectives, technical solutions and advantages of the present application more clearly understood, the following further details the present application in conjunction with specific embodiments and with reference to the accompanying drawings.
[0025] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be the ordinary meanings understood by those of ordinary skill in the art to which the present application pertains. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0026] It can be understood that before using the technical solutions of the various embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0027] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to the software or hardware such as an electronic device, application program, server or storage medium that performs the operations of the technical solutions of the present application according to the prompt message.
[0028] As an optional but non-limiting implementation manner, in response to receiving an active request from a user, the manner of sending a prompt message to the user may be, for example, a pop-up window manner, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0029] It can be understood that the above notification and user authorization acquisition process is only illustrative and does not limit the implementation manner of this application. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of this application.
[0030] As described in the background art section, in the wave of digital transformation, the security of the power communication network is directly related to the development of grid intelligence and the stable operation of the power system. With the continuous evolution and complexity of network attack means, traditional network security defense measures have been difficult to cope with the increasingly severe network security challenges. In order to improve the network security protection ability of the power system and the ability to respond to network security incidents, a power communication network security knowledge graph is constructed to provide artificial intelligence services for power communication network management, risk monitoring, etc.
[0031] The quality of the graph data is a key factor in evaluating the success of constructing the power communication network security knowledge graph. Existing methods screen data manually and then construct a high-quality database, and construct a knowledge graph by obtaining power communication network security data from the high-quality database. However, due to reasons such as the quality problem of data model design and unclear business requirements, the data quality in the high-quality database is not high, and thus the quality of the power communication network security knowledge graph is not good.
[0032] Considering the above, an embodiment of the present application proposes a method for constructing a knowledge graph. Each obtained analysis field is divided into strings. The analysis fields are located in different levels and different security partitions. According to the number of times the string appears in each level text set and the number of times the string appears in the partition text set, the complete index of the string is determined. The level text set is a set composed of strings of all the security partitions within each level. The partition text set is a set composed of strings of each security partition. According to the semantic vector corresponding to the string, the comprehensive importance of the string is determined. Based on the complete index and the comprehensive importance, the quality index of the string is determined. Based on the quality index of the string, a knowledge graph is constructed. In the embodiment of the present application, each obtained analysis field is split into strings and counted in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string in each level text set, but also its performance in a specific security partition. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality assessment. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By normalizing the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that in the process of constructing the knowledge graph, strings with quality indexes higher than the set threshold are preferentially selected. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data basis for knowledge management and application, ultimately promoting the in-depth mining of information and the improvement of application value.
[0033] Hereinafter, the technical solutions of the embodiments of the present application will be described in detail through specific embodiments.
[0034] Referring to Figure 1 , the method for constructing a knowledge graph according to the embodiment of the present application includes the following steps: Step S101, divide each obtained analysis field into strings. The analysis fields are located in different levels and different security partitions. Step S102, determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set. The level text set is a set composed of strings of all the security partitions within each level. The partition text set is a set composed of strings of each security partition. Step S103, determine the comprehensive importance of the string according to the semantic vector corresponding to the string. Step S104, determine the quality index of the string based on the complete index and the comprehensive importance degree. Step S105, construct a knowledge graph based on the quality index of the string.
[0035] Regarding step S101, first, it is necessary to obtain the analysis fields of different levels and different security partitions, and then divide the obtained analysis fields into strings.
[0036] In this embodiment, obtain the analysis fields of different security partitions within each level of the power communication network security system, and denote the set composed of the analysis fields within all levels as the overall text set; divide each analysis field into different strings; denote the set composed of the strings corresponding to the analysis fields of all security partitions within each level as the level text set, and denote the set composed of the strings corresponding to the analysis fields of each security partition as the partition text set; randomly select a string from the strings corresponding to the analysis fields in the overall text set for subsequent description, and denote this string as the target string in the following embodiments.
[0037] Specifically, in the embodiment of the present application, the power communication network security system is divided into three levels: the boundary layer, the core layer, and the access layer, and at the same time, each level is divided into an internal network security area and an external network security area.
[0038] The internal network security area of the boundary layer includes: all devices, systems, and resources within the power communication network, such as power plants, transmission systems, communication devices, etc., and text files such as core system logs, device configuration files, traffic analysis reports, etc. can be obtained; the external network security area includes: security devices such as external entry points, firewalls, intrusion detection systems, anti-virus gateways, etc., and text files such as connection request logs from the outside and anti-virus scan logs can be obtained.
[0039] The internal network security area of the core layer includes: core routers, switches, and important servers and storage devices connected to these devices, and more sensitive and important data such as power production data, real-time sensor data, transmission system status information, etc. can be obtained; the external network security area includes: advanced firewalls, intrusion prevention systems, traffic analysis tools, etc., and performance statistics of core routers and switches, traffic analysis data, operation logs of core devices, etc. can be obtained.
[0040] The internal network security area of the access layer includes various terminal devices, such as end-user computers, smart meters, etc., which can obtain information such as security audit logs, access control lists, and terminal device configurations of end-user devices; the external network security area involves the security control of user access points, such as network access control systems, authentication and authorization systems, and security gateway devices, which may obtain access request records from user terminals, user authentication information, audit logs of network access control systems, etc.
[0041] Obtain the text files of each security partition within each layer of the power communication network security system, and use the Jieba algorithm to perform word segmentation on the text files of each security partition by combining dictionary matching and dynamic programming to obtain the keywords of each security partition. Select any one security partition as the partition to be tested, select any one keyword within the partition to be tested as the target word, and use the brute-force matching algorithm to match the target word with the keywords of the remaining security partitions except the partition to be tested. If there is a keyword in the keywords of the remaining security partitions that matches the target word, then mark the target word as the analysis field of the partition to be tested, and traverse the keywords of the partition to be tested to obtain the analysis field of the partition to be tested. According to the above method, obtain the analysis fields of each security partition within each layer of the power communication network security system.
[0042] Denote the set composed of the analysis fields within all layers as the overall text set; to accurately judge the importance of the analysis fields, randomly divide each analysis field into strings of different lengths. For the convenience of subsequent description, denote the set composed of the strings corresponding to the analysis fields of all security partitions within each layer as the layer text set, and denote the set composed of the strings corresponding to the analysis fields of each security partition as the partition text set; select any one string from the strings corresponding to the analysis fields in the overall text set as the target string.
[0043] Further, for step S102, determine the complete index of the string according to the number of times the string appears in each layer text set and the number of times the string appears in the partition text set; the layer text set is the set composed of the strings of all the security partitions within each layer; the partition text set is the set composed of the strings of each security partition.
[0044] In some embodiments, determining the complete index of the string according to the number of times the string appears in each hierarchical text set and the number of times the string appears in the partitioned text set includes: determining the hierarchical identity index of the string according to the discrete index of the number of times the string appears in each hierarchical text set and the number of times the string appears in all hierarchical text sets; determining the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partitioned text set and the occurrence difference degree of the string in all partitioned text sets corresponding to the hierarchy; determining the complete index of the string based on the hierarchical identity index and the hierarchical isolation index.
[0045] In this embodiment, the core layer of the power domain knowledge graph application architecture bears the capabilities of natural language processing, knowledge extraction, knowledge fusion, and knowledge processing. The power communication network security knowledge data of each hierarchical architecture usually uses a relational database to manage files, and the corresponding relationship of the data is a one-to-many relationship. In the network security knowledge scheduling of power communication, its security pre-plan behavior is carried out according to the requirements of relevant security management regulations, security response manuals and other documents; its security knowledge builds a basic architecture system from top to bottom. The power communication network security knowledge data is relatively complete within the architecture system. The more complete the data is used, the more perfect the connections and nodes of the knowledge graph construction will be; there are many data source channels, and different levels have different requirements for security data. The higher the level of the data, the higher the authority and the higher the data quality, and the data quality is judged accordingly.
[0046] Each level of the terminal and system of the power communication network security system is only allowed to be used within the corresponding level. It is dedicated to the private network, and the equipment designed to connect to the lower-level network cannot intervene in the upper-level network at will. The network operations in the corresponding network knowledge security manual have corresponding and strict specification requirements. Moreover, in the power communication network security system, there is a horizontal security isolation protection between different security partitions, so that the terminal systems in different security partitions cannot directly access each other. Isolation is for security protection and to strengthen the transmission control of data; because the network data between different security partitions at the same level are not connected to each other, the network security-related knowledge requirements for different security partitions are also different, and the integrity of the target string is calculated accordingly.
[0047] In the process of constructing a knowledge graph, two key principles of entity identity and entity isolation need to be followed. Since the security requirements and protection measures of the power communication network security system are common among different levels, there is identity between the strings in the keyword fields of different levels, that is, the dispersion degree of the number of occurrences of the target string in the text sets of all levels is low. Since there is security isolation protection between the terminal systems of different security zones at the same level of the power communication network security system, resulting in differences in the data of different security zones at the same level, there are differences in the number of occurrences of the target string in the text sets of different corresponding zones at the same level. If the target string better conforms to entity identity and entity isolation, it indicates that the analysis field where the target string is located is more suitable for constructing a knowledge graph, and the meaning expressed by the target string has higher integrity in the power communication network security knowledge; analyze the dispersion degree of the number of occurrences of the target string in the text sets of all levels and the differences between the number of occurrences of the target string in the text sets of different corresponding zones at each level to improve the accuracy of the integrity index of the target string.
[0048] In some embodiments, determining the same-level index of the string according to the number of occurrences of the string in the text set of each level and the dispersion index of the number of occurrences of the string in the text sets of all levels includes: determining the quantity proportion according to the ratio of the number of occurrences of the string in the text set of each level to the total number of strings in the text set of that level; calculating the cumulative sum of the quantity proportions of the string in all the text sets of all levels to obtain the importance degree of the level; calculating the dispersion degree of the level according to the dispersion index of the number of occurrences of the string in the text sets of all levels; determining the same-level index according to the importance degree of the level and the dispersion degree of the level.
[0049] Since the security requirements and protection measures of the power communication network security system are common among different levels, there is identity between the strings in the keyword fields of different levels, that is, the number of identical strings in the text sets of different levels is relatively close. The identity of the string reflects the importance degree of the data or information in the system or network. The higher the identity of the string, the more significant the core value and role of the information in a specific context, and it should be in the core position of the knowledge graph of the power communication network security. If the dispersion degree of the number of occurrences of the target string in the text sets of all levels is smaller and the number of occurrences of the target string in the text sets of all levels is larger, it indicates that the target string better conforms to entity identity and entity identity is more important, so as to obtain the same-level index of the target string.
[0050] It should be noted that standard deviation, variance, range, and coefficient of variation are statistical indicators used to measure the dispersion degree of data. In the embodiments of the present application, variance is selected as the dispersion index, that is, the variance of the number of occurrences of the target string in all hierarchical text sets is used as the hierarchical dispersion degree of the target string; the smaller the hierarchical dispersion degree, the more the target string conforms to entity identity. The hierarchical importance presents the proportion of the character string of the target string in all hierarchical text sets, and is used to measure the importance of the target string in all hierarchical text sets. The greater the hierarchical importance, the more important the entity identity of the target string. Therefore, the hierarchical importance and the hierarchical identity index have a positive correlation, and the hierarchical dispersion degree and the hierarchical identity index have a negative correlation. In the embodiments of the present application, the product of the hierarchical importance and the hierarchical dispersion degree of the target string is normalized to obtain the hierarchical identity index of the target string. In the embodiments of the present application, the correlation relationship between the hierarchical importance, the hierarchical dispersion degree, and the hierarchical identity index can also be constructed through other basic mathematical operations, which will not be limited and elaborated here.
[0051] It should be noted that in the embodiments of the present application, the Norm function is used for normalization processing. In the embodiments of the present application, other normalization methods can also be selected, such as function transformation, Sigmoid function and other normalization methods, which will not be limited here.
[0052] In some embodiments, the hierarchical identity index is calculated by the following formula:
[0053] In the formula, Q is the hierarchical identity index of the target string; is the hierarchical dispersion degree of the target string; A is the total number of hierarchical text sets; is the number of occurrences of the target string in the a-th hierarchical text set; is the total number of character strings in the a-th hierarchical text set; is the proportion of the number of the target string in the a-th hierarchical text set; is the hierarchical importance of the target string; exp is the exponential function with the natural constant e as the base; Norm is the normalization function. It should be noted that when the hierarchical identity index Q is larger, the target string conforms to the entity identity more and the entity identity is more important, indicating that the analysis field where the target string is located is more suitable for constructing a knowledge graph, and the integrity of the meaning expressed by the target string in the power communication network security knowledge is higher.
[0054] In some embodiments, determining the hierarchical isolation index of the string according to the discrete index of the number of occurrences of the string in the partitioned text set and the difference degree of the occurrences of the string in all partitioned text sets at the corresponding level includes: calculating the partition dispersion according to the discrete index of the number of occurrences of the string in the partitioned text set; calculating the sum of the difference degrees between the number of occurrences of the string in a certain partitioned text set and the number of occurrences of the string in the remaining partitioned text sets other than the certain partitioned text set in the corresponding hierarchical text set to obtain the occurrence difference degree; calculating the sum of the occurrence difference degrees of all partitioned text sets in the corresponding hierarchical text set to obtain the partition difference degree; and determining the hierarchical isolation index based on the partition dispersion and the partition difference degree.
[0055] There is security isolation protection between terminal systems in different security partitions at the same level of the power communication network security system, so that data in different security partitions cannot be transmitted to each other. The specific reason is that the power communication network security protection systems at the same level are different, the system encryption is different, or different transmission protocols are used, etc., resulting in differences in network security knowledge in different security partitions at the same level. Then the number of occurrences of the target string in the partitioned text sets corresponding to the same level is in a discrete state.
[0056] The discreteness of the number of occurrences of the target string in all partitioned text sets corresponding to each level focuses on comparing the differences and significance between the above-mentioned number of occurrences; the difference between the number of occurrences of the target string in different partitioned text sets corresponding to each level focuses on describing the fluctuation magnitude of the above-mentioned number of occurrences; comprehensively analyzing the discrete state of the number of occurrences of the target string in the partitioned text sets corresponding to the same level from the above two aspects improves the accuracy of the hierarchical isolation index.
[0057] It should be noted that standard deviation, variance, range, and coefficient of variation are statistical indicators used to measure the discreteness of data. In the embodiments of the present application, variance is selected as the discrete index, that is, the variance of the number of occurrences of the target string in all partitioned text sets corresponding to each level is used as the partition dispersion of the target string at each level.
[0058] If the partition dispersion degree is greater, the difference and significance between the number of times the target string appears in the partition text sets at each level are greater; if the partition difference degree is greater, the fluctuation of the number of times the target string appears in the partition text sets at each level is greater, the partition isolation degree of the target string is greater, and the target string better conforms to entity isolation; then both the partition dispersion degree and the partition difference degree are positively correlated with the hierarchical isolation index. In the embodiments of the present application, the product of the partition dispersion degree and the partition difference degree of the target string at each level is normalized to obtain the hierarchical isolation index of the target string. In the embodiments of the present application, the correlation between the partition dispersion degree, the partition difference degree, and the hierarchical isolation index can also be constructed through other basic mathematical operations, which will not be limited and elaborated here.
[0059] In some embodiments, the hierarchical isolation index is calculated by the following formula:
[0060]
[0061] In the formula, E is the hierarchical isolation index of the target string; A is the total number of hierarchical text sets; is the local isolation index of the target string at the a-th level; is the partition dispersion degree of the target string at the a-th level; is the total number of strings in the hierarchical text set at the a-th level; is the number of times the target string appears in the n1-th partition text set at the a-th level; is the number of times the target string appears in the n2-th partition text set at the a-th level; is the appearance difference degree of the target string in the n1-th partition text set at the a-th level; is the partition difference degree of the target string at the a-th level; is the absolute value function; Norm is the normalization function.
[0062] Furthermore, according to the hierarchical identity index and the hierarchical isolation index, the complete index of the target string is obtained.
[0063] In the process of constructing the knowledge graph, two key principles need to be followed: entity identity and entity isolation. The hierarchical identity index reflects the entity identity of the target string, and the hierarchical isolation index reflects the entity isolation of the target string. If both the hierarchical identity index and the hierarchical isolation index are greater, it indicates that the analysis field where the target string is located is more suitable for constructing the knowledge graph, and the integrity of the meaning expressed by the target string in the power communication network security knowledge is higher, then the complete index of the target string is greater. Therefore, both the hierarchical identity index and the hierarchical isolation index are positively correlated with the complete index. In the embodiments of the present application, the product of the hierarchical identity index and the hierarchical isolation index of the target string is normalized to obtain the complete index of the target string.
[0064] In the embodiments of the present application, the correlation between the hierarchical same index, the hierarchical isolation index, and the complete index can also be constructed through other basic mathematical operations, which will not be limited and elaborated herein.
[0065] It should be noted that in the embodiments of the present application, the Sigmoid function is used for normalization processing. In the embodiments of the present application, other normalization methods can also be selected, such as function transformation and other normalization methods, which will not be limited herein.
[0066] For step S103, determine the comprehensive importance of the string according to the semantic vector corresponding to the string.
[0067] In some embodiments, the determining the comprehensive importance of the string according to the semantic vector corresponding to the string includes: determining the local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string at all levels; determining the comprehensive importance based on the local importance.
[0068] In this embodiment, the power communication network security knowledge is multi-source, heterogeneous, and fragmented. At the same time, a large number of useful information fragments such as security knowledge bases and information bases are scattered everywhere in the power communication network, and the data from different sources changes the credibility of the data. For example, for the same relatively complete data, compared with the information data from Baidu web pages, the information data from CNKI has higher authority and better data quality.
[0069] In the power communication network security system, analyzing the authority of a string is a relatively complex task. Generally, it is considered that the higher the level of the subject to which the information source belongs, the higher the data authority; if the data appears more times in relevant laws, regulations, and regulatory documents, making the data more reliable, then the data authority is higher.
[0070] Since the target string may have different but same-meaning expressions in different files due to the reasons of the writers, it is easy to have errors when only using the similarity between strings to analyze the authority of keywords. Therefore, this application takes into account the context semantic features of the keywords where the string is located, and considers the authority of the target string by analyzing the semantic similarity degree between the analysis field where the target string is located and the analysis fields where the corresponding strings in each level of text sets are located, so as to ensure that the entities and relationships in the knowledge graph can accurately reflect the complex relationships and structures in the real world. The higher the authority of the target string, the higher its integrity in the power communication network security system, and integrity is an important indicator to measure the quality of the knowledge graph, which is directly related to whether the knowledge graph can comprehensively cover the knowledge in related fields. Therefore, by combining the semantic similarity degree between the analysis field where the target string is located and the analysis fields where the corresponding strings in each level of text sets are located, and the integrity index for analysis, the accuracy of the quality index of the target string is improved.
[0071] It should be noted that in the embodiments of this application, the analysis fields where the strings are located are vectorized by using the bag-of-words (BoW) model to obtain the semantic vectors of the analysis fields where each string is located; word embedding algorithms, bidirectional encoder representations from transformers (BERT), etc. can also be selected. Since the calculation of semantic vectors involves the semantic understanding of texts, the semantic vectors of the analysis fields where the target string is located and the corresponding analysis fields in the field sets at each level are different due to context, grammar, or phonetic factors.
[0072] The method for obtaining the comprehensive importance is as follows: The sum of the similarity indexes between the semantic vectors of the analysis fields where the target string is located and the target vectors at each level of the target string is used as the comprehensive similarity of the target string at each level; the product of the number of target vectors at each level of the target string and the comprehensive similarity is used as the local importance of the target string at each level; the sum of the local importance of the target string at all levels is recorded as the comprehensive importance of the target string.
[0073] Further, for step S104, the quality index of the string is determined based on the integrity index and the comprehensive importance.
[0074] In some embodiments, the quality index is calculated by the following formula:
[0075] In the formula, T is the quality index of the target string; Z is the integrity index of the target string; A is the total number of text sets at each level, which is equal to the total number of levels in the power communication network security system; is the total number of target vectors of the target string at the a-th level; is the semantic vector of the analysis field where the target string is located; is the b-th target vector of the target string at the a-th level; is the comprehensive similarity of the target string at the a-th level; is the local importance of the target string at the a-th level; is the comprehensive importance of the target string; cos is the cosine function; Sigmoid is the normalization function.
[0076] It should be noted that represents the level of the subject to which the target string belongs. When is larger, the target string has stronger relevance and importance in the data, so the level of the target string in the knowledge graph will be correspondingly higher, the authority of the analysis field where the target string is located is higher, and thus the quality of the analysis field where the target string is located is higher. When is larger, the similarity between the meaning of the analysis field where the target string is located and the semantics of the analysis field at the a-th level is higher, indicating that the analysis field where the target string is located can ensure that the entities and relationships in the knowledge graph can accurately reflect the complex relationships and structures in the real world, the authority of the analysis field where the target string is located is higher, and the quality of the analysis field where the target string is located is higher. When the complete index is larger, it indicates that the reliability of the target string is higher, that is, the energy level of the target string is higher, the analysis field where the target string is located should be more in the core position in the process of constructing the knowledge graph of power communication network security, and the quality index of the target string is larger.
[0077] The method for obtaining the quality index of each string in the set of all-level text is the same as the method for obtaining the quality index of the target string.
[0078] Further, for step S105, based on the quality index of the string, a knowledge graph is constructed.
[0079] In some embodiments, the constructing a knowledge graph based on the quality index of the string includes: using the corresponding strings with quality indexes higher than a preset threshold to construct the knowledge graph.
[0080] In this embodiment, for the quality index of each string in the analysis field of the overall text set, the analysis field where the string with a quality index greater than the preset quality threshold is located is recorded as the target field; the target field is high-quality data, and the target field with a larger quality index should be more in the core position of the knowledge graph of power communication network security; the knowledge graph of the power communication network is constructed using the high-quality target fields, making the knowledge graph more accurate and reliable, which is beneficial to improving the threat intelligence analysis ability, promoting the sharing and inheritance of security knowledge, improving the emergency response efficiency, helping maintenance personnel complete tasks such as decision-making guidance and instruction verification, and reducing the workload and operation risk.
[0081] It should be noted that in the embodiment of this application, the preset quality threshold takes an empirical value of 0.7, and the implementer can set it according to specific circumstances. In the embodiment of this application, the Neo4j graph database is selected to construct the knowledge graph, and community discovery algorithms, shortest path algorithms, node embedding algorithms, etc. can also be selected to construct the knowledge graph.
[0082] As can be seen from the above embodiments, the knowledge graph construction method described in the embodiments of this application divides each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; according to the number of times the string appears in the text set of each level and the number of times the string appears in the text set of the partition, the complete index of the string is determined; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; the comprehensive importance of the string is determined according to the semantic vector corresponding to the string; based on the complete index and the comprehensive importance, the quality index of the string is determined; based on the quality index of the string, a knowledge graph is constructed. In the embodiments of this application, each obtained analysis field is split into strings and counted in different levels and security partitions, so as to generate a complete index for each string. This index not only considers the frequency of the string appearing in the text sets of each level, but also considers its performance in specific security partitions. This multi-dimensional analysis method quantifies the level identity and isolation of each string, providing solid data support for subsequent quality assessment. Then, combined with the semantic vector corresponding to the string, the calculation of the comprehensive importance further enhances the expressiveness of the string in the knowledge graph, making the constructed knowledge graph not only reflect the quantitative characteristics of the data, but also incorporate deep semantic understanding. By normalizing the complete index and the comprehensive importance, the obtained quality index can effectively indicate the relative importance of the string, so as to ensure that in the process of constructing the knowledge graph, strings with a quality index higher than the set threshold are preferentially selected. This process not only improves the accuracy and reliability of the knowledge graph, but also provides a strong data foundation for the management and application of knowledge, ultimately promoting the in-depth mining of information and the improvement of application value.
[0083] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0084] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0085] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides a knowledge graph construction device.
[0086] Referring to Figure 2 , the knowledge graph construction device includes: A division module 21, configured to divide each obtained analysis field into strings; the analysis fields are located in different levels and different security partitions; A complete index determination module 22, configured to determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set composed of strings of all the security partitions within each level; the partition text set is a set composed of strings of each security partition; A comprehensive importance determination module 23, configured to determine the comprehensive importance of the string according to the semantic vector corresponding to the string; A quality index determination module 24, configured to determine the quality index of the string based on the complete index and the comprehensive importance; A construction module 25, configured to construct a knowledge graph based on the quality index of the string.
[0087] In a possible implementation manner, the complete index determination module 22 is further configured to: A level same index determination unit, configured to determine the level same index of the string according to the discrete index of the number of times the string appears in each level text set and the number of times the string appears in all level text sets; A hierarchical isolation index determination unit, configured to determine a hierarchical isolation index of the string according to a discrete index of the number of occurrences of the string in a partitioned text set and a difference degree of the occurrences of the string in all partitioned text sets corresponding to the level; A complete index determination unit, configured to determine a complete index of the string based on the same-level index and the hierarchical isolation index of the level.
[0088] In a possible implementation, the same-level index determination unit of the level is further configured to: Determine a quantity proportion according to a ratio of the number of occurrences of the string in each level text set to the total number of strings in the level text set; Calculate a cumulative sum of the quantity proportions of the string in all the level text sets to obtain a level importance; Calculate a level dispersion degree according to a discrete index of the number of occurrences of the string in all level text sets; Determine the same-level index of the level according to the level importance and the level dispersion degree.
[0089] In a possible implementation, the hierarchical isolation index determination unit of the level is further configured to: Calculate a partition dispersion degree according to a discrete index of the number of occurrences of the string in the partitioned text set; Calculate a cumulative sum of the difference degrees between the number of occurrences of the string in a certain partitioned text set and the number of occurrences of the string in the remaining partitioned text sets other than the certain partitioned text set in the corresponding level text set to obtain an occurrence difference degree; Calculate a cumulative sum of the occurrence difference degrees of all the partitioned text sets in the corresponding level text set to obtain a partition difference degree; Determine the hierarchical isolation index based on the partition dispersion degree and the partition difference degree.
[0090] In a possible implementation, the comprehensive importance determination module 23 is further configured to: Determine a local importance according to the similarity between the semantic vector corresponding to the string and the target vectors of the string in all levels; Determine the comprehensive importance based on the local importance.
[0091] In a possible implementation, the quality index determination module 24 is further configured to: Perform a normalization calculation on the complete index and the comprehensive importance to obtain the quality index.
[0092] In one possible implementation, the building block 25 is further configured to: Construct the knowledge graph by using the corresponding strings whose quality metrics are higher than a preset threshold.
[0093] For convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0094] The device in the above embodiment is used to implement the corresponding knowledge graph construction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.
[0095] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the knowledge graph construction method described in any of the above embodiments.
[0096] Figure 3 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0097] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0098] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0099] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices can include keyboards, mice, touchscreens, microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0100] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0101] The bus 1050 includes a path to transmit information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0102] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not have to include all the components shown in the figure.
[0103] The electronic device of the above embodiment is used to implement the corresponding knowledge graph construction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0104] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the knowledge graph construction method described in any of the foregoing embodiments.
[0105] The computer-readable medium of this embodiment includes both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device.
[0106] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the knowledge graph construction method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0107] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0108] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0109] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0110] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A knowledge graph construction method, characterized in that: include: Dividing each obtained analysis field into strings; the analysis fields are located at different levels and different security partitions; Determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set consisting of all the strings of the security partitions in each level; the partition text set is a set consisting of the strings of each security partition; Determining the comprehensive importance of the character string according to the semantic vector corresponding to the character string; Determining a quality indicator of the character string based on the complete indicator and the comprehensive importance; Based on the quality indicators of the character string, a knowledge graph is constructed.
2. The method according to claim 1, characterized in that: Determining the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set includes: Determining the level identity index of the character string according to the discrete index of the number of times the character string appears in each level text set and the number of times the character string appears in all level text sets; Determine the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference of the string in all partition text sets in the corresponding level; A complete index of the string is determined based on the hierarchical identity index and the hierarchical isolation index.
3. The method according to claim 2, characterized in that Determining the level identity index of the character string according to the discrete index of the number of times the character string appears in each level text set and the number of times the character string appears in all level text sets includes: Determine the quantity proportion according to the ratio of the number of times the character string appears in each of the hierarchical text sets to the total number of character strings in the hierarchical text sets; Calculate the cumulative sum of the number proportions of the character string in all the hierarchical text sets to obtain the hierarchical importance; Calculate the level dispersion according to the discrete index of the number of times the character string appears in all level text sets; The level identity index is determined according to the level importance and the level dispersion.
4. The method according to claim 2, characterized in that: Determining the hierarchical isolation index of the string according to the discrete index of the number of times the string appears in the partition text set and the appearance difference of the string in all partition text sets in the corresponding level includes: The partition discreteness is calculated according to the discrete index of the number of times the character string appears in the partition text set; Calculate the number of times the string appears in a certain partition text set, and add the difference between the number of times the string appears in other partition text sets in the corresponding hierarchical text set except the certain partition text set to obtain the occurrence difference; Calculating the cumulative sum of the occurrence differences of the character string in all partition text sets in the corresponding hierarchical text set to obtain the partition difference; The hierarchical isolation index is determined based on the partition discreteness and the partition difference.
5. The method according to claim 1, characterized in that The step of determining the comprehensive importance of the character string according to the semantic vector corresponding to the character string includes: Determining the local importance according to the similarities between the semantic vectors corresponding to the character string and the target vectors of all levels of the character string; The comprehensive importance is determined based on the local importance.
6. The method according to claim 1, characterized in that The determining the quality index of the character string based on the complete index and the comprehensive importance includes: The complete index and the comprehensive importance are normalized and calculated to obtain the quality index.
7. The method according to claim 1, characterized in that The constructing of a knowledge graph based on the quality index of the character string includes: The knowledge graph is constructed using the corresponding character strings whose quality indicators are higher than a preset threshold.
8. A knowledge graph construction device, characterized in that: include: A partitioning module is configured to partition each acquired analysis field into character strings; the analysis fields are located at different levels and different security partitions; The complete index determination module is configured to determine the complete index of the string according to the number of times the string appears in each level text set and the number of times the string appears in the partition text set; the level text set is a set consisting of all the strings of the security partitions in each level; the partition text set is a set consisting of the strings of each security partition; A comprehensive importance determination module, configured to determine the comprehensive importance of the character string according to the semantic vector corresponding to the character string; a quality indicator determination module, configured to determine a quality indicator of the character string based on the complete indicator and the comprehensive importance; The building module is configured to build a knowledge graph based on the quality index of the character string.
9. The device according to claim 8, characterized in that The complete indicator determination module is further configured as follows: a hierarchical identity index determining unit configured to determine the hierarchical identity index of the character string according to a discrete index of the number of times the character string appears in each hierarchical text set and the number of times the character string appears in all hierarchical text sets; A hierarchical isolation index determination unit is configured to determine the hierarchical isolation index of the string according to a discrete index of the number of times the string appears in a partition text set and a difference in the occurrence of the string in all partition text sets in the corresponding hierarchy; The complete index determining unit is configured to determine the complete index of the character string based on the hierarchical identity index and the hierarchical isolation index.
10. The device according to claim 9, characterized in that The level identity indicator determination unit is further configured to: Determine the quantity proportion according to the ratio of the number of times the character string appears in each of the hierarchical text sets to the total number of character strings in the hierarchical text sets; Calculate the cumulative sum of the number proportions of the character string in all the hierarchical text sets to obtain the hierarchical importance; Calculate the level dispersion according to the discrete index of the number of times the character string appears in all level text sets; The level identity index is determined according to the level importance and the level dispersion.
11. The device according to claim 9, characterized in that The hierarchical isolation index determination unit is further configured to: The partition discreteness is calculated according to the discrete index of the number of times the character string appears in the partition text set; Calculate the number of times the string appears in a certain partition text set, and add the difference between the number of times the string appears in other partition text sets in the corresponding hierarchical text set except the certain partition text set to obtain the occurrence difference; Calculating the cumulative sum of the occurrence differences of the character string in all partition text sets in the corresponding hierarchical text set to obtain the partition difference; The hierarchical isolation index is determined based on the partition discreteness and the partition difference.
12. The device according to claim 8, characterized in that The comprehensive importance determination module is further configured to: Determining the local importance according to the similarities between the semantic vectors corresponding to the character string and the target vectors of all levels of the character string; The comprehensive importance is determined based on the local importance.
13. The device according to claim 8, characterized in that The quality indicator determination module is further configured to: The complete index and the comprehensive importance are normalized and calculated to obtain the quality index.
14. The device according to claim 8, characterized in that The building blocks are further configured to: The knowledge graph is constructed using the corresponding character strings whose quality indicators are higher than a preset threshold.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge network construction method and device, equipment and storage medium
CN114706991A
Text theme segmentation method and device based on knowledge graph and electronic equipment
CN116340525A
Rapid construction and storage system for electric power threat intelligence knowledge graph
CN117453922A
Layered evaluation method, device and equipment for technical defense equipment network and storage medium
CN119011179A
Electric power safety knowledge graph construction method and device, electronic equipment and storage medium
CN119168034A