A Method and Device for Generating an Enhanced Technical Knowledge Graph

By obtaining text data collections, extracting and combining technical entities to form a strengthened technical knowledge graph and associated with general knowledge graphs, the problem of insufficient extraction and association of technical entities in the existing technology is solved, and better support for technical related analysis and clear display of the graph is achieved.

CN115269782BActive Publication Date: 2025-05-27INST OF SCI & TECHN INFORMATION OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210955172.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2025-05-27
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

The existing knowledge graphs have insufficient support for technical related analysis, and it is difficult to effectively extract and associate technical entities, resulting in insufficient support for technical related analysis.

Method used

By obtaining the text data set, technical entities are extracted based on statistical methods and preset extraction rules, and combined to form a reinforced technical knowledge graph and associated with the general knowledge graph.

Benefits of technology

It realizes clear distinction and multi-dimensional query of technical entities, reduces the number of technology-related entities, improves the clarity of the map and retrieval recall, and understands technological development and changes through time factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269782B_ABST
    Figure CN115269782B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligence analysis technology, and in particular, to a method and apparatus for generating an enhanced technical knowledge graph. The method includes extracting technical entities belonging to the types of methods, processes, devices, and functions respectively based on statistical methods and in combination with preset extraction rules; merging the technical entities according to their respective types, performing replacement processing on all technical entities in the specific text data by merging the technical entities, adding relationships to the corresponding merged technical entities, and associating them with the general knowledge graph corresponding to the specific text data to form a knowledge graph of enhanced technology. The present invention has a clear distinction for entities related to the technology itself, can query the knowledge graph from multiple different technical dimensions, and reduces the number of various entities related to the technology through merging, making it clearer in the graph display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligence analysis, and particularly to a method and device for generating an enhanced technical knowledge graph. Background Art

[0002] A knowledge graph is a form of knowledge representation that can be used to describe knowledge such as common sense in academia and industry. Essentially, a knowledge graph is a semantic network, that is, a knowledge base with a directed graph structure, where the nodes of the graph represent entities or concepts, and the edges of the graph represent various semantic relationships between entities / concepts. As a means of knowledge representation and storage, the knowledge graph is considered a means to solve long-term challenges in cognitive intelligence and dilemmas such as the interpretability of deep learning because of its strong expression ability, good scalability, and ability to balance human cognition and machine automatic processing. The knowledge graph has gradually become popular in academia and industry. It was first applied in the search field and now has been applied in multiple fields such as healthcare, e-commerce, finance, military, power, education, and public security. For example, it is applied in credit assessment, risk control, and anti-fraud in the financial field, as well as intelligent consultation in the healthcare field. However, most of the current general knowledge graphs mainly focus on general entities such as people, places, and institutions, and there are deficiencies in the extraction and association methods of technical entities themselves.

[0003] Technical entities themselves have the characteristics of having many requirements and diverse definitions. For example, technology can be a means to achieve human purposes and may be a method, process, or device; technology can also be a collection of practices and components (i.e., a technical body); technology can also be regarded as a collection of devices and engineering practices that can be utilized in a certain culture (i.e., technical elements). Technology itself is an important subject of scientific and technological intelligence work, mainly including a system of knowledge, methods, and skills for designing, manufacturing, adjusting, operating, and monitoring various artificial things and artificial processes. Many types of scientific and technological intelligence work are closely related to technology, such as disruptive technology identification, technology association, technology life cycle analysis, technology tree, and technology efficacy matrix analysis. Therefore, it is necessary to strengthen technical entities on the basis of general knowledge graphs.

[0004] Currently, there are mainly two concepts in knowledge graphs related to technology. One is the knowledge graph of a specific field, which involves the technical part of the field and is also a simple classification of technology; the other is visualization, which uses tools such as CiteSpace to perform technical analysis on the technology in the field. However, the current knowledge graphs all have the following problems: they are relatively general, mainly solving basic problems such as finding people, places, and institutions, and providing insufficient support for technology-related analysis in intelligence analysis. Summary of the Invention

[0005] The object of the present invention is to provide a method and device for generating an enhanced technical knowledge graph, so as to solve the foregoing problems existing in the prior art.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] The present invention provides a method for generating an enhanced technical knowledge graph, and the method includes:

[0008] Obtain a text data set, where the text data set contains a plurality of specific text data;

[0009] Based on statistical methods and combined with preset extraction rules, extract technical entities belonging to the types of methods, processes, devices, and functions respectively;

[0010] Merge the technical entities according to their respective types to obtain merged technical entities;

[0011] Replace all technical entities in the text data set with the merged technical entities to obtain merged text data;

[0012] When the co-occurrence times between different types of merged technical entities in the merged text data exceed the first co-occurrence threshold, add relationships to the corresponding merged technical entities; and associate them with the general knowledge graph corresponding to the specific text data to form a knowledge graph of enhanced technology.

[0013] Preferably, the extraction rules include a stop word list set, and the stop word list set includes a general stop word list and a technical stop word list.

[0014] Preferably, the extraction rules include a relationship matrix; the relationship matrix is a two-dimensional matrix, and the two dimensions are respectively a function module and a technical entity type, and the function module is a set obtained by classifying the text data set; the matrix value represents the possibility that the corresponding technical entity type appears in the corresponding function module, and if there is a possibility, the corresponding matrix value is 1, otherwise it is 0; when extracting technical entities, only extract the corresponding type of technical entities in the function module with a matrix value of 1.

[0015] Preferably, merging the technical entities according to their respective types specifically includes:

[0016] Select any type of technical entity as the first technical entity;

[0017] Obtain a first merging set for the first technical entity based on the literal similarity method;

[0018] Obtain a second merging set for the first technical entity based on the edit distance similarity calculation method;

[0019] Obtain the third merged set regarding the first technical entity based on the deep learning word2vec similarity method;

[0020] Obtain the co-occurrence merged sets of the first technical entity and any other type of technical entity respectively based on the co-occurrence statistics method;

[0021] Perform merged sorting according to the number of occurrences of the merged results in each merged set, and obtain the merged results whose number of occurrences meets the merging threshold.

[0022] Preferably, obtaining the co-occurrence merged sets of the first technical entity and any other type of technical entity respectively based on the co-occurrence statistics method specifically includes:

[0023] Select any other type of technical entity except the first technical entity as the second technical entity;

[0024] Obtain the second technical entity Yi co-occurring with the first technical entity Xi, and obtain the second technical entity Xj co-occurring with the first technical entity Xj. When the second technical entity Yi is the same as the second technical entity Yj, the co-occurrence times of the first technical entity Xi and the first technical entity Xj are incremented by 1;

[0025] Merge the first technical entities whose co-occurrence times meet the second co-occurrence threshold;

[0026] Wherein, the first technical entity Xi and the first technical entity Xj are different technical entities in the first technical entity.

[0027] Preferably, when the second technical entity Yi is different from the second technical entity Yj, obtain the number C1 of the same third technical entities co-occurring with the second technical entity Yi and the second technical entity Yj and the number C2 of the same fourth technical entities;

[0028] If αC1 + βC2 > γ, the co-occurrence times of the first technical entity Xi and the first technical entity Xj are incremented by 1;

[0029] Wherein, 0 < α ≤ 0.5, 0 < β ≤ 0.5, γ ≥ 1, and the third technical entity and the fourth technical entity are other types of technical entities except the first technical entity and the second technical entity respectively.

[0030] Preferably, the method further includes generating a time binary tuple of the merged technical entity based on the release time corresponding to the merged text data, and the time binary tuple includes the earliest release time and the latest release time corresponding to the merged technical entity.

[0031] Correspondingly, the present invention also provides a generating device for an enhanced technical knowledge graph, which is used to implement any of the above methods, including:

[0032] An acquisition module, configured to acquire a text data set, where the text data set contains a plurality of specific text data;

[0033] An extraction module, configured to extract technical entities belonging to the types of methods, processes, devices, and functions respectively based on statistical methods and in combination with preset extraction rules;

[0034] A first processing module, configured to merge the technical entities according to their respective types to obtain merged technical entities;

[0035] A second processing module, configured to perform replacement processing on all technical entities in the text data set through the merged technical entities to obtain merged text data;

[0036] A third processing module, configured to add relationships to the corresponding merged technical entities when the co-occurrence times between different types of merged technical entities in the merged text data exceed a first co-occurrence threshold; and associate with the general knowledge graph corresponding to the specific text data to form a knowledge graph of enhanced technology.

[0037] The beneficial effects of the present invention are:

[0038] The present invention provides a method and a device for generating an enhanced technical knowledge graph, which clearly distinguish entities related to the technology itself, can query the knowledge graph from multiple different technical dimensions of methods, processes, devices, and functions, reduce the number of various entities related to the technology through merging, can be clearer in graph display, and when retrieving, the merged records can be used as entry words to improve the recall rate; and can also more scientifically understand the development and changes of the technology itself through time elements. In addition, the enhanced knowledge graph of this technology does not deviate from the general knowledge graph and is still associated with the existing general knowledge graph, with wider applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flowchart of a method for generating an enhanced technical knowledge graph provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0041] As Figure 1 shown, the present invention provides a method for generating an enhanced technical knowledge graph, and the method includes:

[0042] S101. Obtain a text data set, where the text data set contains multiple specific text data.

[0043] S102. Based on statistical methods and combined with preset extraction rules, extract technical entities belonging to the types of methods, processes, devices, and functions respectively.

[0044] S103. Merge the technical entities according to their respective types to obtain merged technical entities.

[0045] S104. Perform replacement processing on all technical entities in the text data set through the merged technical entities to obtain merged text data.

[0046] S105. When the co-occurrence times between different types of merged technical entities in the merged text data exceed the first co-occurrence threshold, add relationships to the corresponding merged technical entities; and associate with the general knowledge graph corresponding to the specific text data to form a knowledge graph for strengthening technologies.

[0047] In this embodiment, the extraction of technical entities can adopt statistical methods such as based on CRF. The extraction method of technical entities is a conventional means in the prior art, and the specific extraction process will not be elaborated here as long as technical entities regarding methods, processes, devices, and functions can be extracted. The preset extraction rules in this embodiment mainly include a stop word list set and a relationship matrix.

[0048] The stop word list set is several stop word lists serving for technical concept recognition, including several general stop word lists and technical stop word lists. Each stop word list is a text file, and each stop word is a line. The general stop word lists come from the stop words used in general natural language processing, or the union of different stop word lists can be directly taken. The technical stop word list is summarized or reconstructed according to the existing word lists. For example, currently, it can be clearly known that all personal names, place names, and organization names are not technologies and can be part of the technical stop word list. In addition, according to the definition of technologies, those terms that do not belong to methods, processes, devices, and functions can be added and supplemented.

[0049] The relationship matrix is a two-dimensional matrix, and the two dimensions are respectively function modules and technical entity types. The function modules are the sets obtained by classifying the text data set; the matrix value represents the possibility of the corresponding technical entity type appearing in the corresponding function module, and the form is shown in Table 1. If a certain technical entity may appear in a function module, the value in the matrix is 1, otherwise it is 0. When extracting technical entities, only extract the corresponding type of technical entities in the function modules where the matrix value is 1. For example: According to the matrix values in Table 1, only extract method-type technical entities in the function modules of title, keywords, and innovation points, and do not extract them in the function modules of uses and advantages.

[0050] Table 1

[0051]

[0052]

[0053] The classification of functional modules is based on specific text data types, which can be summarized according to existing classification patterns or text characteristics. The classification of functional modules for different types of technical texts can be different. For example, a certain classification of functional modules for patent abstract texts is "innovative points / uses / advantages", and a certain classification of functional modules for paper abstract texts is "purposes / methods / results / conclusions". Usually, the long text parts of specific text data need to be segmented into different classifications, while short texts such as titles and keywords do not need to be segmented and can be directly used as one functional module. For the classification of functional modules, traditional machine learning methods including KNN, SVM, etc. or methods based on deep learning language models can be selected.

[0054] In this embodiment, the technical entities are merged separately according to their respective types. Since the merging methods for technical entities of each type are similar, the following takes the technical entities of the "process" type as an example for illustration, and the specific process is as follows:

[0055] (1) Use the literal similarity method to obtain the first merging set I1 for "process".

[0056] (2) Use the edit distance similarity calculation method to obtain the second merging set I2 for "process".

[0057] (3) Use the deep learning word2vec similarity method to obtain the third merging set I3 for "process".

[0058] (4) Use the process-method co-occurrence statistics to obtain the fourth merging set I4 for "process".

[0059] (5) Use the process-device co-occurrence statistics to obtain the fifth merging set I5 for "process".

[0060] (6) Use the process-function co-occurrence statistics to obtain the sixth merging set I6 for "process".

[0061] Sort the merges in descending order according to the number of occurrences of the merge results in the 6 merge sets (the number of sets involved in the merge results), and preset a merge threshold as needed (the value range of the merge threshold is [1-6]). Only select the merge results greater than or equal to the merge threshold as the merge sets of the process. In addition, for English and other languages, the Stemming stemming method can also be used to obtain the seventh merge set I7 regarding the "process", and 7 merge sets are used in the above merge method. In this embodiment, the co-occurrence statistics methods of I4, I5, and I6 can be calculated separately or jointly, and the process is as follows:

[0062] 1) Represent the method entity by M, the process entity by P, the device entity by U, and the function entity by F.

[0063] 2) When calculating I4, when calculating the co-occurrence times of two process entities Pi and Pj with respect to the method entity, for any co-occurring method entity Mi in Pi and any co-occurring method entity Mj in Pj, if Mi = Mj, the co-occurrence times are incremented by 1. If Mi ≠ Mj, for separate calculation, the co-occurrence times do not change; for joint calculation, it is also necessary to further calculate the number of the same device entities and function entities co-occurring with Mi and Mj, and record them as C1 and C2 respectively. If αC1 + βC2 > γ, the co-occurrence times are incremented by 1, where 0 < α ≤ 0.5, 0 < β ≤ 0.5, and γ ≥ 1. Finally, according to the co-occurrence times, the process entities whose co-occurrence times meet the co-occurrence threshold f1 are merged.

[0064] 3) When calculating I5, when calculating the co-occurrence times of two process entities Pi and Pj with respect to the device entity, for any co-occurring device entity Ui in Pi and any co-occurring device entity Uj in Pj, if Ui = Uj, the co-occurrence times are incremented by 1. If Ui ≠ Uj, for separate calculation, the co-occurrence times do not change; for joint calculation, it is also necessary to further calculate the number of the same method entities and function entities co-occurring with Ui and Uj, and record them as C1 and C2 respectively. If αC1 + βC2 > γ, the co-occurrence times are incremented by 1, where 0 < α ≤ 0.5, 0 < β ≤ 0.5, and γ ≥ 1. Finally, according to the co-occurrence times, the process entities whose co-occurrence times meet the co-occurrence threshold f2 are merged.

[0065] 4) When calculating I6, when calculating the co-occurrence times of two process entities Pi and Pj with respect to a functional entity, for any co-occurring functional entity Fi in Pi and any co-occurring functional entity Fj in Pj, if Fi = Fj, the co-occurrence times are incremented by 1. If Fi ≠ Fj, for individual calculation, the co-occurrence times remain unchanged; for combined calculation, it is also necessary to further calculate the number of identical method entities and device entities in which Fi and Fj co-occur, and denote them as C1 and C2 respectively. If αC1 + βC2 > γ, the co-occurrence times are incremented by 1, where 0 < α ≤ 0.5, 0 < β ≤ 0.5, and γ ≥ 1. Finally, based on the co-occurrence times, the process entities whose co-occurrence times meet the co-occurrence threshold f3 are merged.

[0066] In this embodiment, on the basis of method, process, device, and function merging, it is also necessary to supplement the calculation of the associations between various technical entities to form a technology-centered knowledge graph. The method is as follows:

[0067] Use the merged set of methods, processes, devices, and functions to batch process any specific text data in the text data set, replace all the merged entity words with the unique URI in the merged set, and use the words of the merged entities as the labels of the URI to form a replacement set of the text data set, that is, the merged text data. Finally, calculate the associations of various technical entities in the graph through the co-occurrence method. Specifically:

[0068] For methods (number 1), processes (number 2), devices (number 3), and functions (number 4), set the co-occurrence thresholds f12 (representing the method-process co-occurrence threshold, and the same for the following), f13, f14, f21, f23, f24, f31, f32, f34, f41, f42, f43 for determining relationships. The thresholds are all integers, and fij = fji, i ≠ j. Calculate the co-occurrence times c12 (representing the method-process co-occurrence times in the merged text data, and the same for the following), c13, c14, c21, c23, c24, c31, c32, c34, c41, c42, c43 of different types of technical entities represented by two URIs in the merged text data. If cji ≥ fij, add a relationship for the two technical type entities.

[0069] In this embodiment, it is also necessary to associate the above technology-centered knowledge graph with the general knowledge graph. Specifically, using the merged text data as a medium, associate methods, processes, devices, and functions with a specific text data (if the URI representing a certain technical entity appears in the specific text data of the merged text data, establish an association between this method and the specific text data), and further associate the people, institutions, and locations involved in the specific text data (these information can be obtained from the general knowledge graph or the metadata of the specific text data), so as to form a technology-enhanced knowledge graph.

[0070] In this embodiment, it is also necessary to establish start and end times for various technical entities based on the time of the text data set, so as to more scientifically understand the development and changes of the technology itself through time factors. The text data set is stored in the form of text files. Each specific text data is represented by a text file. All files are in a folder. The files are named by numbers or letters. In addition, there is a file named "time", which includes the time of all files. Each file and its release time are on one line, separated by "#". For example, the format is: abcdefg#2021-12-31 18:30:45, indicating that the content of the text named "abcdefg.txt" was released at "2021-12-31 18:30:45". Finally, in the merged text data, according to the release time of the specific text data, record the binary tuple t=(ts, te) for each method, process, device, and functional entity, where ts is the earliest occurrence time and te is the latest occurrence time. When generating t for the first text, ts = te. As the number of updated texts changes, ts and te are also continuously updated.

[0071] By adopting the above technical solutions disclosed in the present invention, the following beneficial effects are obtained:

[0072] The present invention provides a method and device for generating an enhanced technology knowledge graph, which clearly distinguishes the entities related to the technology itself, can query the knowledge graph from multiple different technical dimensions of methods, processes, devices, and functions. Through merging, the number of various entities related to the technology is reduced, and the graph display can be clearer. And when retrieving, the merged record can be used as an entry word to improve the recall rate; it can also more scientifically understand the development and changes of the technology itself through time factors. In addition, this enhanced knowledge graph technology does not deviate from the general knowledge graph and is still associated with the existing general knowledge graph, with wider applicability.

[0073] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also fall within the protection scope of the present invention.

Claims

1. A method for generating an enhanced technical knowledge graph, characterized in that, the method includes: Obtain a text data set, which contains multiple specific text data; Based on statistical methods and combined with preset extraction rules, extract technical entities belonging to the types of methods, processes, devices, and functions respectively; Merge the technical entities according to their respective types to obtain merged technical entities; Perform replacement processing on all technical entities in the text data set through the merged technical entities to obtain merged text data; When the co-occurrence times between different types of merged technical entities in the merged text data exceed the first co-occurrence threshold, add relationships to the corresponding merged technical entities; and associate with the general knowledge graph corresponding to the specific text data to form a knowledge graph of enhanced technology.

2. The method for generating an enhanced technical knowledge graph according to claim 1, characterized in that, the extraction rules include a stop word list set, and the stop word list set includes a general stop word list and a technical stop word list.

3. The method for generating an enhanced technical knowledge graph according to claim 1, characterized in that, the extraction rules include a relationship matrix; the relationship matrix is a two-dimensional matrix, and the two dimensions are respectively a function module and a technical entity type, and the function module is a set obtained by classifying the text data set; The matrix value represents the possibility that the corresponding technical entity type appears in the corresponding function module. If there is a possibility, the corresponding matrix value is 1, otherwise it is 0; when extracting technical entities, only extract the corresponding type of technical entity in the function module where the matrix value is 1.

4. The method for generating an enhanced technical knowledge graph according to claim 1, characterized in that, Merging the technical entities according to their respective types specifically includes: Select a technical entity of any type as the first technical entity; Obtain a first merged set regarding the first technical entity based on the literal similarity method; Obtain a second merged set regarding the first technical entity based on the edit distance similarity calculation method; Obtain a third merged set regarding the first technical entity based on the deep learning word2vec similarity method; Based on the co-occurrence statistics method, obtain the co-occurrence merged sets of the first technical entity and any other type of technical entity respectively; Perform merging sorting according to the number of occurrences of the merging results in each merged set, and obtain the merging results whose number of occurrences meets the merging threshold.

5. The method for generating an enhanced technical knowledge graph according to claim 4, characterized in that, Based on the co-occurrence statistics method, obtaining the co-occurrence merged sets of the first technical entity and any other type of technical entity respectively specifically includes: Select any other type of technical entity except the first technical entity as the second technical entity; Obtain the second technical entity Yi co-occurring with the first technical entity Xi, and obtain the second technical entity Xj co-occurring with the first technical entity Xj. When the second technical entity Yi is the same as the second technical entity Yj, the co-occurrence times of the first technical entity Xi and the first technical entity Xj are incremented by 1; Merge the first technical entities whose co-occurrence times meet the second co-occurrence threshold; Among them, the first technical entity Xi and the first technical entity Xj are different technical entities in the first technical entities.

6. The method for generating an enhanced technical knowledge graph according to claim 5, characterized in that: When the second technical entity Yi is different from the second technical entity Yj, the number C1 of the same third technical entities co-occurring with the second technical entity Yi and the second technical entity Yj and the number C2 of the same fourth technical entities are obtained; If αC1 + βC2 > γ, the co-occurrence times corresponding to the first technical entity Xi and the first technical entity Xj are incremented by 1; Among them, 0 < α ≤ 0.5, 0 < β ≤ 0.5, γ ≥ 1, and the third technical entity and the fourth technical entity are other types of technical entities other than the first technical entity and the second technical entity respectively.

7. The method for generating an enhanced technical knowledge graph according to claim 1, characterized in that, further comprising: Generating a time tuple of the merged technical entity based on the release time corresponding to the merged text data, where the time tuple includes the earliest release time and the latest release time corresponding to the merged technical entity.

8. An apparatus for generating an enhanced technical knowledge graph, characterized in that, for implementing the method according to any one of claims 1 to 7, including: An acquisition module for acquiring a text data set, where the text data set contains a plurality of specific text data; An extraction module for extracting technical entities belonging to the types of methods, processes, devices, and functions respectively based on statistical methods and in combination with preset extraction rules; A first processing module for merging the technical entities according to their respective types to obtain merged technical entities; A second processing module for performing replacement processing on all technical entities in the text data set through the merged technical entities to obtain merged text data; A third processing module for adding relationships to the corresponding merged technical entities when the co-occurrence times between different types of merged technical entities in the merged text data exceed the first co-occurrence threshold; and associating with the general knowledge graph corresponding to the specific text data to form a knowledge graph of enhanced technology.

Citation Information

Patent Citations

  • Method and device for combining entities in knowledge map

    CN104484459A

  • Data processing method and device and computer readable storage medium

    CN114328799A