A method and system for visualizing and analyzing scientific and technological innovation data
By building a knowledge graph, calculating the stability of properties and update costs, establishing a priority queue, and updating scientific and technological innovation data using TransE model, the problem of untimely update of knowledge graphs is solved, and the accuracy and efficiency of data visualization analysis are improved.
Patent Information
- Application Number
- CN202411106089.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-08-13
AI Technical Summary
In the existing visual analysis of scientific and technological innovation data, there are problems in the process of updating the knowledge graph that the entity is not updated in time, resulting in reduced analysis accuracy, and excessive resource consumption.
By building a knowledge graph, obtain the attribute word matrix and synonyms of the original entity and the new entity, calculate the property stability, determine the update target entity, calculate the update cost and complexity based on the difference in path length and number of connected edges, establish a priority queue, and update using the TransE model.
It improves the accuracy and efficiency of knowledge graph updates, ensures timely updates of important entities, reduces the impact of error processing on the graph, and improves the accuracy of data visualization analysis.
Smart Images

Figure CN119067207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for visualizing and analyzing scientific and technological innovation data. Background Art
[0002] Scientific and technological innovation activities are highly complex and professional, involving a large amount of interdisciplinary and cross-domain data. How to effectively integrate and analyze these data and tap into the value of data is crucial to promoting scientific and technological innovation and promoting industrial upgrading. Data visualization provides a new analytical perspective and research tool for the field of scientific and technological innovation, which helps researchers gain insight into the connotation of data, optimize research methods, and accelerate the process of scientific and technological innovation. Knowledge graphs provide a new form of expression for data visualization by constructing a network of entities and their relationships. It helps to reveal the logical relationships behind complex data and discover potential connections between data. It has been widely used in search engines, recommendation systems and other fields. Introducing knowledge graphs into scientific and technological innovation data visualization can more intuitively show the integration and interaction of cross-domain data, providing a new perspective for scientific and technological innovation analysis.
[0003] When new knowledge is added to the knowledge graph, it is usually necessary to update the knowledge graph. The update of the knowledge graph is divided into data model layer update and data layer update. The data layer update refers to adding new entities and updating the attribute values of entities based on the existing data model. For example, when a new scientific research result is released, it is necessary to add related entities to the knowledge graph and update the relationship between them. The update of the data layer ensures that the content of the knowledge graph remains up-to-date and reflects the latest scientific and technological innovation results. The real-time update of the knowledge graph of scientific and technological innovation data is very important. How to solve the timeliness of the update process and not consume a lot of useless resources is a key issue faced in the process of updating scientific and technological innovation data. For example, in the traditional knowledge graph update and completion process, directly updating the entities to be updated causes important entities to not be updated in time, thereby reducing the accuracy of the visualization analysis of scientific and technological innovation data. Summary of the invention
[0004] The present invention provides a method and system for visualizing and analyzing scientific and technological innovation data to solve the existing problems.
[0005] A method and system for visualizing and analyzing scientific and technological innovation data of the present invention adopts the following technical solutions:
[0006] An embodiment of the present invention provides a method for visualizing and analyzing scientific and technological innovation data, the method comprising the following steps:
[0007] Acquire scientific and technological innovation data, construct a knowledge graph, and obtain a set of original entities and a set of new entities; the entities in the knowledge graph are interconnected, and there are several paths of different lengths between entities, and the length of each path is the number of edges connecting two entities;
[0008] Obtain the attribute word matrix and synonym set of each entity in the original entity set and the new entity set; obtain the property stability of each entity in the original entity set and each entity in the new entity set according to the difference between the attribute word matrix and synonym set of each entity in the original entity set and each entity in the new entity set; obtain several update target entities for each entity in the new entity set according to the property stability of each entity in the original entity set and each entity in the new entity set;
[0009] In the knowledge graph, the update cost of each updated target entity of each entity in the new entity set is obtained according to the difference in path length and number of connecting edges between each updated target entity of each entity in the new entity set and each connected entity; the complexity of each entity to be updated in the new entity set is obtained according to the update cost of each updated target entity of each entity in the new entity set;
[0010] According to the complexity of each entity to be updated in the new entity set, a new entity priority queue is obtained; the link completion prediction time of each entity is obtained; according to the new entity priority queue and the link completion prediction time of each entity, the completed scientific and technological innovation data knowledge graph is obtained.
[0011] Furthermore, the acquisition of the attribute word matrix and the synonym set of each entity in the original entity set and the new entity set includes the following specific steps:
[0012] Use a web crawler to obtain the attribute information of the p-th entity in the new entity set E', and obtain several introductory terms of the p-th entity;
[0013] All the introduction terms of the p-th entity are calculated using the Jieba word segmentation tool to obtain several word segmentation texts of the p-th entity;
[0014] All the segmented texts of the p-th entity are operated using the NLTK method to obtain several characteristic segmented texts of the p-th entity;
[0015] The matrix constructed from all the feature segmentation texts of the p-th entity is recorded as the attribute word matrix of the p-th entity;
[0016] Search the open source Chinese synonym table for the synonyms of each feature segmentation text of the attribute word matrix of the p-th entity in the new entity set E', and construct a set of synonyms of all feature segmentation texts of the attribute word matrix of the p-th entity in the new entity set E', which is recorded as the synonym set of the p-th entity;
[0017] According to the acquisition method of the attribute word matrix and synonym set of each entity in the new entity set, the attribute word matrix and synonym set of each entity in the original entity set are obtained.
[0018] Furthermore, according to the difference between the attribute word matrix and the synonym set of each entity in the original entity set and each entity in the new entity set, the specific calculation formula corresponding to the property stability of each entity in the original entity set and each entity in the new entity set is obtained as follows:
[0019]
[0020] in, Indicates the stability of the properties of the p-th entity in the new entity set E' and the j-th entity in the original entity set E; Represents the attribute word matrix of the pth entity in the new entity set E'; Represents the attribute word matrix of the jth entity in the set E of original entities; represents the number of synonyms in the intersection of the synonym set of the p-th entity in the new entity set E' and the synonym set of the j-th entity in the original entity set E; It represents the number of synonyms in the synonym set of the jth entity in the set E of original entities; Jac() represents the Jaccard correlation coefficient.
[0021] Furthermore, the method of obtaining several update target entities for each entity in the new entity set according to the property stability of each entity in the original entity set and each entity in the new entity set includes the following specific steps:
[0022] If the property stability of the pth entity in the new entity set E' and the jth entity in the original entity set E is greater than v, then the jth entity in the original entity set E is recorded as the update target entity of the pth entity in the new entity set E'; where v is a preset empirical threshold.
[0023] Furthermore, the update cost of each update target entity of each entity in the new entity set is obtained according to the difference in path length and number of connecting edges between each update target entity of each entity in the new entity set and each connected entity, including the following specific steps:
[0024] According to the difference in the number of connection edges between each update target entity of each entity in the new entity set and each connected entity, the association between each update target entity of each entity in the new entity set and each connected entity is obtained;
[0025] According to the association between each update target entity of each entity in the new entity set and each connected entity and the path length, the update cost of each update target entity of each entity in the new entity set is obtained.
[0026] Furthermore, according to the association between each update target entity of each entity in the new entity set and each connected entity and the path length, the specific calculation formula corresponding to the update cost of each update target entity of each entity in the new entity set is obtained as follows:
[0027]
[0028] in, represents the update cost of the sth update target entity of the pth entity in the new entity set E'; represents the number of entities connected to the sth update target entity of the pth entity in the new entity set E'; exp() represents an exponential function with a natural constant as the base; Represents the number of connecting edges between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; Represents the mean number of connection edges between the sth update target entity of the pth entity in the new entity set E' and all connected entities; Update the variance of the number of connected edges between the target entity and all connected entities for the sth pth entity in the new entity set E'; represents the qth path length between the sth update target entity and the ith connected entity of the pth entity in the E' set; represents the number of path lengths between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; Represents the association between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'.
[0029] Furthermore, the updating complexity of each entity in the new entity set is obtained according to the updating cost of each updating target entity of each entity in the new entity set, and the specific steps include the following:
[0030] The average of the update costs of all update target entities of the p-th entity in the new entity set E' is recorded as the update complexity of the p-th entity in the new entity set E'.
[0031] Furthermore, the step of obtaining a new entity priority queue according to the complexity to be updated of each entity in the new entity set includes the following specific steps:
[0032] The queue constructed by ordering each entity in the new entity set in descending order of complexity to be updated is recorded as the new entity priority queue.
[0033] Furthermore, the method of obtaining a completed scientific and technological innovation data knowledge graph based on the new entity priority queue and the link completion prediction time of each entity includes the following specific steps:
[0034] The sequence constructed by the first entity in the new entity priority queue is recorded as the updated entity sequence;
[0035] Step (1): The sum of the link completion prediction times of all entities in the updated entity sequence is recorded as the time threshold;
[0036] Step (2): If the time threshold is greater than the preset link completion time, all entities in the update entity sequence are input into the TransE model for update and completion, and the sequence constructed by the next entity in the new entity priority queue is recorded as the update entity sequence; if the time threshold is less than or equal to the preset link completion time, the next entity in the new entity priority queue is added to the update entity sequence, and steps (1) and (2) are repeated; until all entities in the new entity priority queue are input into the TransE model for update and completion; finally, the completed scientific and technological innovation data knowledge graph is obtained.
[0037] The present invention also proposes a scientific and technological innovation data visualization analysis system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program stored in the memory to implement the steps of the aforementioned scientific and technological innovation data visualization analysis method.
[0038] The beneficial effects of the technical solution of the present invention are:
[0039] Acquire scientific and technological innovation data, construct knowledge graphs, and obtain the original entity set and the new entity set; the entities in the knowledge graph are interconnected, and there are several paths of different lengths between entities, and the length of each path is the number of edges connecting two entities; obtain the attribute word matrix and synonym set of each entity in the original entity set and the new entity set; according to the differences between the attribute word matrix and synonym set of each entity in the original entity set and each entity in the new entity set, obtain the property stability of each entity in the original entity set and each entity in the new entity set; accurately reflect the stability between each entity in the original entity set and each entity in the new entity set. According to the property stability of each entity in the original entity set and each entity in the new entity set, obtain several update target entities for each entity in the new entity set, which improves the accuracy of data update. In the knowledge graph, according to the difference in path length and number of connecting edges between each update target entity of each entity in the new entity set and each connected entity, obtain the update cost of each update target entity of each entity in the new entity set; reduce the impact of the chain reaction caused by the wrong handling of entities on the accuracy and reliability of the entire knowledge graph. According to the update cost of each updated target entity of each entity in the new entity set, the complexity of each entity to be updated in the new entity set is obtained, further improving the accuracy of data update. According to the complexity of each entity to be updated in the new entity set, the new entity priority queue is obtained; the link completion prediction time of each entity is obtained; according to the new entity priority queue and the link completion prediction time of each entity, the completed scientific and technological innovation data knowledge graph is obtained. The present invention improves the accuracy of the visualization analysis of scientific and technological innovation data by analyzing each entity in the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0041] Figure 1 This is a flowchart of the steps of a method for visualizing and analyzing scientific and technological innovation data according to the present invention. DETAILED DESCRIPTION
[0042] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of a method and system for visualizing and analyzing scientific and technological innovation data proposed by the present invention, its specific implementation method, structure, features and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0043] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0044] The following describes in detail a specific scheme of a method and system for visualizing and analyzing scientific and technological innovation data provided by the present invention in conjunction with the accompanying drawings.
[0045] See also Figure 1 , which shows a flowchart of a method for visualizing and analyzing scientific and technological innovation data provided by an embodiment of the present invention, the method comprising the following steps:
[0046] Step S001: Acquire scientific and technological innovation data, construct a knowledge graph, and obtain a set of original entities and a set of new entities; the entities in the knowledge graph are interconnected, and there are several paths of different lengths between entities, and the length of each path is the number of edges connecting two entities.
[0047] Obtain scientific and technological innovation data such as research reports, learning materials, and survey materials from the company's database. Use the JieBa word segmentation tool to segment the acquired scientific and technological innovation data to obtain several segmented texts. Use the top-down method to construct the knowledge graph for all segmented texts.
[0048] It should be noted that the constructed knowledge graph is composed of multiple triples. The structure of the triple is [entity 1, relationship, entity 2]. The top-down method of constructing the knowledge graph and the JieBa word segmentation tool are well-known technologies, and the specific methods are not introduced here.
[0049] The constructed knowledge graph is G = (E, R), E = {E 1 ,E 2 ,…,E M} is the set of original entities in the knowledge graph, which is represented by points in the knowledge graph. 1 ,R 2 ,…,R N} is the set of relations in the knowledge graph, which is represented by edges in the knowledge graph. Among them, M is the number of entities in the set of original entities, E 1 is the first entity in the set of original entities, E2 is the second entity in the set of original entities, E M is the Mth entity in the set of original entities; N is the number of relational data in the set of relations in the knowledge graph, R 1 is the first relation data in the set of relations in the knowledge graph, R 2 is the second relation data in the set of relations in the knowledge graph, R N It is the Nth relation data in the set of relations in the knowledge graph.
[0050] Obtain new technological innovation data, and obtain the new entity set corresponding to the new technological innovation data according to the collection acquisition method of the original entity in the knowledge graph, which is recorded as E'={E' 1 ,E' 2 ,…,E' P}. Where P is the number of entities in the new entity set, E' 1 is the first entity in the new entity set, E' 2 is the second entity in the new entity set, E' P is the Pth entity in the new entity set.
[0051] It should be noted that each entity in the new entity set is an entity to be updated.
[0052] Step S002: Obtain the attribute word matrix and synonym set of each entity in the original entity set and the new entity set; obtain the property stability of each entity in the original entity set and each entity in the new entity set based on the differences between the attribute word matrix and synonym set of each entity in the original entity set and each entity in the new entity set; obtain several update target entities for each entity in the new entity set based on the property stability of each entity in the original entity set and each entity in the new entity set.
[0053] It should be noted that due to the limitation of update resources, the entity data that can be updated at one time in the knowledge graph of scientific and technological innovation data is also limited. Updating a large amount of data at one time may cause performance problems, such as long loading time and query delay. Therefore, it is necessary to determine the maximum number of entities in a single update process. Therefore, it is necessary to first calculate the property stability of each entity in the original entity set E and each entity in the new entity set E'.
[0054] Taking the pth entity in the new entity set E' as an example, a web crawler is used to obtain the attribute information of the pth entity in the new entity set E', and obtain several introduction entries of the pth entity. The web crawler is a well-known technology, and the specific method is not introduced in detail here.
[0055] All the introduction terms of the p-th entity are calculated using the Jieba word segmentation tool to obtain several word segmentation texts of the p-th entity. The Jieba word segmentation tool is a well-known technology, and the specific method is not introduced here.
[0056] All the segmented texts of the p-th entity are operated using the NLTK method to obtain several characteristic segmented texts of the p-th entity. The NLTK method is a well-known technology, and the specific method is not introduced here.
[0057] The matrix constructed from all the feature segmentation texts of the p-th entity is recorded as the attribute word matrix of the p-th entity.
[0058] It should be noted that the attribute word matrix of the p-th entity is a single-row matrix in which all feature segmentation texts of the p-th entity are arranged in the order in which they are acquired.
[0059] According to the method of obtaining the attribute word matrix of the pth entity in the new entity set E', the attribute word matrix of each entity in the original entity set E and the new entity set E' is obtained.
[0060] In the open source Chinese synonym table, search for the synonyms of each feature segmentation text of the attribute word matrix of the p-th entity in the new entity set E', and construct a set of synonyms of all feature segmentation texts of the attribute word matrix of the p-th entity in the new entity set E', which is recorded as the synonym set of the p-th entity.
[0061] According to the method of obtaining the synonym set of the p-th entity in the new entity set E', the synonym set of the original entity set E and each entity in the new entity set E' are obtained.
[0062] Taking the pth entity in the new entity set E' and the jth entity in the E set as an example, the specific calculation formula corresponding to the property stability of the pth entity in the new entity set E' and the jth entity in the original entity set E is:
[0063]
[0064] in, Indicates the stability of the properties of the p-th entity in the new entity set E' and the j-th entity in the original entity set E; Represents the attribute word matrix of the pth entity in the new entity set E'; Represents the attribute word matrix of the jth entity in the set E of original entities; represents the number of synonyms in the intersection of the synonym set of the p-th entity in the new entity set E' and the synonym set of the j-th entity in the original entity set E; Represents the number of synonyms in the synonym set of the jth entity in the set E of original entities; Jac() represents the Jaccard correlation coefficient; wherein, the Jaccard correlation coefficient is a method for comparing the similarities between sample sets, which is a well-known technology and the specific method will not be introduced here.
[0065] What needs to be explained is: The larger the value of , the higher the similarity between the attribute word matrix of the p-th entity in the new entity set E' and the j-th entity in the original entity set E. The larger the value of , the higher the semantic similarity between the p-th entity in the new entity set E' and the j-th entity in the original entity set E. The larger the value of , the greater the property stability of the corresponding p-th entity in the new entity set E' and the j-th entity in the original entity set E.
[0066] Through the above process, the property stability of the p-th entity in the new entity set E' and the j-th entity in the original entity set E is obtained.
[0067] Each entity in the new entity set E' and each entity in the original entity set E are operated according to the above process to obtain the property stability of each entity in the new entity set E' and each entity in the original entity set E.
[0068] If the property stability of the pth entity in the new entity set E' and the jth entity in the original entity set E is greater than v, the jth entity in the original entity set E is recorded as the update target entity of the pth entity in the new entity set E'. Wherein, v is a preset empirical threshold value, and the preset empirical threshold value v in this embodiment is 0.3, which is used as an example for description.
[0069] Through the above process, several update target entities for each entity in the new entity set E' are obtained.
[0070] It should be noted that: when there is no update target entity for a certain entity in the new entity set E', this embodiment sets the complexity to be updated to 0, and takes this as an example for description.
[0071] Step S003: In the knowledge graph, based on the differences in path lengths and the number of connecting edges between each updated target entity of each entity in the new entity set and each connected entity, the update cost of each updated target entity of each entity in the new entity set is obtained; based on the update cost of each updated target entity of each entity in the new entity set, the complexity of updating of each entity in the new entity set is obtained.
[0072] It should be noted that due to the different connection richness of the update target entities in the knowledge graph, the complexity corresponding to each update target is also different. Entities with higher complexity will have a greater impact on the surrounding entities, so more data updates will be involved. Mishandling these entities may lead to a chain reaction, affecting the accuracy and reliability of the entire graph. Therefore, it is necessary to calculate the complexity of these update target entities in the knowledge graph to determine the complexity of the pth entity to be updated in the new entity set E'.
[0073] It is further necessary to explain that: for the sth update target entity of the pth entity in the new entity set E', it is connected to multiple other entities in the knowledge graph. The path lengths between the entities are different, and the relationship strengths between the corresponding entities are different. Therefore, the update cost of the sth update target entity in the knowledge graph is determined by calculating the relationship strengths between the sth update target entity and the multiple other entities connected to it.
[0074] It should be noted that: the entities in the knowledge graph are interconnected, and there can be multiple paths of different lengths between entities. The path length refers to the number of edges connecting two entities.
[0075] In the knowledge graph, taking the sth update target entity of the pth entity in the new entity set E' as an example, the specific calculation formula for the update cost of the sth update target entity of the pth entity in the new entity set E' is:
[0076]
[0077] in, represents the update cost of the sth update target entity of the pth entity in the new entity set E'; represents the number of entities connected to the sth update target entity of the pth entity in the new entity set E'; exp() represents an exponential function with a natural constant as the base; Represents the number of connecting edges between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; Represents the mean number of connection edges between the sth update target entity of the pth entity in the new entity set E' and all connected entities; Update the variance of the number of connected edges between the target entity and all connected entities for the sth pth entity in the new entity set E'; represents the qth path length between the sth update target entity and the ith connected entity of the pth entity in the E' set; Represents the number of path lengths between the sth update target entity and the i-th connected entity of the p-th entity in the new entity set E'.
[0078] What needs to be explained is: Represents the association between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; The larger the value of indicates that the number of connecting edges between the sth update target entity of the pth entity in the new entity set E' and the i-th connected entity is greater than the number of connecting edges between other connected entities; and Inversely proportional, The smaller the value of , the greater the update cost of the sth update target entity of the pth entity in the new entity set E'. and Proportional, The larger the value of , the greater the update cost of the sth update target entity of the pth entity in the new entity set E'.
[0079] Through the above process, the update cost of the sth update target entity of the pth entity in the new entity set E' is obtained.
[0080] Perform the above operation on each update target entity of each entity in the new entity set E' to obtain the update cost of each update target entity of each entity in the new entity set E'.
[0081] The average of the update costs of all update target entities of the p-th entity in the new entity set E' is recorded as the update complexity of the p-th entity in the new entity set E'.
[0082] According to the above process, the complexity to be updated of each entity in the new entity set E' is obtained.
[0083] Step S004: According to the complexity of each entity to be updated in the new entity set, a new entity priority queue is obtained; the link completion prediction time of each entity is obtained; according to the new entity priority queue and the link completion prediction time of each entity, a completed scientific and technological innovation data knowledge graph is obtained.
[0084] It should be noted that: according to the above steps, the update complexity of each entity in the new entity set E' can be calculated. The update complexity of the entity determines the cost of updating the current entity. Therefore, in order to ensure that the entities that have the greatest impact on the structure and function of the knowledge graph can be updated in a timely manner during the subsequent knowledge graph update process, a priority queue is established for all entities to be updated in order of complexity to be updated from large to small, and entities with high complexity are updated first.
[0085] The queue constructed by ordering each entity in the new entity set in descending order of complexity to be updated is recorded as the new entity priority queue.
[0086] The sequence constructed by the first entity in the new entity priority queue is recorded as the updated entity sequence.
[0087] In the Update Entity sequence, do the following:
[0088] Step (1): The sum of the link completion prediction times of all entities in the updated entity sequence is recorded as the time threshold;
[0089] Step (2): If the time threshold is greater than the preset link completion time, all entities in the update entity sequence are input into the TransE model for update and completion, and the sequence constructed by the next entity in the new entity priority queue is recorded as the update entity sequence. If the time threshold is less than or equal to the preset link completion time, the next entity in the new entity priority queue is added to the update entity sequence, and steps (1) and (2) are repeated. Until all entities in the new entity priority queue are input into the TransE model for update and completion. Finally, the completed scientific and technological innovation data knowledge graph is obtained. Among them, the TransE model and the method for calculating the time for link completion prediction in the TransE model are well-known technologies, and the specific method will not be introduced here. Among them, the preset link completion time in this embodiment is 30 seconds, which is described as an example.
[0090] What needs to be explained is: the efficiency of the time representation model of link completion prediction in the TransE model when processing knowledge graph update and completion tasks.
[0091] According to the above method, the completed scientific and technological innovation data knowledge graph is obtained, so as to determine the real-time nature of information in the process of scientific and technological innovation data visualization, providing strong support for the innovation and decision-making of enterprises and researchers.
[0092] So far, the present invention is completed.
[0093] In summary, in an embodiment of the present invention, scientific and technological innovation data is acquired, a knowledge graph is constructed, a set of original entities and a set of new entities are obtained, and the property stability of the set of original entities and the set of new entities is calculated. According to the property stability, several update target entities for each entity in the new entity set are determined. In the knowledge graph, the update cost of each update target entity of each entity in the new entity set is obtained based on the difference in path length and number of connecting edges between each update target entity and the connected entity. According to the update cost of each update target entity of each entity, the complexity of each entity to be updated in the new entity set is calculated, and the completed knowledge graph of scientific and technological innovation data is further obtained. The present invention improves the accuracy of visual analysis of scientific and technological innovation data by analyzing each entity in the knowledge graph.
[0094] The present invention also provides a scientific and technological innovation data visualization analysis system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program stored in the memory to implement the steps of the aforementioned scientific and technological innovation data visualization analysis method.
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for visualizing and analyzing scientific and technological innovation data, characterized in that: The method comprises the following steps: Acquire scientific and technological innovation data, construct a knowledge graph, and obtain a set of original entities and a set of new entities; the entities in the knowledge graph are interconnected, and there are several paths of different lengths between entities, and the length of each path is the number of edges connecting two entities; Obtain the attribute word matrix and synonym set of each entity in the original entity set and the new entity set; obtain the property stability of each entity in the original entity set and each entity in the new entity set according to the difference between the attribute word matrix and synonym set of each entity in the original entity set and each entity in the new entity set; obtain several update target entities for each entity in the new entity set according to the property stability of each entity in the original entity set and each entity in the new entity set; In the knowledge graph, the update cost of each updated target entity of each entity in the new entity set is obtained according to the difference in path length and number of connecting edges between each updated target entity of each entity in the new entity set and each connected entity; the complexity of each entity to be updated in the new entity set is obtained according to the update cost of each updated target entity of each entity in the new entity set; According to the complexity of each entity to be updated in the new entity set, a new entity priority queue is obtained; the link completion prediction time of each entity is obtained; according to the new entity priority queue and the link completion prediction time of each entity, a completed scientific and technological innovation data knowledge graph is obtained; The specific calculation formula corresponding to the update cost of each update target entity of each entity in the new entity set is obtained according to the association between each update target entity of each entity in the new entity set and each connected entity and the path length: in, represents the update cost of the sth update target entity of the pth entity in the new entity set E'; represents the number of entities connected to the sth update target entity of the pth entity in the new entity set E'; exp() represents an exponential function with a natural constant as the base; Represents the number of connecting edges between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; Represents the mean number of connection edges between the sth update target entity of the pth entity in the new entity set E' and all connected entities; Update the variance of the number of connected edges between the target entity and all connected entities for the sth pth entity in the new entity set E'; represents the qth path length between the sth update target entity and the ith connected entity of the pth entity in the E' set; represents the number of path lengths between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'; Represents the association between the sth update target entity and the ith connected entity of the pth entity in the new entity set E'.
2. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: The specific steps of obtaining the attribute word matrix and synonym set of each entity in the original entity set and the new entity set are as follows: Use a web crawler to obtain the attribute information of the p-th entity in the new entity set E', and obtain several introductory terms of the p-th entity; All the introduction terms of the p-th entity are calculated using the Jieba word segmentation tool to obtain several word segmentation texts of the p-th entity; All the segmented texts of the p-th entity are operated using the NLTK method to obtain several characteristic segmented texts of the p-th entity; The matrix constructed from all the feature segmentation texts of the p-th entity is recorded as the attribute word matrix of the p-th entity; Search the open source Chinese synonym table for the synonyms of each feature segmentation text of the attribute word matrix of the p-th entity in the new entity set E', and construct a set of synonyms of all feature segmentation texts of the attribute word matrix of the p-th entity in the new entity set E', which is recorded as the synonym set of the p-th entity; According to the acquisition method of the attribute word matrix and synonym set of each entity in the new entity set, the attribute word matrix and synonym set of each entity in the original entity set are obtained.
3. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: According to the difference between the attribute word matrix and the synonym set of each entity in the original entity set and each entity in the new entity set, the specific calculation formula corresponding to the property stability of each entity in the original entity set and each entity in the new entity set is obtained as follows: in, Indicates the stability of the properties of the p-th entity in the new entity set E' and the j-th entity in the original entity set E; Represents the attribute word matrix of the pth entity in the new entity set E'; Represents the attribute word matrix of the jth entity in the set E of original entities; represents the number of synonyms in the intersection of the synonym set of the p-th entity in the new entity set E' and the synonym set of the j-th entity in the original entity set E; It represents the number of synonyms in the synonym set of the jth entity in the set E of original entities; Jac() represents the Jaccard correlation coefficient.
4. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: The method of obtaining a plurality of update target entities for each entity in the new entity set according to the property stability of each entity in the original entity set and each entity in the new entity set includes the following specific steps: If the property stability of the pth entity in the new entity set E' and the jth entity in the original entity set E is greater than v, then the jth entity in the original entity set E is recorded as the update target entity of the pth entity in the new entity set E'; where v is a preset empirical threshold.
5. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: The method of obtaining the update complexity of each entity in the new entity set according to the update cost of each update target entity of each entity in the new entity set includes the following specific steps: The average of the update costs of all update target entities of the p-th entity in the new entity set E' is recorded as the update complexity of the p-th entity in the new entity set E'.
6. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: The step of obtaining a new entity priority queue according to the complexity to be updated of each entity in the new entity set includes the following specific steps: The queue constructed by ordering each entity in the new entity set in descending order of complexity to be updated is recorded as the new entity priority queue.
7. According to claim 1, a method for visualizing and analyzing scientific and technological innovation data, characterized in that: The specific steps of obtaining a completed scientific and technological innovation data knowledge graph based on the new entity priority queue and the link completion prediction time of each entity are as follows: The sequence constructed by the first entity in the new entity priority queue is recorded as the updated entity sequence; Step (1): The sum of the link completion prediction times of all entities in the updated entity sequence is recorded as the time threshold; Step (2): If the time threshold is greater than the preset link completion time, all entities in the update entity sequence are input into the TransE model for update and completion, and the sequence constructed by the next entity in the new entity priority queue is recorded as the update entity sequence; if the time threshold is less than or equal to the preset link completion time, the next entity in the new entity priority queue is added to the update entity sequence, and steps (1) and (2) are repeated; until all entities in the new entity priority queue are input into the TransE model for update and completion; finally, the completed scientific and technological innovation data knowledge graph is obtained.
8. A system for visualizing and analyzing scientific and technological innovation data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of a method for visualizing and analyzing scientific and technological innovation data as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Time knowledge graph completion method and system based on gated recurrent neural network
CN116108188A
Knowledge graph construction method and device, computer equipment and storage medium
CN117271802A