Knowledge graph generation method and system fusing multi-source science and technology data visualization results

Through intelligent retrieval optimization and graph neural network implicit relationship prediction, the problems of multi-source data fusion and visualization optimization are solved, a highly readable knowledge graph is generated, data consistency and visualization effects are improved, and research topics and cross-domain relationships are revealed.

CN120633807APending Publication Date: 2025-09-12UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510745078.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing knowledge graph generation methods have significant shortcomings in multi-source data fusion, intelligent preprocessing, implicit relationship mining and visualization optimization. It is difficult to integrate multi-source heterogeneous data, process noise, mine semantic associations and generate highly readable charts.

Method used

It uses intelligent retrieval optimization, graph neural network implicit relationship prediction and dynamic layout algorithm to generate highly readable knowledge graphs through multi-source data fusion, data preprocessing and visualization optimization.

Benefits of technology

It achieves comprehensive collection of multi-source data and keyword association expansion, improves data consistency and visualization structure clarity, deeply reveals the overlap of research topics and cross-domain collaborative relationships, and dynamically presents the core nodes and community evolution paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633807A_ABST
    Figure CN120633807A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge graph generation method and system fusing multi-source science and technology data visualization results. The invention provides a systematic solution for solving the problems that in the prior art, multi-source heterogeneous data fusion is difficult, preprocessing intellectualization is insufficient, implicit relation mining is insufficient, and the visualization effect is poor. The method comprises the steps that a data import module obtains data through local and web crawlers, and a keyword is expanded through a GPT model; the preprocessing module is used for cleaning, synonym merging and field standardization; the information processing module generates one-dimensional graph data and a two-dimensional triple, wherein an implicit relationship is predicted by a BERT model and a graph neural network; the graph structure planning module adopts force-oriented layout and dimensionality reduction optimization visualization; and the co-imported network graph is divided into communities through a difference filter and a Louvain algorithm. According to the method, multi-source data is effectively integrated, noise is automatically processed, semantic association is mined, the high-readability chart is generated, and the construction efficiency and the visualization effect of the knowledge graph are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graphs and provides a method and system for generating knowledge graphs by integrating visualization results of multi-source scientific and technological data. Background Art

[0002] Existing technologies for knowledge graph construction and visualization face numerous limitations. First, data sources are limited and processing is complex: Traditional methods typically rely on a single database (such as the Web of Science) or local files, making it difficult to integrate heterogeneous data from multiple sources (e.g., open platform papers retrieved through web crawlers, relational databases, and graph databases). Multi-source data formats vary widely (e.g., CSV files versus triple structures), resulting in inefficient data import and field parsing, and a lack of intelligent search optimization mechanisms.

[0003] Secondly, data preprocessing capabilities are insufficient: existing technologies rely on manual rules or simple statistical methods (such as mean imputation) to handle missing values, outliers, and duplicate records, making them difficult to adapt to the complex noise scenarios in scientific research data. For example, missing paper titles or duplicate records may lead to redundant knowledge graph nodes and frequent pseudo-clustering, affecting the accuracy of subsequent analysis. In addition, synonym merging and term disambiguation rely on static rule bases or low-precision similarity matching, which cannot dynamically adapt to the evolution of domain terminology and semantic association mining.

[0004] Third, field standardization and implicit relationship mining are limited: Traditional methods inadequately standardize fields like time and text, leading to data scale imbalances in visualizations such as trend charts and word frequency histograms. Furthermore, existing technologies rely solely on explicit triples (such as author-institution affiliations) to construct knowledge graphs, neglecting implicit semantic associations (such as overlap in research topics and cross-disciplinary collaborations), making it difficult to fully reflect the deep structure of knowledge networks.

[0005] Furthermore, graph visualization is suboptimal: Existing techniques for designing large-scale knowledge graph layouts often employ static force-directed algorithms or simple dimensionality reduction methods (e.g., PCA). This results in uneven node distribution, excessive edge density, and difficulty in clearly representing the core structure. Traditional methods for community segmentation in co-citation network graphs rely on modularity optimization algorithms (e.g., Louvain), but lack the dynamic identification of high-centrality nodes and cluster label mapping, limiting the visualization of the evolutionary paths of research topics.

[0006] In summary, existing knowledge graph generation methods have significant shortcomings in multi-source data fusion, intelligent preprocessing, implicit relationship mining, and visualization optimization. There is an urgent need for a systematic solution that can integrate multi-source data, automatically process noise, mine semantic relationships, and generate highly readable graphs. This invention effectively addresses these technical challenges by introducing intelligent retrieval optimization, implicit relationship prediction using graph neural networks, and a dynamic layout algorithm. Summary of the Invention

[0007] The purpose of this invention is to solve the technical problems of knowledge graph generation and display caused by multi-source heterogeneous data fusion, inconsistent data quality, difficulty in extracting implicit relationships, and complex visualization structure.

[0008] In order to achieve the above-mentioned purpose, the present invention adopts the following technical means:

[0009] The present invention provides a method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data, comprising the following steps:

[0010] Step 1: Data import module: Loads raw data from local and online sources. Local sources include CSV / TSV files, relational databases, and graph databases. Online sources are collected through automated crawlers based on user-entered keywords. These keywords are then expanded into related terms, synonyms, or domain terms through intelligent search optimization.

[0011] Step 2: Data preprocessing module: cleans the imported data, merges synonyms, disambiguates, and standardizes fields, including missing value processing, outlier detection, and duplicate record merging;

[0012] Step 3: Information processing module: Generate one-dimensional graph data and two-dimensional graph triplets. The one-dimensional graph includes a year-number of papers trend chart, a keyword frequency histogram, and an author / institution activity chart. The two-dimensional graph constructs triplets through explicit entity extraction and implicit relationship prediction. The implicit relationships are used to complete the explicit triplets based on graph neural networks.

[0013] Step 4: Graph structure planning module: Performs visual layout design based on one-dimensional graph data and two-dimensional graph triplets, including scale normalization, sliding average smoothing, and force-directed layout algorithm;

[0014] Step 5: Graph generation module: Render the planned data into visual charts, including scientific knowledge graphs and co-citation network graphs, and map them into two-dimensional space through dimensionality reduction algorithms (such as t-SNE or UMAP).

[0015] In the above scheme, step 1 includes the following steps:

[0016] Step 1.1: Load raw data from local sources and network sources, where the local sources include:

[0017] Step 1.1.1: CSV / TSV files are used to directly import metadata data exported from Web of Science, including author, institution, title, and publication journal fields;

[0018] Step 1.1.2: Extract data from a relational database, such as MySQL or PostgreSQL, by automatically constructing query statements.

[0019] Step 1.1.3: Graph databases, including Neo4j, are used to store complex relationship graphs between papers and extract graph data in the form of triples through query statements;

[0020] Step 1.2: Collect data from the Internet through the following steps:

[0021] Step 1.2.1: Receive keywords input by the user, use the pre-trained GPT model to generate expansion words, synonyms or domain terms related to the keywords, and generate a top-n expansion keyword list;

[0022] Step 1.2.2: Construct a Boolean search formula based on the extended keywords and collect paper data from online database platforms (CNKI, arXiv, Springer) through automatic crawlers;

[0023] Step 1.3: Automatically identify and classify fields in local CSV / TSV files, relational databases, and online data sources, including:

[0024] Step 1.3.1: Parse the header names and content and map the fields to predefined categories, including author field AU, title field TI, institution field C1, keyword field DE, reference field CR, document type field DT, and abstract field AB;

[0025] Step 1.3.2: Format the triple data extracted from the graph database into a (Subject, Predicate, Object) structure for subsequent construction of a two-dimensional graph.

[0026] In the above scheme, step 2 includes the following steps:

[0027] Step 2.1 Data cleaning:

[0028] Step 2.1.1: Missing value processing, including the following two methods:

[0029] Step 2.1.1.1: Delete missing values: If important fields (such as AU and TI) are missing in the record and their contribution to the overall analysis is small, delete the record;

[0030] Step 2.1.1.2: Mean / Median Filling: For numeric fields (e.g., PY, citation count), fill missing values ​​with the mean or median of the field;

[0031] Step 2.1.2: Outlier detection using the box plot method: Let the 25th percentile of the data set be Q1, the 75th percentile be Q3, calculate the interquartile range IQR = Q3-Q1, and define the outlier threshold as:

[0032] Lower Bound=Q1-1.5IQR, Upper Bound=Q3+1.5IQR

[0033] Values ​​outside this range were marked as outliers and extracted for subsequent analysis;

[0034] Step 2.1.3: Merge duplicate records and use the hash function hashlib.blake2b library to generate a unique key: take the DOI field as the priority. If there is a conflict, combine the TI, AU, PY, and SO fields to calculate the hash value h = hash(DOI, TI, AU, PY, SO) and merge records with the same hash value.

[0035] Step 2.2: Synonym merging and disambiguation:

[0036] Step 2.2.1: Convert the text into vectors and use the pre-trained model to calculate word vectors;

[0037] Step 2.2.2: Calculate cosine similarity Among them, v1 and v2 are word vectors;

[0038] Step 2.2.3: Set a threshold θ. If CosSim(v1, v2)>θ, merge them into synonyms (e.g., "AI" and "artificial intelligence").

[0039] Step 2.3 Field Standardization:

[0040] Step 2.3.1: Unify the time format and convert the field PY to the ISO 8601 format YYYY-MM-DD;

[0041] Step 2.3.2: Lowercase the text and remove punctuation. Perform the following operations on the fields TI, AU, and AB:

[0042] Step 2.3.2.1: Convert all letters to lowercase;

[0043] Step 2.3.2.2: Remove punctuation marks “:” and “.”.

[0044] In the above scheme, step 3 includes the following steps:

[0045] Step 3.1: One-dimensional graph data generation:

[0046] Step 3.1.1: Extract the year field PY or date type field DT from the preprocessed data, and calculate the mapping relationship between the year and the number of papers {y i :n i}, where y i Indicates the year, n i Indicates the number of papers published in that year, used to generate a year-paper number trend chart;

[0047] Step 3.1.2: Use the word segmenter to extract high-frequency keywords from the keyword field DE and the abstract AB, and calculate the mapping relationship between keywords and the number of occurrences {k j :c j}, where k j Indicates keywords, c j Indicates the number of occurrences, used to generate keyword frequency histograms;

[0048] Step 3.1.3: Extract the outlier records detected in the data cleaning module, sort them by time, and generate a graph of important nodes;

[0049] Step 3.2: 2D graph triple generation:

[0050] Step 3.2.1: Explicit Entity Extraction:

[0051] Step 3.2.1.1: Co-authorship: Extract the author list from the author field AU and construct a co-authorship triplet by pairing them together. <AU m , CoAuthor, AU n >, the edge weight is the co-authorship frequency w mn , where AU m and AU n For the author;

[0052] Step 3.2.1.2: Institutional Affiliation: Parse the mapping between author and institution from the institution field C1 and construct a triple <AU m ,AffiliatedWith,Institution p >, Institution p Indicates an institution;

[0053] Step 3.2.1.3: Author-research topic relationship: Extract the association between author and keyword from the institution field DE and construct triples <AU m ,ResearchOn,k j >, k j Indicates keywords;

[0054] Step 3.2.2: Implicit Relationship Prediction:

[0055] Step 3.2.2.1: Encode the summary text into a vector v using the pre-trained BERT model q , where v q The semantic vector representing the summary q;

[0056] Step 3.2.2.2: Construct the initial graph based on explicit triples, with node features defined by v q and explicit metrics, where explicit metrics include the frequency of cooperation w mn 、Keyword frequency c j ;

[0057] Step 3.2.2.3: Randomly hide some edges in the training set, use the graph neural network to predict the type of the hidden edges, and add implicit edges of the "ResearchOverlap" or "TopicSimilarity" relationship class in the full graph inference;

[0058] Step 3.3: Constructing the co-citation matrix:

[0059] Step 3.3.1: Parse the reference field CR and the reference set R of each paper i ={r1, r2, ..., r n}, calculate all unordered pairs (R j , R k ) to generate the co-occurrence matrix M, where M[i][j] represents the co-citation counts of documents i and j;

[0060] Step 3.3.2: Parse the reference field CR and the reference set R of each paper i ={r1, r2, ..., r n}, count all unordered pairs (r j , r k )’s co-citation count;

[0061] Step 3.3.3: Construct the co-citation matrix M N×N , where N is the total number of references, and the matrix element M[j][k] represents the number of references r j and r k The number of co-citations;

[0062] Step 3.3.4: Threshold filter the co-occurrence matrix, retain edges with M[i][j] ≥ 2, and normalize the matrix rows to eliminate the popular literature bias:

[0063] The normalization formula is:

[0064]

[0065] Where N is the total number of documents, Mnorm [i][j] represents the normalized co-citation weights.

[0066] In the above scheme, step 4 includes the following steps

[0067] Step 4.1: One-dimensional graph structure planning:

[0068] Step 4.1.1: Use Min-Max normalization for the year-paper number trend graph:

[0069]

[0070] Among them, x i represents the original value of the number of papers in the i-th year, x′ i is the normalized value, max(x) and min(x) are the number of papers in the largest and smallest years in the dataset, respectively;

[0071] Step 4.1.2: Logarithmic transformation of keyword frequency histogram:

[0072] x′ i =log b (x i +c)

[0073] Among them, x i is the number of times the keyword appears, x′ i is the transformed value, b is the logarithmic base, and c is the smoothing constant;

[0074] Step 4.1.3: Z-score normalization of the author / institution activity graph:

[0075]

[0076] Among them, x i is the activity of the author / institution (such as the number of publications), μ is the mean, and σ is the standard deviation;

[0077] Step 4.1.4: Sliding average smoothing of the year-paper number trend graph, using the sliding mean formula with a window size of k:

[0078]

[0079] Among them, s i is the smoothed value of the ith year, x j is the original data point;

[0080] Step 4.1.5: Visual compression strategy of keyword frequency histogram, using color to distinguish frequency levels: according to x′ i The value is divided into intervals and mapped to different colors, transparency or column widths;

[0081] Step 4.1.6: Apply the top-k strategy to the author / institution activity graph, filter the top k largest values ​​and classify them as “others”.

[0082] Step 4.2 Two-dimensional graph structure planning:

[0083] Step 4.2.1: Extract the backbone structure of the scientific knowledge graph using the difference filter formula:

[0084]

[0085] Among them, k i is the number of edges connecting node i, S i is the sum of edge weights of node i, α ij Indicates the significance of edge (i, j), retaining the edges that satisfy α ij <a(, a is the significance threshold;

[0086] Step 4.2.2: Force-directed layout algorithm, calculate the attractive and repulsive forces between nodes:

[0087] Step 4.2.2.1: Initialize node positions: Assign initial two-dimensional coordinates (x i ,y i ), where x i and y i Represents the horizontal and vertical coordinates of the i-th node;

[0088] Step 4.2.2.2 Calculate the repulsive force: For all node pairs (i, j), calculate the repulsive force according to the formula:

[0089]

[0090] Where k is the constant control chart scaling ratio;

[0091] Step 4.2.2.3 Calculate the attraction: For all connected edges (ij,), calculate the attraction according to the formula;

[0092]

[0093] where d ij is the Euclidean distance between nodes, (i, j) represents the connecting edge;

[0094] Step 4.2.2.4 Superimpose the force vectors and update the position: Add the net force F of the repulsive force and the attractive force total (i)=∑ j (F r (i, j)+F a (i, j)) acts on node i and updates its coordinates:

[0095] Where Δt is the time step;

[0096] Step 4.2.2.5 Add cooling factor: In each iteration, add cooling factor in proportion to α t =1-t / T max The number of attenuation movement amplitude is:

[0097] Δt′=αt·Δt

[0098] T max is the maximum number of iterations;

[0099] Step 4.2.2.6 Termination condition judgment: The iteration is terminated if any of the following conditions are met:

[0100] The number of iterations reaches the preset maximum value T max ;

[0101] The node position change amplitude is less than the threshold;

[0102] Step 4.2.2.7 Output node coordinates: Output the final converged two-dimensional coordinates (x′ i , y′ i ) as the visual layout result;

[0103] Step 4.2.3: Dimensionality reduction and mapping to two-dimensional space: Use t-SNE or UMAP algorithm to map the high-dimensional embedding vector to two-dimensional coordinates (x′ i , y′ i ), preserving local / global topology;

[0104] Step 4.2.4: Extract the core structure of the co-citation network graph using the same difference filter formula as step 4.2.1 and retain the top-k largest weight edges of each node;

[0105] Step 4.2.5: Divide the co-citation network graph into communities using the Louvain algorithm, which includes the following sub-steps:

[0106] Step 4.2.5.1: Initialize the community: assign each node to an independent community, that is, each node is initially an independent community;

[0107] Step 4.2.5.2: Node movement optimization: traverse all nodes and move the current node to the community to which its neighboring nodes belong. If the modularity Q is improved, keep this move;

[0108] Step 4.2.5.3: Network reconstruction: Use the currently divided communities as new nodes to rebuild the network graph. The edge weights between the new nodes are the sum of all edge weights between the original communities.

[0109] Step 4.2.5.4: Iterative optimization: Repeat steps 4.2.5.2 to 4.2.5.3 until the modularity Q reaches the convergence condition;

[0110] The calculation formula of modularity Q is:

[0111]

[0112] Among them, w ij is the edge weight between nodes i and j, indicating the number of co-citations, k i is the degree of node i, ∑ i,j w ij is the total edge weight of the network, δ(c i , c j ) is the indicator function, if c i =c j Then take 1, otherwise 0, c i is the community number to which node i belongs.

[0113] In the above scheme, step 5 includes the following sub-steps:

[0114] Step 5.1 One-dimensional graph generation module:

[0115] Step 5.1.1 Generate a trend chart of the number of papers by year:

[0116] Set the year field y i and the corresponding number of papers n i Mapped to a set of data points {(x i ,y i )}, where: x i Indicates the year position, y i Indicates the number of papers in that year. Draw a continuous line connecting all the points. Each data point (x i ,y i ) and mark the number of circles with fixed radius; the horizontal axis is the year X, and the vertical axis is the number of papers Y;

[0117] Step 5.1.2 Generate keyword frequency histogram:

[0118] The keyword k j and the number of occurrences c j Mapped to a set of coordinate points {(x j , h j )),in:

[0119] x j Indicates the position of the jth keyword, h j Indicates the frequency of the keyword, for each keyword in position x j Draw a rectangle at , the parameters are:

[0120] Left border:

[0121] Top edge: top = y base -h i ,

[0122] Width: width=w,

[0123] Height: h j

[0124] The keyword name is displayed in the center below the rectangle, the Y-axis represents the frequency, and the X-axis is the discrete classification coordinate system;

[0125] Step 5.1.3 Generate author / institution activity graph:

[0126] Author or institution A m Activity level w m Mapped to a set of bar data points {(y m , w m )},in:

[0127] y m Indicates the vertical coordinate position of the author / institution, w m Indicates activity, draws a horizontal bar with the starting point fixed on the left and a length of w m , the coordinates are:

[0128] Horizontal axis position: x = x0,

[0129] Vertical axis:

[0130] Bar width: w m ,

[0131] Bar Press w m Sort from largest to smallest, with the author / institution name marked on the left or right;

[0132] Step 5.2: 2D graph generation module:

[0133] Step 5.2.1: Scientific knowledge graph generation:

[0134] Step 5.2.1.1: For each entity e i The two-dimensional coordinates (x′ i , y′ i ) is normalized, and the formula is:

[0135]

[0136] in:

[0137] x i ,yi is the original coordinate, W and H are the width and height of the canvas.

[0138] Step 5.2.1.2: Draw circular nodes: size represents centrality, color represents category, category includes author, institution, keyword, that is, degree size i Map to the diameter of the drawn circle:

[0139]

[0140] where d i is the node degree, r min ,r max is the minimum / maximum node size;

[0141] Step 5.2.1.3: For the triple (e i ,r,e j ), at the coordinate (x′ i , y′ i ) and (x′ j , y′ j ) with edge widths determined by weights:

[0142]

[0143] where w ij is the edge weight, β is the scaling factor;

[0144] Step 5.2.1.4: Add entity names to the nodes to generate a clearly structured scientific knowledge graph;

[0145] Step 5.2.2: Generate co-citation network diagram:

[0146] Step 5.2.2.1: For each document node v i The two-dimensional coordinates (x′ i , y′ i ) for normalization;

[0147] The node size is determined by the degree, and the formula is:

[0148]

[0149] where d i is the node degree, r min , r max is the minimum / maximum node size;

[0150] Step 5.2.2.2: Classify labels c according to community k Mapping colors, different clusters display different color blocks;

[0151] Step 5.2.2.3: Opposite side (vi , v j ) Draw a line segment with edge width determined by weight;

[0152] Step 5.2.2.4: Add background color blocks or dotted borders to different communities and display cluster labels at the centroid.

[0153] The present invention also provides a knowledge graph generation system that integrates the visualization results of multi-source scientific and technological data, characterized in that the system is stored in a storage medium in a program manner, and when the processor executes the program, it implements the knowledge graph generation method that integrates the visualization results of multi-source scientific and technological data.

[0154] Because the present invention adopts the above technical means, it has the following beneficial effects:

[0155] 1. The present invention solves the technical problems of single data source and low retrieval efficiency of traditional methods by adopting intelligent retrieval optimization technology (such as pre-training GPT model to generate extended keywords) and multi-source data fusion mechanism (integration of local database, web crawler and graph database) in step 1, and achieves the effect of comprehensively collecting multi-source heterogeneous data and improving the accuracy of keyword association expansion.

[0156] 2. The present invention solves the technical problems of inconsistent data quality, terminology redundancy and noise interference through the synonym merging and disambiguation technology based on hash functions and pre-trained models in step 2, combined with field standardization processing (such as time format unification and text normalization), and achieves the effect of enhancing data consistency, reducing node redundancy and improving the accuracy of subsequent analysis.

[0157] 3. The present invention introduces a graph neural network in step 3 to predict implicit relationships (such as the "Research Overlap" relationship) of explicit triples, and combines it with the row normalization processing of the co-citation matrix to solve the technical problem of insufficient mining of implicit semantic associations, thereby achieving the effect of deeply revealing the overlap of research topics and cross-domain collaborative relationships and building a high-completeness knowledge network.

[0158] 4. The present invention solves the technical problems of uneven node distribution and excessive edge density in large-scale graphs by adopting a backbone structure extraction algorithm (difference filter) and dynamic force-directed layout optimization (combined with cooling factor and dimensionality reduction mapping) in step 4, thereby achieving the effect of improving the clarity of the visual structure and dynamically presenting the core nodes and community evolution paths.

[0159] 5. The present invention solves the technical problem of ambiguous expression of co-citation network topics by combining dimensionality reduction rendering with community partitioning label mapping (such as the Louvain algorithm to generate clustering background color blocks) in step 5, thereby achieving the effect of enhancing the efficiency of research topic identification and intuitively displaying the relationship between high centrality nodes and knowledge base. BRIEF DESCRIPTION OF THE DRAWINGS

[0160] Figure 1 It is a simplified flow chart of the present invention. DETAILED DESCRIPTION

[0161] The following is a detailed description of the embodiments of the present invention. Although the present invention will be described and illustrated in conjunction with certain specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, modifications or equivalent substitutions of the present invention are intended to fall within the scope of the claims of the present invention.

[0162] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. It will be understood by those skilled in the art that the present invention can also be implemented without these specific details.

[0163] Example

[0164] 1. Data import module

[0165] In the first step, we need to load raw data from different sources (such as CSV files, local databases, network databases) and format the data for subsequent processing. There are two main data sources: local sources and network sources.

[0166] Local source data mainly includes:

[0167] CSV / TSV file: A common format exported by Web of Science. The data includes meta-information of the paper, such as author, institution, title, and publication journal.

[0168] Relational databases: such as MySQL / PostgreSQL, suitable for large data sets.

[0169] Graph database: such as Neo4j, used to store complex relationship graphs between papers.

[0170] Network source data is mainly obtained through automatic crawlers, which crawl paper data from online databases (such as CNKI, Arxiv, Springer and other platforms) based on subject keywords.

[0171] For CSV / TSV files, import them directly. For relational databases and web-sourced data, first construct search keywords. Intelligent search optimization is used to intelligently optimize user-entered keywords, automatically expanding related keywords, synonyms, or domain terms to create more precise search formulas. Pre-trained models (such as GPT) are used to generate expanded terms related to the input keywords.

[0172] GPT (Generative Adversarial Pre-training) is better at generating natural language text related to the input. The GPT model can further expand the user's input keywords and generate more relevant field terms. The main process is as follows:

[0173] The keyword entered by the user is input as a prompt into the GPT model, and the GPT model will generate a series of expanded words or phrases that may be related to the keyword. Select the top-n keywords: Based on the relevant vocabulary output by the GPT model, the generated keywords are sorted according to their relevance or weight to the input keyword, and the top n keywords are selected. These keywords will serve as the basis for constructing the search formula. For relational databases, query statements are automatically constructed and data is extracted. For data from the Internet, a search formula (such as a Boolean search formula) is constructed based on the keywords, and then relevant data is automatically crawled on the corresponding platform (such as CNKI, Arxiv, Springer, etc.).

[0174] After obtaining these three parts of data, the system automatically identifies the fields. This system is mainly for data in the Web of Science format. For Web of Science data, the system can automatically parse the header into predefined field categories through preliminary analysis of field names and content. Automatic classification is based on data content. For example:

[0175] AU: Indicates the author, the field type is "text". TI: Indicates the title of the paper, the field type is "text". C1: Indicates the institution, the field type is "text". DE: Indicates the keywords of the paper, the field type is "text". CR: Indicates the reference list, the field type is "text". DT: Indicates the document type (such as article, conference paper), the field type is "text" ID: Indicates the keyword or subject ID, the field type is "ID". AB: Indicates the abstract, the field type is "text". LA: Indicates the language, the field type is "text". RP: Indicates the corresponding author, the field type is "text". EM: Indicates the email address of the corresponding author, the field type is "text".

[0176] A graph database is a database specifically designed to store and manage graph-structured data. It represents and stores data in the form of triples. A triple is the basic unit of knowledge representation and is in the form of:

[0177] (Subject, Predicate, Object). For example (Beijing is the capital of China).

[0178] In this system, triples are used to construct two-dimensional graphs (such as scientific knowledge graphs). After constructing a graph database query statement based on the keywords above, the triple results are obtained and used to subsequently construct the two-dimensional graph.

[0179] 2. Data Preprocessing Module

[0180] The purpose of data preprocessing is to clean, process, and transform raw data into a format suitable for analysis and visualization. The goal of data cleaning is to address noise or inconsistencies in the data to provide an accurate data foundation for subsequent analysis. Key processes include data cleaning, synonym merging and disambiguation, and field standardization.

[0181] (1) Data cleaning

[0182] First, missing values ​​are processed. Missing values ​​are a common phenomenon in data sets, especially in scientific research data. Missing values ​​may come from the loss of information during the data collection process, or the user has not filled in certain fields. In order to avoid interference with subsequent analysis, missing values ​​must be processed. Delete missing values: If an important field (such as author, title) in a record is missing and the record contributes little to the overall analysis, the record can be deleted. For example, if a paper lacks title information (TI), the row of data can be deleted. Mean / median filling: For numerical data (such as citation number, publication year, etc.), the mean or median of the field can be used to fill in missing values. For example, if the number of citations of a paper is missing, it can be filled with the mean number of citations of all papers.

[0183] Then outlier detection is performed. Outliers refer to values ​​in the data set that deviate significantly from the majority of data, in the paper data. These values ​​may be values ​​with special significance, such as articles with a number of citations significantly higher than other averages, or articles with obvious contributions or significance in the development of a certain field. In this process, this data is extracted for subsequent chart drawing. This system uses the box plot (IQR) method: IQR is a statistic of data distribution, which represents the range between the 75th percentile and the 25th percentile. When the data exceeds the limit of 1.5 times the IQR, it can be considered an outlier.

[0184] Finally, duplicate records are merged. In this project, the data sources include both local databases and automatic web crawling. The sources are diverse and the structures are different, so duplicate records are inevitable. If deduplication is not performed, it will lead to: some nodes in the co-citation network are repeated, resulting in pseudo clustering; entities in the knowledge graph are repeated, seriously affecting the quality of triples;

[0185] The same paper is counted multiple times in the chart, and the data is distorted. It is necessary to identify and merge duplicate records based on specific fields (such as DOI, ID, etc.). This system uses multiple fields to determine uniqueness. Such as TI (title), AU (author list), PY (year), SO (journal / conference), DOI (unique identifier). The hash function adopts the hashlib.blake2b library. Blake2b is a cryptographic hash function provided in the Python standard library hashlib. It divides the data into blocks, mixes rounds, and compresses it to ensure that even if the input changes slightly, the output is completely different. Hash (DOI) is used as the unique key, and other identifiers are used for duplicate detection when a conflict is detected.

[0186] (1) Synonym merging and disambiguation

[0187] There are often problems with synonyms, abbreviations and term disambiguation in the data. In order to ensure the standardization and consistency of the data, these problems must be dealt with. This system uses word vector similarity (Cosine similarity) + threshold merging. For those synonyms that do not have clear standard vocabulary, they can be identified and merged by calculating the similarity of word vectors (for example, using Word2Vec, GloVe and other methods). For example, if the cosine similarity of "AI" and "artificial intelligence" in the word vector space is greater than a certain threshold (such as 0.9), they can be regarded as synonyms and merged. The main process is to convert the text into a vector, calculate the cosine similarity between word vectors, and merge similar terms according to the set threshold.

[0188] (2) Field standardization

[0189] In this step, all data will be further standardized to ensure that it meets the needs of subsequent visualization. First, the time format is unified. All fields involving time (such as publication year) are unified into the ISO8601 format (YYYY-MM-DD). For example, if the publication date of an article is 2019, it will be formatted as 2019-01-01. Then the text is lowercase and punctuation is removed. For text data (such as paper titles, author names, abstracts, etc.), all letters are converted to lowercase and unnecessary punctuation is removed to ensure uniformity. For example, the paper title "AI in Catalysis: A Review" will be converted to "lowercase and punctuation removed": Original text: AI in Catalysis: A Review. After standardization: aiincatalysis a review.

[0190] After data is standardized, it will meet the requirements for subsequent chart construction. This standardized dataset can be used for: One-dimensional graphs: For example, the "Year" field has been standardized to the YYYY-MM-DD format, making it easier to create trend charts. Scientific knowledge graphs: For example, the "Author" and "Institution" fields have been standardized, allowing them to be directly used in graph construction. Co-citation network graphs: For example, the "References" field has been standardized, making it easier to calculate co-citation relationships between documents.

[0191] 3. Information Processing Module

[0192] For one-dimensional graphs, you only need to extract the corresponding information from the processed data. First, extract the year from the field PY (year) or DT (date type). Associate the year with each record to form a year-number of papers relationship. For example: {"2022": 15, "2023": 32, "2024": 20}. The extracted time distribution relationship data is used for trend charts. Then extract keywords from the fields DE (Author Keywords) and abstract AB. Use a word segmenter (such as TF-IDF, jieba, KeyBERT, etc.) to extract high-frequency keywords. Form a keyword-number of occurrences relationship. For example: {"ai": 56, "catalysis": 41, "electrocatalysis": 23}. The extracted keyword distribution is used for a word frequency histogram. Finally, the article data with special contributions detected in the outlier detection step in the first module are arranged in chronological order for drawing important node diagrams.

[0193] For two-dimensional graphs, the cleaned data needs to be converted into triples. The entities and relationships in them need to be extracted, including explicit and implicit entities.

[0194] For display entities, we extract them in a rule-based way. Co-authorship relationships are constructed from the AU (author) field, combining two of them. We introduce "co-authorship frequency" as a weight to take into account the strength of co-authorship (edge ​​weight). For example:<Zhang,CoAuthor4,Wang> Author and institution affiliation: Extract the author and their corresponding institution from the C1 field (author institution mapping). If there are multiple institutions, parse multiple C1 or C3 blocks. For example:<Zhang,AffiliatedWith,Univ Auckland> Authors and Research Topics: From the DE (Keywords) field, associate each author with a keyword. For example:

[0195] <Wang,ResearchOn,Large Language Model>Author and Reference: Create a citation triple from the CR (reference) field, usually using the first author, year, and journal information of the cited article to identify a document. For example:<This Paper,Cites,XX> .

[0196] For implicit entities, relationships that cannot be explicitly expressed by structured fields, AI models can be used to extract deeper triples. For example, if the research topics of authors overlap, even if they do not co-author, if multiple authors are studying "LLM", a semantic collaborative relationship can be constructed:<Zhang,ResearchOverlapWith,Liu> . This system uses graph neural networks to complete the existing set of explicit triples and predict missing edges (implicit relationships). The specific process is as follows: For each entity (author, institution, keyword), a feature vector is constructed using summary embedding (BERT vector) and explicit metrics (number of collaborations, keyword frequency). Explicit relationships such as co-authorship, affiliation, and citation are encoded as different types of edges in the graph. Some known edges are randomly hidden in the training set, and the model is trained to predict the type of hidden edges. The trained model is inferred on the entire graph, and implicit edges such as "ResearchOverlap" and "TopicSimilarity" are added for high-scoring entity pairs.

[0197] After obtaining the complete triples, the system generates a co-occurrence matrix (i.e., co-citation matrix) for the generation of the co-citation network diagram in the subsequent steps. If two papers are cited at the same time in the references of a third paper, there is a "co-citation relationship" between them, and they can be connected by an edge, and the weight of the edge is the "co-citation count". The node is the cited paper (i.e., each reference). The edge is the number of times the two documents are cited together. The co-citation network diagram can reveal the structural evolution, knowledge base, etc. within a research field. The specific construction method is as follows: First, parse the reference field (CR). In each document, the CR field records its reference list. Construct a set of sets R based on this field. i =r1, r2, ..., r n Represents the reference set of the i-th paper. Next, we construct the co-citation pairs. For any paper, its reference set R i The pairwise combinations in the R represent co-citation pairs: i , calculate all the unordered pairs: (R j , R k), where j≠k. Count the frequency of occurrence of all these pairs in the entire data set as the number of co-citations. This step constructs a co-citation statistics table: (Document A, Document B) → Number of co-citations. Finally, construct the co-occurrence matrix (co-citation matrix). Assuming that N different cited documents are extracted from all references, the system constructs an N×N matrix M: M[i][j] = the number of co-citations of two documents i and j. The matrix is ​​symmetric (the co-citation relationship has no direction). After the co-occurrence matrix is ​​constructed, it is necessary to process the data to obtain a co-citation network diagram with a clear structure. First, perform threshold filtering to retain only edges with a co-citation count greater than or equal to 2 to reduce the noise of the sparse network. Then perform row normalization to eliminate the global impact brought by popular documents (avoiding some documents from being co-cited a lot just because of their large number of citations).

[0198] 4. Graph Structure Planning Module

[0199] The function of this module is to plan the graph structure based on the one-dimensional graph data, triples, and co-occurrence matrix obtained in the previous module, so as to obtain clear and well-structured one-dimensional and two-dimensional graphs.

[0200] One-dimensional graphs visualize data for a specific dimension (such as year, keyword, author, or institution) by frequency or quantity, focusing on the following two key points: Scale inconsistency: For example, some keywords may appear hundreds of times, while others only a few. Directly plotting the data will result in long-tail data being invisible. Visual range limitations: Limited visual space prevents the complete display of all data, necessitating compression strategies.

[0201] To address these two issues, we first normalize the data. This aims to map data of varying magnitudes to the same scale (e.g., 0-1) to facilitate comparison and enhance visual contrast. Three one-dimensional graphs are generated, each using a different data normalization method based on its characteristics. The first is a year-by-paper number trend graph, normalized using Min-Max normalization:

[0202]

[0203] This normalization method is suitable for evenly distributed data and preserves the original trend.

[0204] For the keyword frequency histogram, log transformation is used for standardization:

[0205] x′ i =log b (x i +c)

[0206] This method can effectively solve the "long tail effect" and compress extreme large values. It is often used for data with strong skewed distribution, such as keywords and author publication frequency.

[0207] For the author / institution activity graph, use Z-score normalization:

[0208]

[0209] This method is suitable for processing data with outliers and converting the data into a standard normal distribution. It is more suitable when the keyword frequency changes greatly.

[0210] However, when the amount of data is too large or the difference is too large, in order to avoid the graphics being distorted or difficult to recognize, a visualization compression strategy should be adopted to improve the visibility of the graph.

[0211] For the year-paper number trend chart, sliding average smoothing is used:

[0212]

[0213] k is the size of the window. i The points on the original years are replaced by the original values, making the trend of the line chart smoother, reducing fluctuations, and making it easier to identify the turning points and rising periods of research popularity.

[0214] For the keyword frequency histogram, a color-based frequency level distinction strategy is used. Different colors, transparencies, and column widths are used for frequencies of different magnitudes, which can visually enhance the clustering effect (for example, high-frequency keywords are red, and low-frequency keywords are gray).

[0215] For the author / institution activity graph, a top-k strategy is used:

[0216] X′=Top k (X) = {x∈X|x belongs to the first k maximum values}

[0217] Only the top K items are displayed, and the rest are combined into the "Others" category. This makes the visualization clear and focuses on the main information.

[0218] For two-dimensional graphs, scientific knowledge graphs contain multiple relationships between entities (triples representation). After explicit and implicit construction is completed, the number of nodes is large and the edge density is high. The following processing is required to make the graph clear and visible.

[0219] The first step is backbone structure extraction. Edges in a graph network have varying strengths (e.g., co-occurrence frequency, co-citation count, weight, etc.). The goal of backbone extraction is to retain only statistically significant or structurally important edges, reducing edge density and enhancing structural clarity. Backbone structure extraction uses a difference filter to evaluate the significance of the edges connecting each node, retaining statistically significant edges.

[0220] Assume that node i is connected to k iedges, the weight of the edge is W ij , the sum of the weights of all connected edges is:

[0221]

[0222] Calculate the normalized weight of each edge:

[0223]

[0224] Assuming these connections are randomly distributed, the expected distribution of a single edge is:

[0225]

[0226] Retention satisfies: α ij < a, where a is the significance threshold (such as 0.05), and the smaller it is, the stricter it is.

[0227] After extracting the backbone structure, we need to embed the network into a two-dimensional space for visualization. Force-directed layout simulates the "spring force + repulsive force" in physical systems to achieve a structured layout. This avoids the undesirable situation of densely populated nodes in some parts of the scientific knowledge graph and sparse nodes in others, enhancing the graph's visibility.

[0228] Each node is considered a charged particle and has the following two forces: Attraction: connected nodes are pulled together. Repulsion: all nodes repel each other and avoid overlapping. For the connecting edge (i, j), the attraction is:

[0229]

[0230] where d ij is the Euclidean distance between nodes.

[0231] There is a repulsive force between all nodes, which is expressed as:

[0232]

[0233] Where k is the scaling ratio of the constant control chart, d ij For distance.

[0234] To prevent the graph from spreading out too much, add a central gravity:

[0235] F g (i) = -g·pos i

[0236] Where g is the gravity coefficient, pos i is the node position vector.

[0237] The iterative process of the force-directed layout algorithm is as follows: initialize the positions of all nodes; calculate the repulsive force for all nodes; calculate the attractive force for all edges; update the node positions after superimposing the force vectors; add a cooling factor (gradually reduce the movement amplitude) to converge the graph; if the iteration reaches stability or reaches the maximum number of steps, output the node coordinates.

[0238] Finally, a dimensionality reduction algorithm is used to map the high-dimensional embeddings into a two-dimensional space, preserving the local and global topological structure, which is suitable for graph visualization. Models such as t-SNE or UMAP can be used for dimensionality reduction.

[0239] For the co-citation network graph, we first construct a basic co-citation network based on the co-occurrence matrix obtained in the previous module. Each node represents a document, and the edge weight is the number of co-citations. Let G = (V, E), where V is the document set and E is the edge set. For the co-citation network, we first perform core structure extraction, extracting statistically significant edges, and then perform community segmentation.

[0240] In core structure extraction, we use a difference filter, similar to the scientific knowledge graph, to retain only the “structurally significant” edges to reflect the main connection threads:

[0241]

[0242] The method is the same as above, only the parameters need to be modified. It is worth noting that for the co-citation network, the number of retained edges should be greater than that of the scientific knowledge graph to better reflect the research relationships. Not only should the parameter selection be more conservative, but to prevent important nodes from being isolated, for each node i, the top k edges with the largest edge weights should be retained:

[0243] E′ i =TopK j∈N(i) (w ij )

[0244] Then, community segmentation is performed. The purpose of community segmentation is to identify document clusters, which usually correspond to a research topic or research path. The Louvain algorithm is applied. Louvain is a greedy community detection method whose core goal is to maximize the modularity Q. The modularity formula is as follows:

[0245]

[0246] Among them, w ij is the edge weight between nodes i and j, k i is the degree of node i, m is the total edge weight of the network, c i is the community number to which node i belongs, δ(c i , c j ) The same community is 1, and different communities are 0.

[0247] The steps for community partitioning are as follows: First, initialize each node as a community. Each time a node is moved, observe whether moving to a neighboring community improves modularity. Then, construct a new graph, using the current community as a new node, and iteratively rebuild the network. Repeat the optimization until modularity converges.

[0248] The network after core structure extraction is a sparse graph, retaining only those edges with frequent co-citations and significant structures. Community partitioning can divide the co-citation graph into multiple "research topic clusters," each of which is highly interconnected.

[0249] Finally, a dimensionality reduction algorithm is used to map the high-dimensional embedding to a two-dimensional space. Models such as t-SNE or UMAP can be used for dimensionality reduction.

[0250] 4. Graph Generation Module

[0251] The graph generation module renders the graph embedding vectors designed in the previous module into visual graphs and charts. First, for the one-dimensional graph, following the layout of the previous module, three types of graphs can be generated: a year-by-paper trend chart, a keyword frequency histogram, and an author / institution activity chart.

[0252] For the year-paper number trend graph, the graph generation scheme is as follows: the data form is a number of points (x i ,y i ), where x i Indicates the year position, y i Indicates the number of documents. Indicates the number of papers. Connect all points in chronological order and draw a continuous line. Each segment is a line segment (x i ,y i ) to (x i+1 ,y i+1 ); then draw data point circles, at each point (x i ,y i ) to facilitate annotation, and add labels to show the quantity values. Finally, draw the coordinate axes and grid lines, with the horizontal axis representing the year (X-axis) and the vertical axis representing the number of papers (Y-axis). The coordinate axis scales automatically align with the data.

[0253] For the keyword frequency histogram, the graph generation scheme is as follows: the data format is a set of coordinate points (x i , h i ). where x i is the i-th keyword position, h i is the height (frequency) of its column. Each keyword is a category, usually arranged horizontally. For each keyword, in x i Draw a rectangle at:

[0254] top=y base-h i

[0255] height=h i ,width=w

[0256] The color of the rectangle can represent different categories. The keyword name is displayed in the center below the column. The Y-axis represents the frequency, and the X-axis is the category coordinate system (discrete).

[0257] For the author / institution activity graph, the graph generation scheme is as follows: the data format is (y i , w i ) means: ordinate position y i , activity (bar length) w i .

[0258] First draw a horizontal bar, one for each author, with the starting point fixed on the left and the length w i , the bar coordinates are:

[0259] x=x0,

[0260]

[0261] width=w i ,

[0262] Then label and sort, add the author / institution name to the left or right of each bar, and sort the bars by w i Sort from largest to smallest to increase contrast.

[0263] Then for the scientific knowledge graph, after the layout of the previous step is completed, each entity e can be obtained i The two-dimensional coordinates (x i ,y i ) and the relationship triples between each pair of entities (e i ,r,e j ). First, normalize the nodes to ensure that they are distributed within a reasonable range of the canvas:

[0264]

[0265] Where W and H are the width and height of the canvas. At each position (x′ i , y′ i ) draws circular nodes. The size represents the centrality (degree). The color represents the category (author / institution / keyword), and the border represents a special role (such as a core word or a bridge node). Then for the triple (e i ,r,e j ). , in (x′ i , y′ i ) and (x′j , y′ j ). Use edge width to reflect edge strength or weight (such as co-occurrence frequency):

[0266]

[0267] Among them, w ij For node e i and e j β is the scaling factor of the scientific knowledge graph, which can be adjusted to change the maximum edge width. Finally, by adding entity names to the nodes, a clearly structured visual scientific knowledge graph can be automatically generated.

[0268] For the co-citation network diagram, after the layout design in the previous step, we can get: each document node v i The two-dimensional coordinates (x′ i , y′ i ), each edge (v i , v j ) and weight w ij , represents the number of co-citations. And the community c to which each node belongs k , that is, the cluster labels obtained by Louvain.

[0269] First, normalize the embedded coordinates to the canvas range:

[0270]

[0271] Use the normalized coordinates (x′ i , y′ i ) Place nodes and determine the size of nodes based on their degree:

[0272]

[0273] Classify labels according to community k Mapping colors, different clusters display different color blocks, showing the titles or author names of high centrality nodes (i.e. important nodes with large co-citation counts). Finally, for all edges (v i , v j ) Draw a line segment between corresponding coordinate points. As with the scientific knowledge graph, the width of the edge is determined by the weight:

[0274]

[0275] Where β is the scaling factor of the co-citation network graph, which can be adjusted to change the maximum edge width.

[0276] Finally, perform visual enhancement of the clustering by adding background color blocks (such as transparent bubbles) or dotted borders to different communities, and select the "cluster label" display: display the topic name (such as "artificial intelligence", "material chemistry", etc.) at the center of the community.

Claims

1. A knowledge graph generation method integrating visualization results of multi-source scientific and technological data, characterized in that: The following steps are involved: Step 1: Data import module: Loads raw data from local and online sources. Local sources include CSV / TSV files, relational databases, and graph databases. Online sources are collected through automated crawlers based on user-entered keywords. These keywords are then expanded into related terms, synonyms, or domain terms through intelligent search optimization. Step 2: Data preprocessing module: cleans the imported data, merges synonyms, disambiguates, and standardizes fields, including missing value processing, outlier detection, and duplicate record merging; Step 3: Information processing module: Generate one-dimensional graph data and two-dimensional graph triplets. The one-dimensional graph includes a year-number of papers trend chart, a keyword frequency histogram, and an author / institution activity chart. The two-dimensional graph constructs triplets through explicit entity extraction and implicit relationship prediction. The implicit relationships are used to complete the explicit triplets based on graph neural networks. Step 4: Graph structure planning module: Performs visual layout design based on one-dimensional graph data and two-dimensional graph triplets, including scale normalization, sliding average smoothing, and force-directed layout algorithm; Step 5: Graph generation module: Render the planned data into visual charts, including scientific knowledge graphs and co-citation network graphs, and map them to two-dimensional space through dimensionality reduction algorithms.

2. The method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data according to claim 1, characterized in that: Step 1 includes the following steps: Step 1.1: Load raw data from local sources and network sources, where the local sources include: Step 1.1.1: CSV / TSV files are used to directly import metadata data exported from Web of Science, including author, institution, title, and publication journal fields; Step 1.1.2: Extract data from a relational database, such as MySQL or PostgreSQL, by automatically constructing query statements. Step 1.1.3: Graph databases, including Neo4j, are used to store complex relationship graphs between papers and extract triples of graph data through query statements. Step 1.2: For network source data, collect it through the following steps: Step 1.2.1: Receive keywords input by the user, use the pre-trained GPT model to generate expansion words, synonyms or domain terms related to the keywords, and generate a top-n expansion keyword list; Step 1.2.2: Construct a Boolean search formula based on the extended keywords, and collect paper data in the online database platform through automatic crawler; Step 1.3: Automatically identify and classify fields in local CSV / TSV files, relational databases, and online data sources, including: Step 1.3.1: Parse the header names and content and map the fields to predefined categories, including author field AU, title field TI, institution field C1, keyword field DE, reference field CR, document type field DT, and abstract field AB; Step 1.3.2: Format the triple data extracted from the graph database into a (Subject, Predicate, Object) structure for subsequent construction of a two-dimensional graph.

3. The method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data according to claim 1, characterized in that: Step 2 includes the following steps: Step 2.1 Data cleaning: Step 2.1.1: Missing value processing, including the following two methods: Step 2.1.1.1: Remove missing values: If important fields in a record are missing and contribute little to the overall analysis, delete the record; Step 2.1.1.2: Mean / Median Filling: Fill missing values ​​in numeric fields with the mean or median of the field; Step 2.1.2: Outlier detection using the box plot method: Let the 25th percentile of the data set be Q1, the 75th percentile be Q3, calculate the interquartile range IQR = Q3-Q1, and define the outlier threshold as: Lower Bound=Q1-1.5IQR, Upper Bound=Q3+1.SIQR Values ​​outside this range were marked as outliers and extracted for subsequent analysis; Step 2.1.3: Merge duplicate records and use the hash function hashlib.blake2b library to generate a unique key: take the DOI field as the priority. If there is a conflict, combine the TI, AU, PY, and SO fields to calculate the hash value h = hash(DOI, TI, AU, PY, SO) and merge records with the same hash value. Step 2.2: Synonym merging and disambiguation: Step 2.2.1: Convert the text into vectors and use the pre-trained model to calculate word vectors; Step 2.2.2: Calculate cosine similarity Among them, v1 and v2 are word vectors; Step 2.2.3: Set a threshold θ. If CosSim(v1, v2)>θ, merge them into synonyms. Step 2.3 Field Standardization: Step 2.3.1: Unify the time format and convert the field PY to the ISO8601 format YYYY-MM-DD; Step 2.3.2: Lowercase the text and remove punctuation. Perform the following operations on the fields TI, AU, and AB: Step 2.3.2.1: Convert all letters to lowercase; Step 2.3.2.2: Remove punctuation marks “:” and “.”.

4. The method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: One-dimensional graph data generation: Step 3.1.1: Extract the year field PY or date type field DT from the preprocessed data, and calculate the mapping relationship between the year and the number of papers {y i :n i }, where y i Indicates the year, n i Indicates the number of papers published in that year, used to generate a year-paper number trend chart; Step 3.1.2: Use the word segmenter to extract high-frequency keywords from the keyword field DE and the abstract AB, and calculate the mapping relationship between keywords and the number of occurrences {k j :c j }, where k j Indicates keywords, c j Indicates the number of occurrences, used to generate keyword frequency histograms; Step 3.1.3: Extract the outlier records detected in the data cleaning module, sort them by time, and generate a graph of important nodes; Step 3.2: 2D graph triple generation: Step 3.2.1: Explicit Entity Extraction: Step 3.2.1.1: Co-authorship: Extract the author list from the author field AU and construct a co-authorship triplet by pairing them together. <AU m , CoAuthor, AU n >, the edge weight is the co-authorship frequency w mn , where AU m and AU n For the author; Step 3.2.1.2: Institutional Affiliation: Parse the mapping between author and institution from the institution field C1 and construct a triple <AU m ,AffiliatedWith,Institution p >, Institution p Indicates an institution; Step 3.2.1.3: Author-research topic relationship: Extract the association between author and keyword from the institution field DE and construct triples <AU m ,ResearchOn,k j >, k j Indicates keywords; Step 3.2.2: Implicit Relationship Prediction: Step 3.2.2.1: Encode the summary text into a vector v using the pre-trained BERT model q , where v q The semantic vector representing the summary q; Step 3.2.2.2: Construct the initial graph based on explicit triples, with node features defined by v q and explicit metrics, where explicit metrics include the frequency of cooperation w mn 、Keyword frequency c j ; Step 3.2.2.3: Randomly hide some edges in the training set, use the graph neural network to predict the type of the hidden edges, and add implicit edges of the "ResearchOverlap" or "TopicSimilarity" relationship class in the full graph inference; Step 3.3: Constructing the co-citation matrix: Step 3.3.1: Parse the reference field CR and the reference set R of each paper i ={r1, r2, ..., r n }, calculate all unordered pairs (R j , R k ) to generate the co-occurrence matrix M, where M[i][j] represents the co-citation counts of documents i and j; Step 3.3.2: Parse the reference field CR and the reference set R of each paper i ={r1, r2, ..., r n }, count all unordered pairs (r j , r k )’s co-citation count; Step 3.3.3: Construct the co-citation matrix M N×N , where N is the total number of references, and the matrix element M[j][k] represents the number of references r j and r k The number of co-citations; Step 3.3.4: Threshold filter the co-occurrence matrix, retain edges with M[i][j] ≥ 2, and normalize the matrix rows to eliminate the popular literature bias: The normalization formula is: Where N is the total number of documents, M norm [i][j] represents the normalized co-citation weights.

5. The method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data according to claim 1, characterized in that: Step 4 includes the following steps Step 4.1: One-dimensional graph structure planning: Step 4.1.1: Use Min-Max normalization for the year-paper number trend graph: Among them, x i represents the original value of the number of papers in the i-th year, x i is the normalized value, max(x) and min(x) are the number of papers in the largest and smallest years in the dataset, respectively; Step 4.1.2: Logarithmic transformation of keyword frequency histogram: x′ i =log b (x i +c) Among them, x i is the number of times the keyword appears, x i is the transformed value, b is the logarithmic base, and c is the smoothing constant; Step 4.1.3: Z-score normalization of the author / institution activity graph: Among them, x i is the activity of the author / institution (such as the number of publications), μ is the mean, and σ is the standard deviation; Step 4.1.4: Sliding average smoothing of the year-paper number trend graph, using the sliding mean formula with a window size of k: Among them, s i is the smoothed value of the i-th year, x j is the original data point; Step 4.1.5: Visual compression strategy of keyword frequency histogram, using color to distinguish frequency levels: according to x i The value is divided into intervals and mapped to different colors, transparency or column widths; Step 4.1.6: Apply the top-k strategy to the author / institution activity graph, filter the top k largest values ​​and classify them as "others"; Step 4.2 Two-dimensional graph structure planning: Step 4.2.1: Extract the backbone structure of the scientific knowledge graph using the difference filter formula: Among them, k i is the number of edges connecting node i, S i is the sum of edge weights of node i, α ij Indicates the significance of edge (i, j), retaining the edge that satisfies α ii <a(, a is the significance threshold; Step 4.2.2: Force-directed layout algorithm, calculate the attractive and repulsive forces between nodes: Step 4.2.2.1: Initialize node positions: Assign initial two-dimensional coordinates (x i ,y i ), where x i and y i Represents the horizontal and vertical coordinates of the i-th node; Step 4.2.2.2 Calculate the repulsive force: For all node pairs (i, j), calculate the repulsive force according to the formula: Where k is the constant control chart scaling ratio; Step 4.2.2.3 Calculate the attraction: For all connected edges (i, j), calculate the attraction according to the formula; where d ij is the Euclidean distance between nodes, (i, j) represents the connecting edge; Step 4.2.2.4 Superimpose the force vectors and update the position: Add the resultant force F of the repulsive force and the attractive force total (i)=∑ j (F r (i,j)+F a (i,j)) acts on node i and updates its coordinates: x′ i =x i +Δt·F total (i) x y′ i =y i +Δt·F total (i) y Where Δt is the time step; Step 4.2.2.5 Add cooling factor: In each iteration, add cooling factor in proportion to α t =1-t / T max The number of attenuation movements is: Δt′=αt·Δt T max is the maximum number of iterations; Step 4.2.2.6 Termination condition judgment: The iteration is terminated if any of the following conditions are met: The number of iterations reaches the preset maximum value T max ; The node position change is less than the threshold; Step 4.2.2.7 Output node coordinates: Output the final converged two-dimensional coordinates (x′ i , y′ i ) as the visual layout result; Step 4.2.3: Dimensionality reduction and mapping to two-dimensional space: Use t-SNE or UMAP algorithm to map the high-dimensional embedding vector to two-dimensional coordinates (x′ i , y′ i ), preserving local / global topology; Step 4.2.4: Extract the core structure of the co-citation network graph using the same difference filter formula as step 4.2.1 and retain the top-k largest weight edges of each node; Step 4.2.5: Divide the co-citation network graph into communities using the Louvain algorithm, which includes the following sub-steps: Step 4.2.5.1: Initialize the community: assign each node to an independent community, that is, each node is initially an independent community; Step 4.2.5.2: Node movement optimization: traverse all nodes and move the current node to the community to which its neighboring nodes belong. If the modularity Q is improved, keep this move; Step 4.2.5.3: Network reconstruction: Use the currently divided communities as new nodes to rebuild the network graph. The edge weights between the new nodes are the sum of all edge weights between the original communities. Step 4.2.5.4: Iterative optimization: Repeat steps 4.2.5.2 to 4.2.5.3 until the modularity Q reaches the convergence condition; The calculation formula of modularity Q is: Among them, w ij is the edge weight between nodes i and j, indicating the number of co-citations, k i is the degree of node i, is the total edge weight of the network, δ(c i , c j ) is the indicator function, if c i =c j Then take 1, otherwise 0, c i is the community number to which node i belongs.

6. The method for generating a knowledge graph by integrating visualization results of multi-source scientific and technological data according to claim 1, characterized in that: Step 5 includes the following sub-steps: Step 5.1 One-dimensional graph generation module: Step 5.1.1 Generate a trend chart of the number of papers by year: Set the year field y i and the corresponding number of papers n i Mapped to a set of data points {(x i ,y i )}, where: x i Indicates the year position, y i Indicates the number of papers in that year. Draw a continuous line connecting all the points. Each data point (x i ,y i ) and mark the number of circles with fixed radius; the horizontal axis is the year X, and the vertical axis is the number of papers Y; Step 5.1.2 Generate keyword frequency histogram: The keyword k j and the number of occurrences c j Mapped to a set of coordinate points {(x j , h j )),in: x j Indicates the position of the jth keyword, h j Indicates the frequency of the keyword, for each keyword in position x j Draw a rectangle at , the parameters are: Left border: Top edge: top = y base -h i , Width: width=w, Height: h j The keyword name is displayed in the center below the rectangle, the Y-axis represents the frequency, and the X-axis is the discrete classification coordinate system; Step 5.1.3 Generate author / institution activity graph: Author or institution A m Activity level w m Mapped to a set of bar data points {(y m , w m )},in: y m Indicates the vertical coordinate position of the author / institution, w m Indicates activity, draws a horizontal bar with the starting point fixed on the left and a length of w m , the coordinates are: Horizontal axis position: x = x0, Vertical axis: Bar width: w m , Bar Press w m Sort from largest to smallest, with the author / institution name marked on the left or right; Step 5.2: 2D graph generation module: Step 5.2.1: Scientific knowledge graph generation: Step 5.2.1.1: For each entity e i The two-dimensional coordinates (x′ i , y′ i ) is normalized, and the formula is: in: x i ,y i is the original coordinate, W and H are the width and height of the canvas. Step 5.2.1.2: Draw circular nodes: size represents centrality, color represents category, category includes author, institution, keyword, that is, degree size i Map to the diameter of the drawn circle: where d i is the node degree, R min , r max is the minimum / maximum node size; Step 5.2.1.3: For the triple (e i ,r,e j ), at the coordinate (x′ i , y′ i ) and (x′ j , y′ j ) with edge widths determined by weights: where w ij is the edge weight, β is the scaling factor; Step 5.2.1.4: Add entity names to the nodes to generate a clearly structured scientific knowledge graph; Step 5.2.2: Generate co-citation network diagram: Step 5.2.2.1: For each document node v i The two-dimensional coordinates (x′ i , y′ i ) for normalization; The node size is determined by the degree, and the formula is: where d i is the node degree, r min , r max is the minimum / maximum node size; Step 5.2.2.2: Classify labels c according to community k Mapping colors, different clusters display different color blocks; Step 5.2.2.3: Opposite side (v i , v j ) Draw a line segment with edge width determined by weight; Step 5.2.2.4: Add background color blocks or dotted borders to different communities and display cluster labels at the centroid.

7. A knowledge graph generation system integrating visualization results of multi-source scientific and technological data, characterized by: The system is stored in a storage medium in a program manner. When the processor executes the program, it implements a knowledge graph generation method based on the fusion of multi-source data visualization results as described in any one of claims 1-6.

Citation Information

Cited By

  • Standard reference relationship identification system and method based on gravitation model

    CN121479716A