Large-scale Knowledge Graph Quasi-real-time Construction Method, Device, Equipment and Storage Medium
By designing point, edge, and attribute data standards, static and dynamic data are processed, static and dynamic point-edge hive tables are generated, and incremental data is updated in real time, which solves the problems of data inconsistency and low processing efficiency in large-scale relationship map construction, and realizes efficient data fusion and real-time requirements.
Patent Information
- Application Number
- CN202510561844.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing technology has problems such as data inconsistency, low processing efficiency and difficulty in data governance in the construction of large-scale relationship maps. Especially in the process of multi-source data fusion, it is difficult to achieve real-time and efficient data requirements.
The point, edge and attribute data standards are designed, and the static and dynamic point-edge hive tables are generated by processing static and dynamic data separately, and the incremental data is processed in real time through serial task arrangement, and the associated data is automatically updated, and the data is finally imported into the hugegraph graph database to build a complete graph model data.
It significantly shortens the time for large-scale data construction, meets real-time needs, supports the continuous growth of data scale, stable performance, reduces data management and maintenance costs, and realizes the timely construction of various data sources and the integration of similar data.
Smart Images

Figure CN120086387B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large-scale relational graph construction, and particularly to a method, apparatus, device, and storage medium for quasi-real-time construction of a large-scale relational graph. Background Art
[0002] With the explosive growth of data, knowledge graphs, as semantic networks revealing relationships between entities, provide a more effective way for the expression, organization, management, and utilization of massive, heterogeneous, and dynamic big data on the Internet, making the network more intelligent and closer to human cognitive thinking. The applications of knowledge graphs are increasingly widespread, and efficient knowledge graph construction and management technologies are indispensable in fields such as social networks, financial risk control, and intelligent recommendation.
[0003] As the data type and scale continue to increase, traditional graph construction methods are inefficient in ultra-large-scale knowledge graphs, and data governance methods and graph construction methods need to be optimized. The present invention proposes a method for quasi-real-time construction of a relational graph with a scale of hundreds of billions based on multi-source data fusion. This method models and governs multi-source data through a data modeling platform to ensure data consistency and integrity, and improve the efficiency, accuracy, and real-time performance of graph construction.
[0004] Existing ultra-large-scale graph construction mainly relies on modifying graph databases or retrieval methods. Document 1 (a Chinese patent application for invention with the application number CN202210874965.8) proposes an intelligent indexing method and system for ultra-large-scale knowledge graphs based on deep learning intelligent hashing. High-efficiency indexing is achieved through a deep learning model to improve retrieval efficiency. The specific steps include dividing the index input into entity, relationship triples, and attribute triples, encoding the input using a BERT-compatible model, and regressing the starting position of data storage and the length of physical storage through an aggregation network and a multi-layer perceptron. It is applicable to the storage of large-scale semantic knowledge graphs, improves retrieval efficiency, and realizes simple retrieval, complex multi-hop retrieval, and complex analysis with knowledge reasoning. Document 2 (a Chinese patent application for invention with the application number CN202110677218.0) proposes a distributed cluster framework for graph databases based on docker-compose technology, as well as a method for joint storage and retrieval of graph databases and document databases. Through distributed storage, indexing, and calculation, the efficiency of knowledge graph construction and retrieval under the background of large-scale massive data is improved. It effectively solves the problem that a single-machine graph database cannot meet the storage and retrieval requirements of massive data, and first proposes a plug-and-play distributed cluster framework for graph databases. The proposed heterogeneous database storage and retrieval solution significantly improves the overall retrieval efficiency and alleviates the problem of reduced retrieval performance of super nodes in graph databases.
[0005] When constructing a relationship graph, one is for a single data source, and the other is that the data volume is small, and it is basically impossible to process ultra-large-scale data regularly. Document 3 (a Chinese invention patent application with the application number CN202410252439.7) provides a method for constructing and optimizing a person relationship graph based on multi-source data fusion. The steps include real-time collecting and generating basic person information, preprocessing the basic person information, using machine learning to identify and extract the preprocessed basic person information to obtain person relationship data, constructing a person relationship graph based on the person relationship data, and optimizing and updating the person relationship graph based on the real-time obtained basic person information. Integrate multiple data sources to provide comprehensive person relationship information and enhance the accuracy and reliability of the graph. Through data preprocessing, machine learning, and knowledge graph technology, improve the efficiency and quality of constructing a person relationship graph. Document 4 (a Chinese invention patent application with the application number: CN202111268566.9) provides a method for constructing an ultra-large-scale spectrum knowledge graph for a 6G communication network. It includes spectrum knowledge graph ontology construction, spectrum data acquisition, spectrum knowledge extraction, spectrum knowledge fusion, spectrum knowledge quality assessment, spectrum knowledge graph construction and storage, spectrum knowledge reasoning, and spectrum knowledge graph update. From the perspective of wireless communication, it provides a framework for the core technology to achieve intelligent spectrum resource management. The constructed spectrum knowledge graph has high quality and strong logic between nodes, and can perform efficient spectrum knowledge query and utilization.
[0006] The existing technologies have the following deficiencies:
[0007] Data inconsistency: In the process of multi-source data fusion, the data formats and qualities of different data sources vary greatly, resulting in data inconsistency and affecting the accuracy and reliability of the graph.
[0008] Low processing efficiency: Processing large-scale data requires a large amount of computing resources and time. There are bottlenecks in the existing technologies in terms of processing efficiency and cannot meet the requirements of real-time and high efficiency.
[0009] Difficult data governance: Governing and maintaining large-scale data requires complex technical means. There are deficiencies in the existing technologies in data governance, resulting in high costs for data management and maintenance.
[0010] In the process of constructing a paper citation data - relationship graph, the data sources are obtained from different platforms such as VIP, Wanfang, and CNKI, and the data formats and contents are different. How to achieve both constructing and generating respective graph data for various data sources in a timely manner and fusing the same type of data in different data sources is a technical problem. Summary of the Invention
[0011] Based on this, it is necessary to provide a method, apparatus, device, and storage medium for quasi-real-time construction of a large-scale relational graph in response to the above technical problems.
[0012] A method for quasi-real-time construction of a large-scale relational graph, the method comprising:
[0013] Design point, edge, and attribute data standards.
[0014] For static data, according to the point, edge, and attribute data standards, construct corresponding point models, edge models, and point-edge models for the paper citation data to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables.
[0015] For dynamic data, model the stock data according to the point, edge, and attribute data standards, and perform real-time processing on the incremental data to form a dynamic point-edge hive table.
[0016] Through serial task scheduling, run the incremental data governance model at regular intervals to process the incremental data and automatically update the associated data.
[0017] At a preset moment, import the static point-edge hive table and the dynamic point-edge hive table into the HugeGraph graph database to obtain citation graph data.
[0018] Extract edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merge the citation relationship table and the paper information entity table to form complete graph model data.
[0019] In one embodiment, designing the point, edge, and attribute data standards includes: defining a corresponding number of entity points and a number of point-edge relationships for multiple data sources, setting the storage total to the TB level, setting the number of entity points to tens of billions, and setting the number of relationship edges to hundreds of billions.
[0020] The identification of points is represented in the form of an identification prefix + value; point attributes include point identification, point type, display value, update time, and partition field; edge attributes include subcategory, start point, end point, update time, remarks, and partition field.
[0021] In one embodiment, the specific process of extracting edge information from the paper citation data includes:
[0022] Filter the paper citation data in the data modeling platform to extract the data of the previous day.
[0023] Filter the data of the previous day to obtain the valid citing paper IDs.
[0024] The cited party's paper IDs are split from one piece of data into multiple pieces by SQL operators, with each paper ID being one piece of data.
[0025] The same citing party's paper ID and cited party's paper ID on the same day are aggregated and counted by a data aggregation operator.
[0026] The aggregation result is processed by a table structure processing operator to generate a citation relationship table with the fine category being citation, the starting point being the citing party's paper ID, the ending point being the cited party's paper ID, and the update time being the citation date.
[0027] In one embodiment, the specific process of extracting point information from paper citation data includes:
[0028] In the data modeling platform, data filtering is performed on the paper information data to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information.
[0029] The paper information of the previous day is de-duplicated by paper ID through a data de-duplication operator.
[0030] For the removal result, a field addition operator is used to add 'paper_' in front of the paper ID to construct a point identifier, the point type is paper, and the display value is the paper title.
[0031] An all-merge operator is used to merge the point identifier and the paper information entity table formed historically.
[0032] For the merge result, a de-duplication operator is used to de-duplicate the point identifier and the display value, and data filtering is performed on the de-duplication result to extract the data of the previous day to obtain the paper information entity table.
[0033] In one embodiment, the extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; merging the citation relationship table and the paper information entity table includes:
[0034] In the data modeling platform, data filtering is performed on the citation relationship table to extract the citation relationship information of the previous day.
[0035] The table structure of the citation relationship information of the previous day is processed to obtain the citing party's paper ID and the cited party's paper ID, which are renamed as point identifiers.
[0036] After processing the citing party's paper ID and the cited party's paper ID with an all-merge operator, a data set operator is used to aggregate the paper point identifier and the update time to obtain the daily paper address entity table.
[0037] For the daily paper address entity table, a field addition operator is used to set the sorting field to 1, and the paper ID field is renamed as the display value.
[0038] After processing the paper information entity table through data filtering and adding field operators, extract the data from the previous day and set the sorting field to 3.
[0039] After processing the paper graph entity table formed historically through data filtering and adding field operators, extract the data from the previous day and set the sorting field to 2.
[0040] For the daily paper address entity table with the sorting field set, and the data from the previous day extracted from the paper information entity table and the paper graph entity table formed historically, perform a full merge using the full merge operator.
[0041] For the obtained full merge result, perform deduplication on the point identifiers using the data deduplication operator, sort according to the sorting field, and then filter out the data where the sorting field is not 2 using the data filtering operator to obtain the paper graph entity table.
[0042] In one embodiment, for dynamic data, model the stock data according to the point, edge, and attribute data standards, perform real-time processing on the incremental data, and form a dynamic point-edge hive table, including:
[0043] For dynamic data, model the stock data according to the point, edge, and attribute data standards, and obtain all the stock data through the model; build a data governance model according to the partition field to process the incremental data and form different types of dynamic point-edge hive tables.
[0044] In one embodiment, the method further includes: optimizing the data governance model through machine learning and deep learning methods.
[0045] A quasi-real-time construction device for a large-scale relational graph, the device includes:
[0046] A relational graph standard design module for designing point, edge, and attribute data standards.
[0047] A construction module for graph data, for static data, according to the point, edge, and attribute data standards, construct corresponding point models, edge models, and point-edge models for paper citation data to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables; for dynamic data, model the stock data according to the point, edge, and attribute data standards, perform real-time processing on the incremental data, and form a dynamic point-edge hive table; through serial task scheduling, run the incremental data governance model regularly to process the incremental data and automatically update the associated data; at a preset moment, import the static point-edge hive table and the dynamic point-edge hive table into the hugegraph graph database to obtain citation graph data.
[0048] The data governance model construction module is used to extract edge information and vertex information from the paper citation data, and merge the edge information and vertex information to form complete graph model data.
[0049] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0050] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the above method are implemented.
[0051] The above large-scale relational graph quasi-real-time construction method, device, equipment and storage medium. The method includes: designing data standards for vertices, edges, and attributes; constructing static vertex-edge Hive tables and dynamic vertex-edge Hive tables through the separate processing of static data and dynamic data, and the real-time processing of incremental data; importing the static vertex-edge Hive tables and dynamic vertex-edge Hive tables into the HugeGraph graph database at a preset moment to obtain citation graph data; extracting edge information and vertex information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merging the citation relationship table and the paper information entity table to form complete graph model data. This method significantly shortens the construction time when processing data on a scale of hundreds of billions, meets the real-time requirements; supports the continuous growth of the data scale and has stable performance; can not only construct and generate respective graph data for various data sources in a timely manner, but also fuse the same type of data in different data sources. Description of the Drawings
[0052] Figure 1 It is a schematic flowchart of the large-scale relational graph quasi-real-time construction method in an embodiment;
[0053] Figure 2 It is a flowchart of edge extraction from citation data in another embodiment;
[0054] Figure 3 It is a flowchart of vertex extraction from citation data in another embodiment;
[0055] Figure 4 It is a flowchart of edge-to-vertex merging of citation data in another embodiment;
[0056] Figure 5 It is a structural block diagram of the large-scale relational graph quasi-real-time construction device in an embodiment;
[0057] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] In one embodiment, Figure 1 As shown, a method for constructing a large-scale relationship graph in quasi-real time is provided, and the method comprises the following steps:
[0060] Step 100: Design point, edge, and attribute data standards.
[0061] Step 102: For static data, construct corresponding point models, edge models, and point-edge models for paper citation data according to point, edge, and attribute data standards to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables.
[0062] Specifically, for static data: through different data sources, build specific point models, edge models, and point-edge models to process the data, obtain static point and edge data, and form different types of static point-edge hive tables.
[0063] Among them, each type of point or edge has multiple different data sources, such as retrieval platforms such as VIP, Wanfang, and CNKI, which should be considered to be merged together when constructing graph data.
[0064] For paper citation data, the data source is the paper and citation information obtained from different databases such as VIP, Wanfang, and CNKI. Each point in the graph is a paper, and the point model is to build models for different databases to form point (paper) data. Each edge in the graph represents the citation relationship between papers, and the edge model is to build models for different databases to form edge (citation relationship) data. The point-edge model is to model the same type of points (or edges), and merge the model results built by different databases to form a type of point (or edge) model.
[0065] Among them, static data refers to data that is no longer updated, such as paper citation data purchased / acquired once.
[0066] Step 104: For dynamic data, the stock data is modeled according to the point, edge, and attribute data standards, and the incremental data is processed in real time to form a dynamic point-edge hive table.
[0067] Specifically, for the existing data, all the existing data is obtained by constructing a data governance model to ensure the integrity and consistency of the initial data. For the incremental data, a data governance model is constructed according to the partition fields to process the incremental data, ensuring the timeliness and accuracy of the incremental data. Finally, different types of dynamic vertex-edge Hive tables are formed.
[0068] Dynamic data refers to data sources where the data is constantly being updated, such as current mainstream paper retrieval databases. The citation data in such data sources has both historical accumulation and is increasing every day. Existing data refers to all the data from this moment back in time. For the existing data, a one-time modeling is performed to form vertex and edge data.
[0069] Incremental data refers to the new paper citation data obtained every day. For the incremental data, we construct an incremental model that runs once a day at a scheduled time. Each time it runs, it takes the difference set with the existing data to form the incremental vertex and edge data for the day.
[0070] By separately processing the static data and dynamic data, and the real-time processing of the incremental data, the efficiency and timeliness of graph construction are improved.
[0071] Step 106: Through serial task scheduling, the incremental data governance model is run at a scheduled time to process the incremental data and automatically update the associated data.
[0072] Specifically, a data governance model is constructed for the incremental data. Through serial task scheduling, the constructed incremental data governance model is run at a scheduled time to process the incremental data and automatically update the associated data, ensuring the consistency and integrity of the data.
[0073] All models are data governance models. The incremental data governance model needs to run daily at a scheduled time. When there are multiple incremental data governance models, task scheduling should be carried out in a certain order to ensure that each data governance model can run once a day. Some of these models have dependencies, so when performing serial task scheduling, the order of running the data governance models needs to be designed well.
[0074] Step 108: At a preset moment, the static vertex-edge Hive table and the dynamic vertex-edge Hive table are imported into the HugeGraph graph database to obtain the citation graph data.
[0075] Specifically, the well-governed vertex-edge Hive table is imported into the HugeGraph graph database daily at a scheduled time through a written script program.
[0076] Step 110: Extract the edge information and vertex information from the paper citation data to obtain the citation relationship table and the paper information entity table, and merge the citation relationship table and the paper information entity table to form the complete graph model data.
[0077] Specifically, a multi-source data modeling and governance is carried out through a data modeling platform to ensure data consistency and integrity, and improve the accuracy and reliability of the graph. The data governance process is simplified through automated and visual data governance tools, and the data management and maintenance costs are reduced.
[0078] In the above method for quasi-real-time construction of a large-scale relational graph, the method includes: designing data standards for points, edges, and attributes; constructing static point-edge Hive tables and dynamic point-edge Hive tables through the separate processing of static data and dynamic data, as well as the real-time processing of incremental data; importing the static point-edge Hive tables and dynamic point-edge Hive tables into the HugeGraph graph database at a preset moment to obtain citation graph data; extracting edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merging the citation relationship table and the paper information entity table to form complete graph model data. This method significantly shortens the construction time when processing data on a scale of hundreds of billions, meets the real-time requirements; supports the continuous growth of the data scale with stable performance; through modular design and automated management, the system maintenance cost is significantly reduced, and the user operation is simple. It can not only construct and generate their respective graph data from various data sources in a timely manner, but also fuse the same type of data from different data sources.
[0079] In one embodiment, step 100 includes: defining a corresponding number of entity points and a number of point-edge relationships for multiple data sources, setting the total storage capacity to the TB level, setting the number of entity points to tens of billions, and setting the number of relationship edges to hundreds of billions; representing the identifier of a point in the form of an identifier prefix + value; the point attributes include point identifier, point type, display value, update time, and partition field; the edge attributes include sub-category, starting point, ending point, update time, remarks, and partition field.
[0080] In one embodiment, the specific process of extracting edge information from the paper citation data in step 110 includes: filtering the paper citation data in the data modeling platform to extract the data of the previous day; filtering the data of the previous day to obtain the valid citing paper IDs; splitting the cited paper IDs from one piece of data into multiple pieces through SQL operators, with each paper ID being a piece of data; aggregating and counting the same citing paper IDs and cited paper IDs on the same day through a data aggregation operator; processing the aggregation result using a table structure processing operator to generate a citation relationship table with the sub-category being citation, the starting point being the citing paper ID, the ending point being the cited paper ID, and the update time being the citation date. The edge extraction process of the citation data is as Figure 2 shown.
[0081] In one embodiment, the specific process of extracting point information from the paper citation data in step 110 includes: filtering the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; de-duplicating the paper information of the previous day by the paper ID using the data de-duplication operator; using the add field operator to add 'paper_' in front of the paper ID to construct a point identifier for the removal result, the point type is a paper, and the display value is the paper title; using the all merge operator to merge the point identifier and the paper information entity table formed by history; using the de-duplication operator to de-duplicate the point identifier and the display value for the merge result, and performing data filtering on the de-duplication result to extract the data of the previous day to obtain the paper information entity table. The citation data point extraction process is as Figure 3 shown.
[0082] In one embodiment, the extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; merging the citation relationship table and the paper information entity table in step 110 includes: filtering the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day; performing table structure processing on the citation relationship information of the previous day to obtain the citing paper ID and the cited paper ID, and renaming them as point identifiers; using the all merge operator to process the citing paper ID and the cited paper ID, and then using the data set operator to aggregate the paper point identifiers and the update time to obtain the daily paper address entity table; using the add field operator for the daily paper address entity table to set the sorting field to 1 and rename the paper ID field as the display value; after processing the paper information entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 3; after processing the paper graph entity table formed by history through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 2; using the all merge operator to perform an all merge on the daily paper address entity table with the sorting field set, and the data of the previous day extracted from the paper information entity table and the paper graph entity table formed by history; using the data de-duplication operator to de-duplicate the point identifier for the obtained all merge result, sorting according to the sorting field, and then filtering out the data with the sorting field not equal to 2 through the data filtering operator to obtain the paper graph entity table. The citation data edge generating point merging process is as Figure 4 shown.
[0083] Specifically, taking the citation data as an example, edge extraction, point extraction, and edge generating point merging through the data modeling platform are introduced respectively.
[0084] (1)Edge extraction: For citation data, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the citation data of the previous day; through the data filtering operator, extract the valid citing party paper IDs; through the SQL operator, split the cited party paper ID from one piece of data into multiple, with each paper ID being a piece of data; through the data aggregation operator, perform aggregation and counting on the same citing party paper ID and cited party paper ID on the same day; through the table structure processing operator, finally form a citation relationship table with the fine category being citation, the starting point being the citing party paper ID, the ending point being the cited party paper ID, and the update time being the citation date.
[0085] (2)Node extraction: For paper information data, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the paper information of the previous day, including the paper ID and the corresponding paper information; through the data deduplication operator, deduplicate by paper ID; through the add field operator, add 'paper_' in front of the paper ID to construct a node identifier, with the node type being paper and the display value being the paper title; through the all merge operator, merge with the paper information entity table formed historically; through the data deduplication operator, aggregate the node identifier and the display value; through the data filtering operator, take the data of the previous day to obtain the paper information entity table.
[0086] (3)Edge-node merging: For the citation relationship table formed in the first step, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the citation relationship information of the previous day; through table structure processing, obtain the citing party paper ID and the cited party paper ID, and rename them as node identifiers; through the all merge operator and the data aggregation operator, aggregate the paper node identifiers and the update time to form a daily paper address entity table; through the add field operator, set the sorting field to 1 and rename the paper ID field as the display value; through the data filtering and add field operators, respectively extract the data of the previous day from the paper information entity table formed in step two and set the sorting field to 3, and extract the data of the previous day from the historically formed paper graph entity table and set the sorting field to 2; through the all merge operator, merge the three types of data; through the data deduplication operator, deduplicate the node identifiers; sort by the sorting field; through the data filtering operator, filter out the data with the sorting field not equal to 2, and finally obtain the paper graph entity table.
[0087] In one embodiment, step 104 includes: For dynamic data, model the stock data according to the node, edge, and attribute data standards, and obtain all the stock data through the model; construct a data governance model according to the partition field for incremental data processing to form different types of dynamic node-edge hive tables.
[0088] In one embodiment, the method further includes: optimizing the data governance model through machine learning and deep learning methods to improve the intelligent level of data processing.
[0089] The method can expand data sources: it can support more data source types, further improving the diversity and richness of data.
[0090] The relationship graph constructed by using this method can be used in fields such as financial risk control, intelligent recommendation, and social network analysis. Specifically:
[0091] Financial risk control: used for risk control in the financial field, discovering potential risk points through graph analysis.
[0092] Intelligent recommendation: used for building a recommendation system, providing personalized recommendation services through graph analysis.
[0093] Social network analysis: used for analyzing social networks, discovering key nodes and community structures in social networks through graph analysis.
[0094] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0095] In one embodiment, as Figure 5 shown, a large-scale relationship graph quasi-real-time construction device is provided, including: a relationship graph standard design module, a graph data construction module, and a data governance model construction module. Among them:
[0096] The relationship graph standard design module is used to design the standards for point, edge, and attribute data.
[0097] The graph data construction module is used for static data. According to the point, edge, and attribute data standards, it constructs corresponding point models, edge models, and point-edge models for the paper citation data to process the static data, obtaining static point and edge data, and generating different types of static point-edge Hive tables. For dynamic data, it models the stock data according to the point, edge, and attribute data standards, and processes the incremental data in real time to form dynamic point-edge Hive tables. Through serial task scheduling, it runs the incremental data governance model regularly to process the incremental data and automatically update the associated data. At a preset moment, it imports the static point-edge Hive tables and dynamic point-edge Hive tables into the HugeGraph graph database to obtain citation graph data.
[0098] The data governance model construction module is used to extract edge information and point information from the paper citation graph data, and merge the edge information and point information to form complete graph model data.
[0099] In one embodiment, the relationship graph standard design module is further used to define a corresponding number of entity points and a number of point-edge relationships for multiple data sources, set the storage total to the TB level, set the number of entity points to tens of billions, and set the number of relationship edges to hundreds of billions. The identification of points is represented in the form of an identification prefix + value. The point attributes include point identification, point type, display value, update time, and partition field. The edge attributes include subcategory, start point, end point, update time, remarks, and partition field.
[0100] In one embodiment, the specific process of extracting edge information from the paper citation data in the data governance model construction module includes: filtering the paper citation data in the data modeling platform to extract the data of the previous day; filtering the data of the previous day to obtain the valid cited party paper IDs; using SQL operators to split the cited party paper IDs from one piece of data into multiple pieces, with each paper ID being a piece of data; using data aggregation operators to aggregate and count the same citing party paper IDs and cited party paper IDs on the same day; processing the aggregation result using a table structure processing operator to generate a citation relationship table with the subcategory being citation, the start point being the citing party paper ID, the end point being the cited party paper ID, and the update time being the citation date.
[0101] In one embodiment, the specific process of extracting point information from the paper citation data in the data governance model construction module includes: filtering the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; de-duplicating the paper information of the previous day by the paper ID through the data de-duplication operator; using the add field operator on the removal result to add 'paper_' in front of the paper ID to construct a point identifier, the point type is a paper, and the display value is the paper title; using the all merge operator to merge the point identifier and the paper information entity table formed historically; using the de-duplication operator on the merge result to de-duplicate the point identifier and the display value, and filtering the de-duplication result to extract the data of the previous day to obtain the paper information entity table.
[0102] In one embodiment, the extracted edge information is the citation relationship table, and the extracted point information is the paper information entity table; the merging of the citation relationship table and the paper information entity table in the data governance model construction module includes: filtering the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day; performing table structure processing on the citation relationship information of the previous day to obtain the citing paper ID and the cited paper ID, and renaming them as point identifiers; using the all merge operator on the citing paper ID and the cited paper ID, and then using the data set operator to aggregate the paper point identifiers and the update time to obtain the daily paper address entity table; using the add field operator on the daily paper address entity table to set the sorting field to 1 and rename the paper ID field as the display value; after processing the paper information entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 3; after processing the historically formed paper graph entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 2; using the all merge operator to perform all merges on the daily paper address entity table with the sorting field set, and the data of the previous day extracted from the paper information entity table and the historically formed paper graph entity table; using the data de-duplication operator to de-duplicate the point identifiers on the obtained all merge result, sorting according to the sorting field, and then filtering out the data with the sorting field not equal to 2 through the data filtering operator to obtain the paper graph entity table.
[0103] In one embodiment, the graph data construction module is further configured to, for dynamic data, model the stock data according to the point, edge, and attribute data standards, and obtain all the stock data through the model; construct a data governance model according to the partition field to process the incremental data, and form different types of dynamic point-edge hive tables.
[0104] In one embodiment, the device further includes an optimization module for optimizing the data governance model through machine learning and deep learning methods.
[0105] For the specific limitations of the large-scale relational graph quasi-real-time construction device, reference can be made to the limitations of the large-scale relational graph quasi-real-time construction method in the foregoing text, which will not be elaborated herein. Each module in the above large-scale relational graph quasi-real-time construction device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0106] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a large-scale relational graph quasi-real-time construction method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0107] Those skilled in the art can understand that Figure 6 the structure shown in
[0108] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0109] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above method embodiment.
[0110] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0111] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0112] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for constructing a large-scale relational graph in near real-time, characterized in that The method includes: Designing data standards for points, edges, and attributes; For static data, according to the data standards for points, edges, and attributes, construct corresponding point models, edge models, and point-edge models for the paper citation data to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables; For dynamic data, according to the data standards for points, edges, and attributes, model the stock data and perform real-time processing on the incremental data to form a dynamic point-edge hive table; Through serial task scheduling, run the incremental data governance model at regular intervals to process the incremental data and automatically update the associated data; At a preset moment, import the static point-edge hive table and the dynamic point-edge hive table into the HugeGraph graph database to obtain citation graph data; Extract edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merge the citation relationship table and the paper information entity table to form complete graph model data.
2. The method for constructing a large-scale relational graph in near real-time according to claim 1, wherein Designing data standards for points, edges, and attributes, including: Defining corresponding several entity points and several point-edge relationships for multiple data sources, setting the storage total to the TB level, setting the number of entity points to tens of billions, and setting the number of relationship edges to hundreds of billions; Representing the identifier of a point in the form of an identifier prefix + value; point attributes include point identifier, point type, display value, update time, and partition field; edge attributes include subcategory, start point, end point, update time, remarks, and partition field.
3. The method for constructing a large-scale relational graph in near real time according to claim 1, wherein The specific process of extracting edge information from the paper citation data includes: Filter the paper citation data in the data modeling platform to extract the data of the previous day; Filter the data of the previous day to obtain the valid citing party paper IDs; Use SQL operators to split the cited party paper ID from one piece of data into multiple pieces, with each paper ID being one piece of data; Use data aggregation operators to aggregate and count the same citing party paper ID and cited party paper ID on the same day; Process the aggregation result using a table structure processing operator to generate a citation relationship table with the subcategory being citation, the start point being the citing party paper ID, the end point being the cited party paper ID, and the update time being the citation date.
4. The method for constructing a large-scale relational graph in near real time according to claim 1, wherein The specific process of extracting point information from the paper citation data includes: Filter the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; Deduplicate the paper information of the previous day by paper ID using a data deduplication operator; Use an add field operator to add 'paper_' in front of the paper ID to construct a point identifier for the removal result, with the point type being paper and the display value being the paper title; Use an all merge operator to merge the point identifier and the paper information entity table formed by history; Use a deduplication operator to deduplicate the point identifier and the display value for the merge result, and filter the deduplication result to extract the data of the previous day to obtain the paper information entity table.
5. The method for constructing a large-scale relational graph in near real time according to claim 1, wherein The extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; Merging the citation relationship table and the paper information entity table includes: Filter the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day; Process the citation relationship information of the previous day for table structure to obtain the citing paper ID and the cited paper ID, and rename them as point identifiers; After processing the citing paper ID and the cited paper ID using the all-merge operator, use the data set operator to aggregate the paper point identifiers and the update time to obtain the daily paper address entity table; For the daily paper address entity table, set the sorting field to 1 through the add field operator, and rename the paper ID field as the display value; After processing the paper information entity table through data filtering and the add field operator, extract the data of the previous day and set the sorting field to 3; After processing the historically formed paper graph entity table through data filtering and the add field operator, extract the data of the previous day and set the sorting field to 2; Perform an all-merge on the daily paper address entity table with the sorting field set, and the data of the previous day extracted from the paper information entity table and the historically formed paper graph entity table using the all-merge operator; Perform deduplication on the point identifiers for the obtained all-merge result through the data deduplication operator, sort according to the sorting field, and then filter out the data with the sorting field not equal to 2 through the data filtering operator to obtain the paper graph entity table; 6. The method for constructing a large-scale relational graph in near real time according to claim 1, characterized in that For dynamic data, model the stock data according to the point, edge, and attribute data standards, and perform real-time processing on the incremental data to form a dynamic point-edge hive table, including: For dynamic data, model the stock data according to the point, edge, and attribute data standards, and obtain all the stock data through the model; construct a data governance model according to the partition field to process the incremental data, and form different types of dynamic point-edge hive tables; 7. The method for constructing a large-scale relational graph in near real time according to claim 1, wherein The method further includes: optimizing the data governance model through machine learning and deep learning methods; 8. An apparatus for quasi-real-time construction of a large-scale relational graph, characterized in that The device includes: A relational graph standard design module for designing point, edge, and attribute data standards; A graph data construction module for, for static data, processing the static data according to the point, edge, and attribute data standards to construct corresponding point models, edge models, and point-edge models for the paper citation data to obtain static point and edge data, and generating different types of static point-edge hive tables; for dynamic data, modeling the stock data according to the point, edge, and attribute data standards, performing real-time processing on the incremental data to form a dynamic point-edge hive table; through serial task scheduling, running the incremental data governance model at regular intervals to process the incremental data and automatically update the associated data; importing the static point-edge hive table and the dynamic point-edge hive table into the hugegraph graph database at a preset moment to obtain citation graph data; A data governance model construction module for extracting edge information and point information from the paper citation data, and merging the edge information and the point information to form complete graph model data; 9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the large-scale relational graph quasi-real-time construction method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for constructing a large-scale relational graph in near real time according to any one of claims 1 to 7.
Citation Information
Patent Citations
6G communication network-oriented super-large-scale spectrum knowledge graph construction method
CN113992288A
Knowledge graph construction method and system for large-scale mass data
CN114297173A
Indexing methods, systems, and computer equipment for ultra-large-scale knowledge graph storage
CN114936296A
Method for constructing and optimizing figure relation graph based on multi-source data fusion
CN117992555A
Multi-source heterogeneous data driven domain knowledge graph construction system and method
CN116204660A