Large-scale relation graph quasi-real-time construction method and device, equipment and storage medium
Through the design of data standards and multi-source data fusion methods, a relationship map of 100 billion scale was built, which solved the problems of data inconsistency and low processing efficiency in the existing technology, real-time construction and data fusion were realized, and the needs of efficiency and real-time were met.
Patent Information
- Application Number
- CN202510561844.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing technology has problems such as data inconsistency, low processing efficiency, and difficulty in data governance when building a hyper-large-scale relationship map, which is difficult to meet the requirements of real-time and efficientness.
A quasi-real-time construction method of a 100 billion-scale relationship map based on multi-source data fusion is proposed. By designing point, edge, and attribute data standards, static and dynamic data are processed, static and dynamic point-edge hive tables are formed, and graph data is imported through hugegraph graph database, thereby realizing real-time fusion and update of data.
It significantly shortens the construction time, meets real-time requirements, supports the continuous growth of data scale, stable performance, reduces system maintenance costs, and realizes data fusion of different data sources.
Smart Images

Figure CN120086387A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large-scale relational graph construction, and particularly to a method, device, equipment, and storage medium for quasi-real-time construction of a large-scale relational graph. Background Art
[0002] With the explosive growth of data, the knowledge graph, as a semantic network revealing the relationships between entities, provides a more effective way for expressing, organizing, managing, and utilizing massive, heterogeneous, and dynamic big data on the Internet, making the network more intelligent and closer to human cognitive thinking. The applications of the knowledge graph are becoming increasingly widespread, and efficient knowledge graph construction and management technologies are indispensable in fields such as social networks, financial risk control, and intelligent recommendation.
[0003] As the data type and scale continue to increase, traditional graph construction methods are inefficient in ultra-large-scale knowledge graphs, and data governance methods and graph construction methods need to be optimized. The present invention proposes a method for quasi-real-time construction of a relational graph of hundreds of billions of scale based on multi-source data fusion. This method models and governs multi-source data through a data modeling platform to ensure data consistency and integrity, and improve the efficiency, accuracy, and real-time performance of graph construction.
[0004] Existing ultra-large-scale graph construction mainly relies on modifying graph databases or retrieval methods. Document 1 (Chinese patent application with application number CN202210874965.8) proposes an intelligent indexing method and system for an ultra-large-scale knowledge graph based on deep learning intelligent hashing. High-efficiency indexing is achieved through a deep learning model to improve retrieval efficiency. The specific steps include dividing the index input into entity, relationship triples, and attribute triples, encoding the input using a BERT-compatible model, and regressing the starting position of data storage and the length of physical storage through an aggregation network and a multi-layer perceptron. It is applicable to the storage of large-scale semantic knowledge graphs, improves retrieval efficiency, and realizes simple retrieval, complex multi-hop retrieval, and complex analysis with knowledge reasoning. Document 2 (Chinese patent application with application number: CN202110677218.0) proposes a distributed cluster framework for a graph database based on docker-compose technology, as well as a method for joint storage and retrieval of a graph database and a document database. Through distributed storage, indexing, and calculation, the efficiency of knowledge graph construction and retrieval under the background of large-scale massive data is improved. It effectively solves the problem that a single-machine graph database cannot meet the storage and retrieval requirements of massive data, and first proposes a plug-and-play distributed cluster framework for a graph database. The proposed heterogeneous database storage and retrieval solution significantly improves the overall retrieval efficiency and alleviates the problem of reduced retrieval performance by super nodes in the graph database.
[0005] The construction of a relational graph has two problems. One is that it targets a single data source, and the other is that the data volume is small, and it is basically impossible to process ultra-large-scale data regularly. Document 3 (a Chinese invention patent application with the application number CN202410252439.7) provides a method for constructing and optimizing a person relational graph based on multi-source data fusion. The steps include real-time collecting and generating basic person information, preprocessing the basic person information, using machine learning to identify and extract the preprocessed basic person information to obtain person relationship data, constructing a person relational graph based on the person relationship data, and optimizing and updating the person relational graph based on the real-time obtained basic person information. Integrating multiple data sources provides comprehensive person relationship information and enhances the accuracy and reliability of the graph. Through data preprocessing, machine learning, and knowledge graph technology, the efficiency and quality of constructing a person relational graph are improved. Document 4 (a Chinese invention patent application with the application number: CN202111268566.9) provides a method for constructing an ultra-large-scale spectrum knowledge graph for a 6G communication network. It includes spectrum knowledge graph ontology construction, spectrum data acquisition, spectrum knowledge extraction, spectrum knowledge fusion, spectrum knowledge quality assessment, spectrum knowledge graph construction and storage, spectrum knowledge reasoning, and spectrum knowledge graph update. From the perspective of wireless communication, it provides a framework for the core technology to achieve intelligent spectrum resource management. The constructed spectrum knowledge graph has high quality, strong logic between nodes, and can perform efficient spectrum knowledge query and utilization.
[0006] The existing technologies have the following deficiencies: Data inconsistency: During the fusion process of multi-source data, the data formats and qualities of different data sources vary greatly, resulting in data inconsistency, which affects the accuracy and reliability of the graph.
[0007] Low processing efficiency: Processing large-scale data requires a large amount of computing resources and time. The existing technologies have bottlenecks in processing efficiency and cannot meet the requirements of real-time and high efficiency.
[0008] Difficult data governance: Governing and maintaining large-scale data requires complex technical means. The existing technologies have deficiencies in data governance, resulting in high management and maintenance costs of data.
[0009] In the process of constructing a relational graph from paper citation data, the data sources are obtained from different platforms such as VIP, Wanfang, and CNKI, and the data formats and contents are different. How to achieve both timely construction and generation of respective graph data for various data sources and fusion of similar data in different data sources is a technical problem. Summary of the Invention
[0010] Based on this, it is necessary to provide a method, device, equipment, and storage medium for quasi-real-time construction of a large-scale relational graph to address the above technical problems.
[0011] A quasi-real-time construction method for large-scale relational graphs, the method comprising: Designing data standards for points, edges, and attributes.
[0012] For static data, according to the data standards for points, edges, and attributes, construct corresponding point models, edge models, and point-edge models for paper citation data to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables.
[0013] For dynamic data, model the existing data according to the data standards for points, edges, and attributes, and perform real-time processing on the incremental data to form a dynamic point-edge hive table.
[0014] Through serial task scheduling, the incremental data governance model is run regularly to process the incremental data and automatically update the associated data.
[0015] At a preset moment, import the static point-edge hive table and the dynamic point-edge hive table into the HugeGraph graph database to obtain citation graph data.
[0016] Extract edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merge the citation relationship table and the paper information entity table to form complete graph model data.
[0017] In one embodiment, designing data standards for points, edges, and attributes includes: defining a corresponding number of entity points and a number of point-edge relationships for multiple data sources, setting the storage total to the TB level, setting the number of entity points to tens of billions, and setting the number of relationship edges to hundreds of billions.
[0018] The identification of points is represented in the form of an identification prefix + value; point attributes include point identification, point type, display value, update time, and partition field; edge attributes include subcategory, starting point, ending point, update time, remarks, and partition field.
[0019] In one embodiment, the specific process of extracting edge information from paper citation data includes: Filter the paper citation data in the data modeling platform to extract the data of the previous day.
[0020] Filter the data of the previous day to obtain the valid cited paper IDs.
[0021] Through SQL operators, split the cited paper ID from one piece of data into multiple pieces, with each paper ID being one piece of data.
[0022] Through data aggregation operators, perform aggregation counting on the same cited paper ID and the same citing paper ID on the same day.
[0023] The aggregated results are processed using a table structure processing operator to generate a citation relationship table with the fine category being reference, the starting point being the citing party's paper ID, the ending point being the cited party's paper ID, and the update time being the citation date.
[0024] In one embodiment, the specific process of extracting point information from paper citation data includes: Filter the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information.
[0025] Deduplicate the paper information of the previous day by paper ID using a data deduplication operator.
[0026] For the removal result, use an add field operator to add 'paper_' in front of the paper ID to construct a point identifier, with the point type being paper and the display value being the paper title.
[0027] Use an all merge operator to merge the point identifier and the paper information entity table formed historically.
[0028] For the merged result, use a deduplication operator to deduplicate the point identifier and the display value, and perform data filtering on the deduplication result to extract the data of the previous day to obtain the paper information entity table.
[0029] In one embodiment, the extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; merging the citation relationship table and the paper information entity table includes: Filter the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day.
[0030] Perform table structure processing on the citation relationship information of the previous day to obtain the citing party's paper ID and the cited party's paper ID, and rename them as point identifiers.
[0031] After processing the citing party's paper ID and the cited party's paper ID using an all merge operator, use a data set operator to aggregate the paper point identifier and the update time to obtain the daily paper address entity table.
[0032] For the daily paper address entity table, use an add field operator to set the sorting field to 1 and rename the paper ID field as the display value.
[0033] After processing the paper information entity table through data filtering and an add field operator, extract the data of the previous day and set the sorting field to 3.
[0034] After processing the historically formed paper graph entity table through data filtering and an add field operator, extract the data of the previous day and set the sorting field to 2. For the daily paper address entity table with sorting fields set, all the data extracted from the paper information entity table and the paper graph entity table formed historically on the previous day is fully merged using the full merge operator.
[0035] For the obtained full merge result, the point identifiers are de-duplicated through the data de-duplication operator, sorted according to the sorting fields, and then the data with sorting fields not equal to 2 is filtered out through the data filtering operator to obtain the paper graph entity table.
[0036] In one embodiment, for dynamic data, the stock data is modeled according to the point, edge, and attribute data standards, and the incremental data is processed in real time to form a dynamic point-edge hive table, including: For dynamic data, the stock data is modeled according to the point, edge, and attribute data standards, and all the data of the stock is obtained through the model; a data governance model is constructed according to the partition fields for incremental data processing to form different types of dynamic point-edge hive tables.
[0037] In one embodiment, the method further includes: optimizing the data governance model through machine learning and deep learning methods.
[0038] A large-scale relational graph quasi-real-time construction device, the device includes: A relational graph standard design module for designing point, edge, and attribute data standards.
[0039] A graph data construction module for, for static data, processing the static data according to the point, edge, and attribute data standards to construct corresponding point models, edge models, and point-edge models for the paper citation data to obtain static point and edge data, and generating different types of static point-edge hive tables; for dynamic data, modeling the stock data according to the point, edge, and attribute data standards, processing the incremental data in real time to form a dynamic point-edge hive table; through serial task scheduling, running the incremental data governance model at regular intervals to process the incremental data and automatically update the associated data; and importing the static point-edge hive table and the dynamic point-edge hive table into the hugegraph graph database at a preset moment to obtain citation graph data.
[0040] A data governance model construction module for extracting edge information and point information from the paper citation data and merging the edge information and point information to form complete graph model data.
[0041] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0042] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0043] The above large-scale relational graph quasi-real-time construction method, device, equipment and storage medium, the method includes: designing point, edge, and attribute data standards; constructing static point-edge hive tables and dynamic point-edge hive tables through the separate processing of static data and dynamic data, and the real-time processing of incremental data; importing the static point-edge hive tables and dynamic point-edge hive tables into the HugeGraph graph database at a preset moment to obtain citation graph data; extracting edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merging the citation relationship table and the paper information entity table to form complete graph model data. This method significantly shortens the construction time when processing data on the scale of hundreds of billions, meets the real-time requirement; supports the continuous growth of the data scale and has stable performance; can not only construct and generate their respective graph data from various data sources in a timely manner, but also fuse the same type of data in different data sources. Description of the Drawings
[0044] Figure 1 It is a schematic flowchart of the large-scale relational graph quasi-real-time construction method in an embodiment; Figure 2 It is a flowchart of edge extraction of citation data in another embodiment; Figure 3 It is a flowchart of point extraction of citation data in another embodiment; Figure 4 It is a flowchart of edge-point merging of citation data in another embodiment; Figure 5 It is a structural block diagram of the large-scale relational graph quasi-real-time construction device in an embodiment; Figure 6 It is an internal structural diagram of a computer device in an embodiment. Detailed Embodiments
[0045] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0046] In one embodiment, as Figure 1 shown, a large-scale relational graph quasi-real-time construction method is provided, and the method includes the following steps: Step 100: Design point, edge, and attribute data standards.
[0047] Step 102: For static data, construct corresponding point models, edge models, and point-edge models for paper citation data according to point, edge, and attribute data standards to process the static data, obtain static point and edge data, and generate different types of static point-edge hive tables.
[0048] Specifically, for static data: through different data sources, build specific point models, edge models, and point-edge models to process the data, obtain static point and edge data, and form different types of static point-edge hive tables.
[0049] Among them, each type of point or edge has multiple different data sources, such as retrieval platforms such as VIP, Wanfang, and CNKI, which should be considered to be merged together when constructing graph data.
[0050] For paper citation data, the data source is the paper and citation information obtained from different databases such as VIP, Wanfang, and CNKI. Each point in the graph is a paper, and the point model is to build models for different databases to form point (paper) data. Each edge in the graph represents the citation relationship between papers, and the edge model is to build models for different databases to form edge (citation relationship) data. The point-edge model is to model the same type of points (or edges), and merge the model results built by different databases to form a type of point (or edge) model.
[0051] Among them, static data refers to data that is no longer updated, such as paper citation data purchased / acquired once.
[0052] Step 104: For dynamic data, the stock data is modeled according to the point, edge, and attribute data standards, and the incremental data is processed in real time to form a dynamic point-edge hive table.
[0053] Specifically, for existing data, we build a data governance model to obtain all existing data to ensure the integrity and consistency of the initial data. For incremental data, we build a data governance model based on the partition field to process incremental data to ensure the real-time and accuracy of incremental data. Finally, different types of dynamic point-edge hive tables are formed.
[0054] Dynamic data refers to data sources that are constantly updated, such as the current mainstream paper retrieval databases. The citation data in this type of data source is both historically accumulated and continuously increasing every day. Stock data refers to all data from now on. For stock data, a one-time model is built to form point and edge data.
[0055] Incremental data refers to the new paper citation data that is continuously acquired every day. For incremental data, we build an incremental model, which runs once a day. Each time it runs, it takes the difference set with the existing data to form the incremental point and edge data of the day.
[0056] By separately processing static data and dynamic data, as well as real-time processing of incremental data, the efficiency and real-time performance of graph construction are improved.
[0057] Step 106: Through serial task scheduling, the incremental data governance model is run regularly to process incremental data and automatically update the associated data.
[0058] Specifically, a data governance model for incremental data is constructed. Through serial task scheduling, the constructed incremental data governance model is run regularly to process incremental data and automatically update the associated data, ensuring data consistency and integrity.
[0059] All models are data governance models. The incremental data governance model needs to be run regularly every day. When there are multiple incremental data governance models, task scheduling should be carried out in a certain order to ensure that each data governance model can run once a day. Some of these models have dependencies, so when performing serial task scheduling, the order of running the data governance models needs to be designed well.
[0060] Step 108: At a preset moment, the static point-edge hive table and the dynamic point-edge hive table are imported into the HugeGraph graph database to obtain citation graph data.
[0061] Specifically, the well-governed point-edge hive table is imported into the HugeGraph graph database through a written script program every day at a fixed time.
[0062] Step 110: Edge information and point information are extracted from the paper citation data to obtain a citation relationship table and a paper information entity table, and the citation relationship table and the paper information entity table are merged to form complete graph model data.
[0063] Specifically, multi-source data is modeled and governed through a data modeling platform to ensure data consistency and integrity, and improve the accuracy and reliability of the graph. Through automated and visual data governance tools, the data governance process is simplified, and the data management and maintenance costs are reduced.
[0064] In the above method for quasi-real-time construction of large-scale relational graphs, the method includes: designing data standards for points, edges, and attributes; constructing static point-edge Hive tables and dynamic point-edge Hive tables through the separate processing of static data and dynamic data, as well as the real-time processing of incremental data; importing the static point-edge Hive tables and dynamic point-edge Hive tables into the HugeGraph graph database at a preset moment to obtain citation graph data; extracting edge information and point information from the paper citation data to obtain a citation relationship table and a paper information entity table, and merging the citation relationship table and the paper information entity table to form complete graph model data. This method significantly shortens the construction time when processing data on the scale of hundreds of billions, meets the real-time requirements; supports the continuous growth of the data scale with stable performance; through modular design and automated management, the system maintenance cost is significantly reduced, and the user operation is simple. It can not only construct and generate respective graph data for various data sources in a timely manner, but also fuse the same type of data in different data sources.
[0065] In one embodiment, step 100 includes: defining a corresponding number of entity points and a number of point-edge relationships for multiple data sources, setting the total storage capacity to the TB level, setting the number of entity points to tens of billions, and setting the number of relationship edges to hundreds of billions; representing the identifier of a point in the form of an identifier prefix + value; point attributes include point identifier, point type, display value, update time, and partition field; edge attributes include subcategory, start point, end point, update time, remarks, and partition field.
[0066] In one embodiment, the specific process of extracting edge information from the paper citation data in step 110 includes: filtering the paper citation data in the data modeling platform to extract the data of the previous day; filtering the data of the previous day to obtain the valid cited paper IDs; using SQL operators to split the cited paper ID from one piece of data into multiple pieces, with each paper ID being a piece of data; aggregating and counting the same citing paper ID and cited paper ID on the same day through a data aggregation operator; processing the aggregation result using a table structure processing operator to generate a citation relationship table with the subcategory being citation, the start point being the citing paper ID, the end point being the cited paper ID, and the update time being the citation date. The process of extracting citation data edges is as Figure 2 shown.
[0067] In one embodiment, the specific process of extracting point information from the paper citation data in step 110 includes: filtering the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; de-duplicating the paper information of the previous day by the data de-duplication operator according to the paper ID; using the add field operator to add 'paper_' in front of the paper ID to construct a point identifier for the de-duplication result, the point type is a paper, and the display value is the paper title; using the all merge operator to merge the point identifier and the paper information entity table formed by history; using the de-duplication operator to de-duplicate the point identifier and the display value for the merge result, and performing data filtering on the de-duplication result to extract the data of the previous day to obtain the paper information entity table. The citation data point extraction process is as Figure 3 shown.
[0068] In one embodiment, the extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; merging the citation relationship table and the paper information entity table in step 110 includes: filtering the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day; performing table structure processing on the citation relationship information of the previous day to obtain the citing paper ID and the cited paper ID, and renaming them as point identifiers; using the all merge operator to process the citing paper ID and the cited paper ID, and then using the data set operator to aggregate the paper point identifiers and the update time to obtain the daily paper address entity table; using the add field operator for the daily paper address entity table to set the sorting field to 1 and rename the paper ID field as the display value; after processing the paper information entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 3; after processing the paper graph entity table formed by history through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 2; using the all merge operator to perform all merges on the daily paper address entity table with the sorting field set, and the data of the previous day extracted from the paper information entity table and the paper graph entity table formed by history; using the data de-duplication operator to de-duplicate the point identifiers for the obtained all merge result, sorting according to the sorting field, and then filtering out the data with the sorting field not equal to 2 through the data filtering operator to obtain the paper graph entity table. The citation data edge generating point merging process is as Figure 4 shown.
[0069] Specifically, taking the citation data as an example, the edge extraction, point extraction, and edge generating point merging through the data modeling platform are introduced respectively.
[0070] (1)Edge extraction: For citation data, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the citation data of the previous day; through the data filtering operator, extract the valid citing party paper IDs; through the SQL operator, split the cited party paper ID from one piece of data into multiple pieces, with each paper ID being a piece of data; through the data aggregation operator, aggregate and count the same citing party paper IDs and cited party paper IDs on the same day; through the table structure processing operator, finally form a citation relationship table with the fine category being citation, the starting point being the citing party paper ID, the ending point being the cited party paper ID, and the update time being the citation date.
[0071] (2)Node extraction: For paper information data, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the paper information of the previous day, including the paper ID and the corresponding paper information; through the data deduplication operator, deduplicate by paper ID; through the add field operator, add 'paper_' in front of the paper ID to construct the node identifier, with the node type being paper and the display value being the paper title; through the all merge operator, merge with the paper information entity table formed historically; through the data deduplication operator, aggregate the node identifier and the display value; through the data filtering operator, take the data of the previous day to obtain the paper information entity table.
[0072] (3)Edge-node merging: For the citation relationship table formed in the first step, in the data modeling platform, through the data filtering operator, set the partition field = T - 1 to extract the citation relationship information of the previous day; through table structure processing, obtain the citing party paper ID and the cited party paper ID, and rename them as the node identifier; through the all merge operator and the data aggregation operator, aggregate the paper node identifier and the update time to form the daily paper address entity table; through the add field operator, set the sorting field to 1 and rename the paper ID field as the display value; through the data filtering and add field operators, respectively extract the data of the previous day from the paper information entity table formed in step two and set the sorting field to 3, and extract the data of the previous day from the historically formed paper graph entity table and set the sorting field to 2; through the all merge operator, merge the three types of data; through the data deduplication operator, deduplicate the node identifier; sort by the sorting field; through the data filtering operator, filter out the data with the sorting field not equal to 2, and finally obtain the paper graph entity table.
[0073] In one embodiment, step 104 includes: For dynamic data, model the stock data according to the node, edge, and attribute data standards, and obtain all the stock data through the model; construct a data governance model according to the partition field for incremental data processing to form different types of dynamic node-edge hive tables.
[0074] In one embodiment, the method further includes: optimizing the data governance model through machine learning and deep learning methods to improve the intelligent level of data processing.
[0075] This method can expand the data sources: it can support more data source types, further improving the diversity and richness of data.
[0076] The relationship graph constructed by using this method can be used in fields such as financial risk control, intelligent recommendation, and social network analysis. Specifically: Financial risk control: It is used for risk control in the financial field, and potential risk points are discovered through graph analysis.
[0077] Intelligent recommendation: It is used for the construction of a recommendation system, and personalized recommendation services are provided through graph analysis.
[0078] Social network analysis: It is used for the analysis of social networks, and key nodes and community structures in social networks are discovered through graph analysis.
[0079] It should be understood that although Figure 1 the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0080] In one embodiment, as Figure 5 shown, a large-scale relationship graph quasi-real-time construction device is provided, including: a relationship graph standard design module, a graph data construction module, and a data governance model construction module, where: The relationship graph standard design module is used to design the standards for point, edge, and attribute data.
[0081] The graph data construction module is used for static data. According to the point, edge, and attribute data standards, it constructs corresponding point models, edge models, and point-edge models for the paper citation data to process the static data, obtaining static point and edge data, and generating different types of static point-edge hive tables. For dynamic data, it models the stock data according to the point, edge, and attribute data standards and processes the incremental data in real time to form dynamic point-edge hive tables. Through serial task scheduling, it runs the incremental data governance model at regular intervals to process the incremental data and automatically update the associated data. At a preset moment, it imports the static point-edge hive table and the dynamic point-edge hive table into the HugeGraph graph database to obtain citation graph data.
[0082] The data governance model construction module is used to extract edge information and point information from the paper citation graph data and merge the edge information and point information to form complete graph model data.
[0083] In one embodiment, the relationship graph standard design module is further used to define a corresponding number of entity points and a corresponding number of point-edge relationships for multiple data sources, set the storage total to the TB level, set the number of entity points to the tens of billions level, and set the number of relationship edges to the hundreds of billions level. The identifier of the point is represented in the form of an identifier prefix + value. The point attributes include point identifier, point type, display value, update time, and partition field. The edge attributes include subcategory, start point, end point, update time, remarks, and partition field.
[0084] In one embodiment, the specific process of extracting edge information from the paper citation data in the data governance model construction module includes: filtering the paper citation data in the data modeling platform to extract the data of the previous day; filtering the data of the previous day to obtain the valid cited party paper IDs; using SQL operators to split the cited party paper IDs from one piece of data into multiple pieces, with each paper ID being a piece of data; aggregating and counting the same citing party paper IDs and cited party paper IDs on the same day through data aggregation operators; processing the aggregation result using a table structure processing operator to generate a citation relationship table with the subcategory being citation, the start point being the citing party paper ID, the end point being the cited party paper ID, and the update time being the citation date.
[0085] In one embodiment, the specific process of extracting point information from the paper citation data in the data governance model construction module includes: filtering the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; de-duplicating the paper information of the previous day by the paper ID using the data de-duplication operator; using the add field operator to add 'paper_' in front of the paper ID to construct a point identifier for the removal result, the point type is a paper, and the display value is the paper title; using the all merge operator to merge the point identifier and the paper information entity table formed by history; using the de-duplication operator to de-duplicate the point identifier and the display value for the merge result, and performing data filtering on the de-duplication result to extract the data of the previous day to obtain the paper information entity table.
[0086] In one embodiment, the extracted edge information is the citation relationship table, and the extracted point information is the paper information entity table; the data governance model construction module merges the citation relationship table and the paper information entity table, including: filtering the citation relationship table in the data modeling platform to extract the citation relationship information of the previous day; performing table structure processing on the citation relationship information of the previous day to obtain the citing paper ID and the cited paper ID, and renaming them as point identifiers; using the all merge operator to process the citing paper ID and the cited paper ID, and then using the data set operator to aggregate the paper point identifier and the update time to obtain the daily paper address entity table; using the add field operator for the daily paper address entity table to set the sorting field to 1 and rename the paper ID field as the display value; after processing the paper information entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 3; after processing the historically formed paper graph entity table through data filtering and the add field operator, extracting the data of the previous day and setting the sorting field to 2; using the all merge operator to perform all merges on the daily paper address entity table with the sorting field set, and the data of the previous day extracted from the paper information entity table and the historically formed paper graph entity table; using the data de-duplication operator to de-duplicate the point identifier for the obtained all merge result, sorting according to the sorting field, and then filtering out the data with the sorting field not equal to 2 through the data filtering operator to obtain the paper graph entity table.
[0087] In one embodiment, the graph data construction module is further configured to, for dynamic data, model the stock data according to the point, edge, and attribute data standards, and obtain all the stock data through the model; construct a data governance model according to the partition field to process the incremental data, and form different types of dynamic point-edge hive tables.
[0088] In one embodiment, the device further includes an optimization module for optimizing the data governance model through machine learning and deep learning methods.
[0089] For the specific limitations of the large-scale relational graph quasi-real-time construction device, reference can be made to the limitations of the large-scale relational graph quasi-real-time construction method in the above text, which will not be elaborated here. Each module in the above large-scale relational graph quasi-real-time construction device can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0090] In one embodiment, a computer device is provided. This computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a large-scale relational graph quasi-real-time construction method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0091] Those skilled in the art can understand that Figure 6 the structure shown in
[0092] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0093] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above method embodiment.
[0094] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0095] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0096] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for constructing a large-scale relationship graph in quasi-real time, characterized in that: The method comprises: Design point, edge, and attribute data standards; For static data, we construct corresponding point models, edge models, and point-edge models for paper citation data according to the point, edge, and attribute data standards to process static data, obtain static point and edge data, and generate different types of static point-edge hive tables; For dynamic data, we model the stock data according to the point, edge, and attribute data standards, process the incremental data in real time, and form a dynamic point-edge hive table; Through serial task orchestration, the incremental data governance model is run regularly to process incremental data and automatically update related data. Import the static point-edge hive table and the dynamic point-edge hive table into the hugegraph database at a preset time to obtain the citation graph data; Edge information and point information are extracted from the paper citation data to obtain the citation relationship table and the paper information entity table, which are then merged to form a complete graph model data.
2. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: Design point, edge, and attribute data standards, including: Define a number of entity points and a number of point-edge relationships corresponding to various data sources, set the total storage volume to TB level, set the number of entity points to tens of billions, and set the number of relationship edges to hundreds of billions; The point identification is expressed in the form of identification prefix + value; point attributes include point identification, point type, display value, update time, and partition field; edge attributes include sub-category, starting point, end point, update time, remarks, and partition field.
3. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: The specific process of extracting side information from paper citation data includes: Filter the paper citation data in the data modeling platform and extract the data from the previous day; Filter the data from the previous day to obtain the valid citing paper ID; SQL operators are used to split the cited paper ID from one piece of data into multiple pieces, with each paper ID being one piece of data. The same citing paper ID and cited paper ID on the same day are aggregated and counted through the data aggregation operator; The aggregation results are processed using a table structure processing operator to generate a citation relationship table with the subcategory being citation, the starting point being the citing party's paper ID, the ending point being the cited party's paper ID, and the update time being the citation date.
4. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: The specific process of extracting point information from paper citation data includes: Filter the paper information data in the data modeling platform to extract the paper information of the previous day; the paper information of the previous day includes the paper ID and the corresponding paper information; The paper information of the previous day is deduplicated by paper ID using the data deduplication operator; For the removed results, use the Add Field operator to add 'paper_' in front of the paper ID to construct a point identifier. The point type is paper, and the displayed value is the paper title. The all-merge operator is used to merge the point identifier and the paper information entity table formed by history; The deduplication operator is used to deduplicate the point identifiers and display values of the merged results, and the deduplication results are filtered to extract the data of the previous day to obtain the paper information entity table.
5. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: The extracted edge information forms a citation relationship table, and the extracted point information forms a paper information entity table; Merge the citation relationship table and the paper information entity table, including: Performing data filtering on the citation relationship table in the data modeling platform to extract citation relationship information of the previous day; Process the table structure of the citation relationship information of the previous day to obtain the citing paper ID and the cited paper ID, and rename them as point identifiers; After processing the citing paper ID and the cited paper ID with the all-merge operator, the paper point identifier and update time are aggregated with the data set operator to obtain the paper address entity table for each day. For the daily paper address entity table, add a field operator, set the sort field to 1, and rename the paper ID field to display value; After processing the paper information entity table through data filtering and adding field operators, the data of the previous day is extracted and the sorting field is set to 3; After filtering the data and adding field operators to the historical paper graph entity table, the data from the previous day is extracted and the sorting field is set to 2. The daily paper address entity table with sorting fields set, the data of the previous day extracted from the paper information entity table and the paper graph entity table formed historically are all merged using the all merge operator; The point identifiers of all the merged results are deduplicated using the data deduplication operator and sorted according to the sorting field. Then, the data whose sorting field is not 2 is filtered out using the data filtering operator to obtain the paper graph entity table.
6. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: For dynamic data, we model the stock data according to the point, edge, and attribute data standards, process the incremental data in real time, and form a dynamic point-edge hive table, including: For dynamic data, the stock data is modeled according to the point, edge, and attribute data standards, and all the stock data is obtained through the model; a data governance model is built according to the partition field to process incremental data and form different types of dynamic point and edge hive tables.
7. The method for constructing a large-scale relationship graph in quasi-real time according to claim 1, characterized in that: The method also includes: optimizing the data governance model through machine learning and deep learning methods.
8. A device for constructing a large-scale relationship graph in quasi-real time, characterized in that: The device comprises: Relationship graph standard design module, used to design point, edge, and attribute data standards; The graph data construction module is used to construct corresponding point models and edge models for paper citation data according to point, edge, and attribute data standards for static data, and to process static data with point-edge models to obtain static point and edge data and generate different types of static point-edge hive tables; for dynamic data, the stock data is modeled according to point, edge, and attribute data standards, and the incremental data is processed in real time to form dynamic point-edge hive tables; through serial task scheduling, the incremental data governance model is scheduled to run, the incremental data is processed, and the associated data is automatically updated; at the preset time, the static point-edge hive table and the dynamic point-edge hive table are imported into the hugegraph graph database to obtain citation graph data; The data governance model building module is used to extract edge information and point information from paper citation data, and merge the edge information and point information to form a complete graph model data.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the quasi-real-time construction method of a large-scale relationship graph described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a large-scale relationship graph in quasi-real time according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
6G communication network-oriented super-large-scale spectrum knowledge graph construction method
CN113992288A
Knowledge graph construction method and system for large-scale mass data
CN114297173A
Indexing methods, systems, and computer equipment for ultra-large-scale knowledge graph storage
CN114936296A
Method for constructing and optimizing figure relation graph based on multi-source data fusion
CN117992555A
Method, medium and equipment for predicting quoted quantity of achievements based on dynamic knowledge graph
CN114817571A